Navigation auf zora.uzh.ch

Search

ZORA (Zurich Open Repository and Archive)

Encoder-Decoder Methods for Text Normalization

Lusetti, Massimo; Ruzsics, Tatyana; Göhring, Anne; Samardžić, Tanja; Stark, Elisabeth (2018). Encoder-Decoder Methods for Text Normalization. In: Proceedings of the Fifth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2018), Santa Fe, New Mexico, USA, 20 August 2018. Association for Computational Linguistics, 18-28.

Abstract

Text normalization is the task of mapping non-canonical language, typical of speech transcription and computer-mediated communication, to a standardized writing. It is an up-stream task necessary to enable the subsequent direct employment of standard natural language processing tools and indispensable for languages such as Swiss German, with strong regional variation and no written standard. Text normalization has been addressed with a variety of methods, most successfully with character-level statistical machine translation (CSMT). In the meantime, machine translation has changed and the new methods, known as neural encoder-decoder (ED) models, resulted in remarkable improvements. Text normalization, however, has not yet followed. A number of neural methods have been tried, but CSMT remains the state-of-the-art. In this work, we normalize Swiss German WhatsApp messages using the ED framework. We exploit the flexibility of this framework, which allows us to learn from the same training data in different ways. In particular, we modify the decoding stage of a plain ED model to include target-side language models operating at different levels of granularity: characters and words. Our systematic comparison shows that our approach results in an improvement over the CSMT state-of-the-art.

Additional indexing

Item Type:Conference or Workshop Item (Paper), original work
Communities & Collections:06 Faculty of Arts > Institute of Romance Studies
06 Faculty of Arts > Institute of Computational Linguistics
08 Research Priority Programs > Language and Space
Dewey Decimal Classification:000 Computer science, knowledge & systems
410 Linguistics
Language:English
Event End Date:20 August 2018
Deposited On:26 Sep 2018 09:42
Last Modified:06 Apr 2022 15:18
Publisher:Association for Computational Linguistics
OA Status:Green
Official URL:http://aclweb.org/anthology/W18-3902
Project Information:
  • Funder: SNSF
  • Grant ID: CRSII1_160714
  • Project Title: What’s Up, Switzerland? Language, Individuals and Ideologies in mobile messaging.
  • : Project Websitehttps://www.whatsup-switzerland.ch
Download PDF  'Encoder-Decoder Methods for Text Normalization'.
Preview
  • Content: Published Version
  • Language: English
  • Licence: Creative Commons: Attribution 4.0 International (CC BY 4.0)

Metadata Export

Statistics

Citations

Downloads

384 downloads since deposited on 26 Sep 2018
72 downloads since 12 months
Detailed statistics

Authors, Affiliations, Collaborations

Similar Publications