Multi-Dialect Arabic POS Tagging: A CRF Approach

Kareem Darwish; Hamdy Mubarak; Ahmed Abdelali; Mohamed Eldesouki; Younes Samih; Randah Alharbi; Mohammed Attia; Walid Magdy; Laura Kallmeyer

Multi-Dialect Arabic POS Tagging: A CRF Approach

Kareem Darwish

Hamdy Mubarak

Ahmed Abdelali

Mohamed Eldesouki

Younes Samih

Randah Alharbi

Mohammed Attia

Walid Magdy

Laura Kallmeyer

Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), European Language Resources Association (ELRA), Miyazaki, Japan (2018), pp. 93-98

Download Google Scholar

Abstract

This paper introduces a new dataset of POS-tagged Arabic tweets in four major dialects along with tagging guidelines. The data, which we are releasing publicly, includes tweets in Egyptian, Levantine, Gulf, and Maghrebi, with 350 tweets for each dialect with appropriate train/test/development splits for 5-fold cross validation. We use a Conditional Random Fields (CRF) sequence labeler to train POS taggers for each dialect and examine the effect of cross and joint dialect training, and give benchmark results for the datasets. Using clitic n-grams, clitic metatypes, and stem templates as features, we were able to train a joint model that can correctly tag four different dialects with an average accuracy of 89.3%.

Research Areas

Natural language processing

Explore our many areas of focus

Building a collaborative ecosystem

Shaping the future together

Translating discovery into real-world impact

Multi-Dialect Arabic POS Tagging: A CRF Approach

Abstract

Research Areas

Meet the teams driving innovation

Google AI

Google Cloud

Google DeepMind

Google Labs