ResearcharXiv AI / CL
Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya

Summary
Multilingual pre-trained language models (PLMs) exhibit degraded performance on low-resource, non-Latin-script languages, driven by high out-of-vocabulary (OOV) rates and excessive subword fragmentation that result from Latin-script-centric tokenizer training.
Original Article
Captured source content or English translation, normalized into this reading format.
This story does not yet have captured source text. Open the source link to read it.
Region
Global
Heat Score
76
Category
Research
Language
en
