Empowering Sentiment Analysis in African Low-Resource Languages Through Transformer Models and Strategic Language Selection
Nilanjana Raychawdhary, Sutanu Bhattacharya, Cheryl Seals, Gerry Dozier
Auburn University Auburn University at Montgomery
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
This research addresses the significant challenges of sentiment analysis in low-resource African languages by utilizing advanced transformer-based models to bridge gaps in natural language processing (NLP) for underrepresented linguistic communities. Covering 12 diverse languages: Hausa, Yoruba, Igbo, Nigerian Pidgin, Amharic, Algerian Arabic, Moroccan Arabic Darija, Swahili, Kinyarwanda, Twi, Mozambican Portuguese, and Xitsonga. Our work includes pre-trained language models (PLMs) such as XLM-R, AfroXLMR, AfriBERTa, and mDeBERTaV3, fine-tuning them for the specific task of sentiment classification. In particular, using datasets from the AfriSenti SemEval 2023 Shared Task 12, this process involves tailoring multilingual transformer models to low-resource African languages, enabling them to capture complex linguistic nuances and sentiment polarities effectively. The results demonstrate significant improvements across key metrics, including accuracy, precision, recall, and weighted F1 score. Notably, our work using AfroXLMR achieves the top 1 rank, with a weighted F1 score of 75.8%, outperforming all other submissions. In addition to multilingual transformer models, we evaluated traditional machine learning models, including CNN, Naïve Bayes, and SVM, to compare their performance. The evaluation focused on zero-shot sentiment classification for Tigrinya, a low-resource African language not seen during model training. The results indicate that while traditional models produced meaningful results, multilingual transformer models achieved higher performance in the zero-shot setting. This highlights both the challenges and potential of multilingual transfer approaches for sentiment analysis in African low-resource languages.Moreover, this research highlights the transformative role of Artificial Intelligence (AI) in addressing linguistic diversity, fostering digital inclusion, and enabling reliable sentiment analysis. Furthermore, it lays the groundwork for future advancements in processing low-resource languages and emphasizes the importance of promoting cultural inclusivity through AI technologies.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AITopic Modeling
Natural Language Processing Techniques · Sentiment Analysis and Opinion Mining
参考文献 23
此处列出前 3 条
引用本文 6
按被引量排序,此处列出前 3 条