Advancing Large Language Models for Low-Resource Languages: A Systematic Review of Pretraining, Adaptation, and Ethical Challenges

Citation

Hossain, Ismail and Banik, Mridul and Farid, Fahmid Al and Uddin, Jia and Abdul Karim, Hezerul (2026) Advancing Large Language Models for Low-Resource Languages: A Systematic Review of Pretraining, Adaptation, and Ethical Challenges. Computer Modeling in Engineering & Sciences, 148 (1). pp. 1-10. ISSN 1526-1506

[img] Text
TSP_CMES_75507.pdf - Published Version
Restricted to Repository staff only

Download (6MB)

Abstract

In recent years, the rapid advancement of Large Language Models (LLMs) has significantly transformed natural language processing (NLP), enabling impressive performance across a wide range of tasks. However, these developments have largely benefited high-resource languages, leaving many low-resource and underrepresented languages at risk of further digital marginalization. Addressing this imbalance is crucial to building more inclusive and culturally sustainable AI systems, which is motivating growing research interest in adapting LLMs for linguistically diverse and resource-scarce communities. This systematic review examines recent progress (2020–2025) in the pretraining and adaptation of LLMs for Low-Resource Languages (LRLs). Analysed 812 records obtained in the large databases and using PRISMA criteria, 140 core studies were identified. The innovations in data augmentation and parameter-efficient fine-tuning approaches can be outlined in this selection process. It combines major innovations on data-driven augmentation, parameter-efficient fine-tuning and morphologically rich and underrepresented language script-sensitive tokenization. The results highlight the growing effectiveness of culturally aware standards such as IrokoBench and BLEnD and show that approaches to lightweight adaptation eliminate high computational costs while maintaining language accuracy. The review focuses on the ethics in AI practice, the development of corpora through communities, and interdisciplinary research collaboration among computational linguists, social scientists, and digital humanists. The task of generating a diversified dataset, typology-conscious modelling strategies, and open-source multilingual benchmarks should be prioritized in future research as one possible solution to the existing digital language gap worldwide.

Item Type: Article
Uncontrolled Keywords: LLMs, low-resource languages (LRLs
Subjects: T Technology > T Technology (General)
Divisions: Faculty of Artificial Intelligence & Engineering (FAIE)
Depositing User: Ms Rosnani Abd Wahab
Date Deposited: 04 Aug 2026 01:08
Last Modified: 04 Aug 2026 01:08
URII: http://shdl.mmu.edu.my/id/eprint/16444

Downloads

Downloads per month over past year

View ItemEdit (login required)