CLRLC Logo

Advancing AI:
For Every Language

YorubaHausaIgboSwahiliAmharicZuluTwiWolofShonaSomaliOromoXhosaQuechuaGuaraniNahuatlMaoriCherokeeTagalogUyghurTibetanMaithiliSinhalaKhmerLaoFulaBambaraTigrinyaKinyarwanda

Our research focuses on advancing AI for low-resource languages through the development of high-quality speech and text dataset. We combine interdisciplinary research with community-driven approaches to build technologies that preserve linguistic and cultural diversity.

What We Focus On

Data Curation

We curate high-quality speech and text datasets for African and other low-resource languages, powering NLP tasks such as machine translation, ASR, and speech-to-speech translation, with expert annotation for linguistic and cultural accuracy.

Language Models & Resources

We develop language models, benchmarks, and open resources that bring underrepresented languages and their cultures into modern AI systems, in collaboration with researchers and institutions worldwide.

Projects

Dataset / Machine Translation

YorGe-CS Corpus

A Yoruba–German code-switched machine translation corpus covering five domains: Education, Food, Business, Agriculture, and Sport.

5 DomainsComing Soon

Langture

Coming Soon

Publications & Papers

Publications coming soon.