Advancing AI:
For Every Language
Our research focuses on advancing AI for low-resource languages through the development of high-quality speech and text dataset. We combine interdisciplinary research with community-driven approaches to build technologies that preserve linguistic and cultural diversity.
What We Focus On
Data Curation
We curate high-quality speech and text datasets for African and other low-resource languages, powering NLP tasks such as machine translation, ASR, and speech-to-speech translation, with expert annotation for linguistic and cultural accuracy.
Language Models & Resources
We develop language models, benchmarks, and open resources that bring underrepresented languages and their cultures into modern AI systems, in collaboration with researchers and institutions worldwide.
Projects
YorGe-CS Corpus
A Yoruba–German code-switched machine translation corpus covering five domains: Education, Food, Business, Agriculture, and Sport.
Langture
Publications & Papers
Publications coming soon.
