About the Kodava takk Project
A small team of Kodavas set out to give our mother tongue a digital home — a translator, a dictionary, and lessons where none existed. This is the story of that work, the people behind it, and how you can be part of it.
Background
The Language
Kodava takk (also called Kodava thakk, or the Coorgi language by non-Kodavas) is spoken by roughly 114,000–200,000 people in the Kodagu (Coorg) district of Karnataka, India. It is a Dravidian language related to Kannada, Tulu, Tamil, and Malayalam — yet distinct from all of them.
kfa
South Dravidian I
Kannada, Kodava Lipi
~150,000
Why This Project
Our children are growing up in cities and countries far from Kodagu. Many understand Kodava takk but struggle to speak it fluently. Without tools to learn and practice, we risk losing the language that connects us to our roots.
There was no Kodava dictionary app, no translation tool, no online course — anywhere. So we built them: a free translator, a comprehensive dictionary, grammar guides, and lessons with native-speaker audio.
The Challenge
Building language technology for Kodava takk is unlike building it for Kannada or Hindi. It is what researchers call a low-resource language — and among low-resource languages, it is one of the hardest cases.
Zero existing datasets
When we started, there was not a single Kodava takk dataset on any of the world's NLP platforms — HuggingFace, OPUS, Tatoeba, FLORES-200, NLLB, AI4Bharat. No parallel corpus, no translation model, nothing to build on. Every sentence pair in our system was collected, translated, or verified by our team. Before this project, fewer than 2,500 Kodava words existed in digitally usable form, scattered across a handful of websites.
Almost no digital footprint
Compare us to our nearest neighbour, Tulu — itself considered a minority language:
| Kodava takk | Tulu | |
|---|---|---|
| Speakers | ~150,000 | ~1.8 million |
| Wikipedia | None | Own edition since 2016 |
| Script in Unicode | Not yet | Yes (Tulu-Tigalari, 2024) |
| Digital books & sources | A handful of scattered word lists (~2,500 words before this project) | Digitised 6-volume Tulu Lexicon, online dictionaries, digitised literature |
| Cinema & media | A handful of films | Active film industry, TV |
| Taught in schools | Rarely | Optional subject in coastal Karnataka |
If Tulu is under-resourced, Kodava takk is starting from almost nothing. That is exactly why this work matters.
A language without a settled script
Kodava takk has historically been written in the Kannada script — a borrowed alphabet. The Kodava Lipi was officially adopted in 2022 but is not yet in Unicode, so it cannot be typed, searched, or stored on ordinary devices. Older sources use ad-hoc English spellings that disagree with each other. We standardised on Kannada script plus a consistent Latin transliteration and ISO 15919, but every word we digitise involves choices no dictionary can settle for us.
Sounds the script cannot write
This is the deepest challenge. Kodava takk has vowel sounds that simply do not exist in Kannada — like the short central vowel ë at the end of words such as nīrë (water), which the Kannada script has no letter for. Written in Kannada, these sounds silently disappear. Some of them do exist in Malayalam and Tamil, so our pronunciation work borrows from those languages where it helps — but that introduces its own problems (Tamil, for instance, cannot distinguish sounds like b and p the way we say them). Getting Kodava takk to sound right in a digital tool means stitching together three scripts, and still falling short of a native voice.
Native audio doesn't scale — yet
Because no text source captures true Kodava pronunciation, native-speaker recordings are the only ground truth. Our lessons already carry real native audio, recorded by our team — but a few people cannot record a whole language. Reaching every word and phrase will only happen with active community involvement. That is the goal we are building towards.
Doing with 50,000 what normally takes 4 million
Training a translation model for even a low-resource language typically takes around 4 million sentence pairs. Sourcing that much content for Kodava takk is nearly impossible — it would take roughly 1,000 native speakers volunteering their time. We built our translator with under 50,000 sentence pairs by applying several optimizations, which means some translations will need correction. That is exactly where native speakers can help: every correction from a fluent speaker makes the model better. Reach out to us at info.kodavatakk@gmail.com if you would like to get involved.
Our Story
The project began quietly, about two years ago — collecting sentence pairs, hunting down word lists, and discovering just how little of Kodava takk existed in digital form. Progress was slow and part-time, but the conviction was there: we had to start somewhere.
The founding team set out to build it properly, making time for it alongside work, school, and family. Over the following months we sourced, collected, and digitised dictionary entries word by word, hand-corrected a large corpus of English–Kodava sentence pairs with native speakers, trained the first-ever machine translation models for the language, and launched kodavathakk.org.
The site now holds the largest digital collection of Kodava takk anywhere — a comprehensive dictionary, verified sentence pairs, grammar guides, and audio for pronunciation. The translator combines a purpose-trained model with AI retrieval over everything the team has built.
The foundation was built by a small team. The next chapter — more voices, more corrections, audio for every word — can only be written by the community.
The Team
This project was built by a handful of people who love this language — not by a company, not by a grant, entirely funded by the founders. Everything you see here is a labour of love, built in our spare time.
Founders
The people who started this — building the platform, the datasets, and the translation models around jobs, school, and everything else.
Content Team
Native speakers and language lovers who transcribed the dictionary, corrected sentence pairs, and recorded audio.
Contributors
Community members who have submitted words, corrections, and phrases through the site.
Ambassadors
Kodavas around the world spreading the word — in family groups, kootas, and community events.
How You Can Help
The team laid the foundation. Making it complete — and keeping it alive — needs you.
Translate & Correct
Try translating everyday phrases and fix what doesn't sound right. You know our language best.
Share Words You Know
Remember a word your Thaatha or Avvayya used? Add it before it's forgotten. Every word matters.
Record Your Voice
Native audio is the only true record of how our language sounds. If you speak Kodava takk, your voice is a gift to the next generation.
Spread the Word
Share this with family and friends from Kodagu. The more of us who contribute, the better it gets.
Every word you add, every correction you make, every phrase you record helps preserve our language for the next generation of Kodavas.