About the Kodava takk Project

A small team of Kodavas set out to give our mother tongue a digital home — a translator, a dictionary, and lessons where none existed. This is the story of that work, the people behind it, and how you can be part of it.

Background

The Language

Kodava takk (also called Kodava thakk, or the Coorgi language by non-Kodavas) is spoken by roughly 114,000–200,000 people in the Kodagu (Coorg) district of Karnataka, India. It is a Dravidian language related to Kannada, Tulu, Tamil, and Malayalam — yet distinct from all of them.

ISO Code

kfa

Family

South Dravidian I

Scripts

Kannada, Kodava Lipi

Speakers

~150,000

Why This Project

Our children are growing up in cities and countries far from Kodagu. Many understand Kodava takk but struggle to speak it fluently. Without tools to learn and practice, we risk losing the language that connects us to our roots.

There was no Kodava dictionary app, no translation tool, no online course — anywhere. So we built them: a free translator, a comprehensive dictionary, grammar guides, and lessons with native-speaker audio.

The Challenge

Building language technology for Kodava takk is unlike building it for Kannada or Hindi. It is what researchers call a low-resource language — and among low-resource languages, it is one of the hardest cases.

1

Zero existing datasets

When we started, there was not a single Kodava takk dataset on any of the world's NLP platforms — HuggingFace, OPUS, Tatoeba, FLORES-200, NLLB, AI4Bharat. No parallel corpus, no translation model, nothing to build on. Every sentence pair in our system was collected, translated, or verified by our team. Before this project, fewer than 2,500 Kodava words existed in digitally usable form, scattered across a handful of websites.

2

Almost no digital footprint

Compare us to our nearest neighbour, Tulu — itself considered a minority language:

Kodava takkTulu
Speakers~150,000~1.8 million
WikipediaNoneOwn edition since 2016
Script in UnicodeNot yetYes (Tulu-Tigalari, 2024)
Digital books & sourcesA handful of scattered word lists (~2,500 words before this project)Digitised 6-volume Tulu Lexicon, online dictionaries, digitised literature
Cinema & mediaA handful of filmsActive film industry, TV
Taught in schoolsRarelyOptional subject in coastal Karnataka

If Tulu is under-resourced, Kodava takk is starting from almost nothing. That is exactly why this work matters.

3

A language without a settled script

Kodava takk has historically been written in the Kannada script — a borrowed alphabet. The Kodava Lipi was officially adopted in 2022 but is not yet in Unicode, so it cannot be typed, searched, or stored on ordinary devices. Older sources use ad-hoc English spellings that disagree with each other. We standardised on Kannada script plus a consistent Latin transliteration and ISO 15919, but every word we digitise involves choices no dictionary can settle for us.

4

Sounds the script cannot write

This is the deepest challenge. Kodava takk has vowel sounds that simply do not exist in Kannada — like the short central vowel ë at the end of words such as nīrë (water), which the Kannada script has no letter for. Written in Kannada, these sounds silently disappear. Some of them do exist in Malayalam and Tamil, so our pronunciation work borrows from those languages where it helps — but that introduces its own problems (Tamil, for instance, cannot distinguish sounds like b and p the way we say them). Getting Kodava takk to sound right in a digital tool means stitching together three scripts, and still falling short of a native voice.

5

Native audio doesn't scale — yet

Because no text source captures true Kodava pronunciation, native-speaker recordings are the only ground truth. Our lessons already carry real native audio, recorded by our team — but a few people cannot record a whole language. Reaching every word and phrase will only happen with active community involvement. That is the goal we are building towards.

6

Doing with 50,000 what normally takes 4 million

Training a translation model for even a low-resource language typically takes around 4 million sentence pairs. Sourcing that much content for Kodava takk is nearly impossible — it would take roughly 1,000 native speakers volunteering their time. We built our translator with under 50,000 sentence pairs by applying several optimizations, which means some translations will need correction. That is exactly where native speakers can help: every correction from a fluent speaker makes the model better. Reach out to us at info.kodavatakk@gmail.com if you would like to get involved.

Our Story

2024 — The seed

The project began quietly, about two years ago — collecting sentence pairs, hunting down word lists, and discovering just how little of Kodava takk existed in digital form. Progress was slow and part-time, but the conviction was there: we had to start somewhere.

Early 2026 — Taking shape

The founding team set out to build it properly, making time for it alongside work, school, and family. Over the following months we sourced, collected, and digitised dictionary entries word by word, hand-corrected a large corpus of English–Kodava sentence pairs with native speakers, trained the first-ever machine translation models for the language, and launched kodavathakk.org.

Today

The site now holds the largest digital collection of Kodava takk anywhere — a comprehensive dictionary, verified sentence pairs, grammar guides, and audio for pronunciation. The translator combines a purpose-trained model with AI retrieval over everything the team has built.

Next — You

The foundation was built by a small team. The next chapter — more voices, more corrections, audio for every word — can only be written by the community.

The Team

This project was built by a handful of people who love this language — not by a company, not by a grant, entirely funded by the founders. Everything you see here is a labour of love, built in our spare time.

Founders

The people who started this — building the platform, the datasets, and the translation models around jobs, school, and everything else.

Coming soon.

Content Team

Native speakers and language lovers who transcribed the dictionary, corrected sentence pairs, and recorded audio.

Coming soon.

Contributors

Community members who have submitted words, corrections, and phrases through the site.

Every correction you submit on the translator makes you one of us. Your name could be here.

Ambassadors

Kodavas around the world spreading the word — in family groups, kootas, and community events.

Share the project with your family and community. Write to us to become an ambassador.

How You Can Help

The team laid the foundation. Making it complete — and keeping it alive — needs you.

1

Translate & Correct

Try translating everyday phrases and fix what doesn't sound right. You know our language best.

2

Share Words You Know

Remember a word your Thaatha or Avvayya used? Add it before it's forgotten. Every word matters.

3

Record Your Voice

Native audio is the only true record of how our language sounds. If you speak Kodava takk, your voice is a gift to the next generation.

4

Spread the Word

Share this with family and friends from Kodagu. The more of us who contribute, the better it gets.

Every word you add, every correction you make, every phrase you record helps preserve our language for the next generation of Kodavas.