In Javanese there is no general word for "fall." There is geblak, to fall backwards. Nyungsep, to fall forwards. Ceblok, to fall from above. Goleng, to fall sideways. Learning them taught me two things: that Javanese draws distinctions the languages I was schooled in simply discard, and that we Javanese must be unusually clumsy.
That density is not unique to Javanese. Indonesia holds 728 living languages, close to a tenth of every language spoken on Earth, and each one encodes distinctions its speakers found worth making. Most have never been written down at scale. Almost none of them exist inside the systems that increasingly decide what can be searched, translated, transcribed, or answered.
Indonesiaku is a non-profit research lab building the datasets, models, and open infrastructure to change that. This is the first note; the site went live today.
The constraint is data, not speakers
Javanese has 68 million speakers. Sundanese has 32 million. Buginese has four million. These are not small languages — several of them outrank the national languages of European countries that have mature language technology.
What they lack is text. They are spoken far more than they are typed, and what does get typed is informal, scattered, and unlabelled. Around 90% of Indonesia's languages have no corpus large enough to train a translation model by conventional means, however many people speak them.
The result is easy to state. There is no machine translation for Buginese. There is none for Banjarese, Sasak, or Ngaju. Not from us, not from Google, not from anyone. Four million Buginese speakers meet the digital world in a language that is not their own, or they do not meet it at all.
What we are building
A corpus with provenance. More than 12,000 verified translations across 12 languages, contributed by over 600 speakers and checked by other speakers. No scraping, no synthetic filler. Every row records who wrote it and when, and the whole thing is published under CC BY-SA 4.0.
Models built for scarcity. Standard training recipes assume data we do not have. NusaMT-7B is our attempt at the opposite problem: usable translation quality from thousands of examples rather than millions. The method is in our paper.
Tools that work now. Our translator covers 16 languages of Indonesia. It is deliberately honest about its edges — where a language has no engine behind it, we show the gap rather than fill it with a guess.
Open, or it does not count
An archive that depends on us to survive is not preservation. Every dataset we build is openly licensed, every model public, every note published. If Indonesiaku disappears, the corpus should outlive it. That constraint decides what we build and what we decline to build.
What comes next
The four languages with no translation anywhere are where we go first — they are the clearest case for doing this work at all. Beyond that: more languages into the corpus, better evaluation for the pairs where we still fail, and notes here on what we learn as we go, including the parts that do not work.
If you speak one of these languages, five translated sentences is a real contribution. That is how the corpus was built and how it will keep growing.