A language is more than a dataset.

Work around Kichwa connects endangered-language preservation with translation, dataset quality, model training, and expert validation.

Scarce data makes provenance especially consequential. A model cannot resolve uncertainty about where its examples came from, what they mean, or whose judgment should establish whether an output is useful.

The people who can tell

Expert validation belongs in the process, alongside the technical work. Language preservation asks for attention to the knowledge already held by speakers, not just the patterns that can be extracted from text.

The open question is what responsible machine learning should look like when the dataset is small and the stakes are not.

A working entry. The connections may change.