KLaSAn
Corpus · Analytics

About corpus analytics

How we interpret language resource data — what we measure, why it matters, and where to see live numbers on this site.

Corpus analytics is how we interpret language resource data — not just counts, but completeness, gaps, growth, and readiness.

Detailed operational analytics (contributor workflows, audit trails, export tools) live in the internal admin dashboard. On this public site you can read how we analyze the language and view live status numbers on the Observatory.

Completeness

How much of each resource type exists compared to its target ceiling.

  • Eight public dimensions: core vocabulary, grammar, audio, examples, translations, phonology, cultural terms, and verb paradigms.
  • Each dimension uses LaSAn scale rules — logarithmic for counts, linear for coverage.
  • Gap analysis highlights categories furthest from their documentation targets.

Growth

Whether documentation is expanding over time.

  • Tracks new dictionary entries, audio clips, sentences, and translations month over month.
  • Growth rate shows if effort is accelerating, steady, or slowing.
  • Used alongside the Impact Radar to see which levers are moving.

Language health

A composite view of how well Kifuliiru is documented and supported.

  • Language Resource Index (LRI) — average score across 68 LaSAn resource units.
  • Status bands from Critically Underdocumented to Fully Digitized (0–100).
  • Dimension scores cover lexicon, audio, documentation, verification, and digital presence.

NLP & grammar readiness

Signals that the corpus is usable for tools, not just human reading.

  • Verification rates, tense coverage, and structured metadata quality.
  • NLP readiness — whether there is enough labeled data for taggers, ASR, or MT.
  • Orthography and standardisation progress for digital tools.

Explore the data