KLaSAn
Methodology

How We Measure Language Health

Our scoring, status bands, and projections are grounded in the LaSAn (Language Simulator and Analyser) framework and calibrated against well-resourced minority languages.

Primary Framework: LaSAn

We use the LaSAn Numerical Reference v1.0 (Wekify LLC, March 2026) as the backbone for measuring documentation and language resource health. LaSAn defines 68 resource units across eight categories. Each unit is assigned a numeric score from 0 to 100; the score drives the status label—never the other way around.

Scale types:

  • Scale A (Logarithmic): For count-based resources (entries, words, hours). Score = log(n+1) / log(ceiling+1) × 100 — reflects that going from 0→100 entries is far more meaningful than 9,900→10,000.
  • Scale B (Linear): For coverage and performance (%, accuracy). Score = (value / ceiling) × 100 — 90% coverage is proportionally better than 80%.
  • Scale C (Milestone): For discrete states (e.g. ISO 639 code: 0→25→50→75→100).

Example (LEX-01, ceiling 50,000): 0 entries → 0; 50 entries → 11; 500 → 27; 3,500 → 52 (Functional); 9,500 → 71 (Comprehensive); 50,000 → 100.

Reference Ceiling Philosophy

All ceilings are calibrated against well-resourced minority languages—those that have achieved functional infrastructure through dedicated effort, without the resources of a national majority language.

Reference languages:

  • Welsh — strong NLP, educational system, active standardisation
  • Basque — comprehensive grammar, significant corpus, functional NLP tools
  • Māori — strong oral and educational resources, growing digital presence
  • Faroese — complete literary tradition, functional digital tools
  • Luxembourgish — standardised orthography, growing corpus and NLP resources

The ceiling is the honest aspiration, not the comfortable comparison.

The Eight Resource Categories

LaSAn organises the 68 units into eight categories:

LEXLexical ResourcesDictionaries, thesauri, wordlists, etymologies, dialect vocabularies, and neologism registries.
GRMGrammatical ResourcesReference grammars, phonology, morphology, syntax, tone documentation, and interlinear text collections.
CORText CorporaGeneral, parallel, annotated, news, religious, literary, synthetic, and historical corpora.
AUDAudio and VisualElicited recordings, spontaneous speech, oral literature, transcribed audio, video, and pronunciation references.
EDUEducationalLiteracy primers, children's books, textbooks, teacher training, learning apps, flashcards, and curriculum guides.
NLPDigital and NLPKeyboards, spell checkers, tokenizers, POS taggers, morphological analysers, TTS, ASR, MT, language models.
STDStandardisationOrthography, tone marking, spelling standards, style guides, terminology, fonts, Unicode, ISO 639 codes.
ETHCultural and EthnographicProverbs, folktales, songs, riddles, kinship terminology, place names, ethnobotany, ritual language.

Language Resource Index (LRI)

The LRI is the average of all 68 resource unit scores. Each unit contributes 1/68 to the total. Range: 0–100. It drives our primary status label and projections.

Formula: LRI = (Sum of all 68 scores) / 68

RCR (Resource Coverage Rate) = (count of units with score ≥ 46) / 68 × 100%. Measures how many resource types are at Functional level or above.

CGC (Critical Gap Count) = count of units where score < 46 and the unit is a hard dependency for at least one other unit. Identifies missing resources that block downstream development.

LaSAn mapping: We currently map corpus data to six units:

  • LEX-01 General Dictionary — dictionary entries (words), ceiling 50,000
  • LEX-02 Bilingual Dictionary — bilingual pairs (translations), ceiling 30,000
  • AUD-01 Elicited Recordings — recorded items (audio clips), ceiling 20,000
  • GRM-07 Interlinear Texts — glossed sentences, ceiling 5,000
  • COR-01 General Corpus — words + sentences×6 (estimated tokens), ceiling 5,000,000
  • LEX-04 Specialised Glossaries — numbers + word problems + math expressions, ceiling 5,000

The remaining 62 units are placeholders (score 0) until additional data is available. See the Resource Units page for the full 68-unit view.

All 68 Resource Units — Full Reference

The LaSAn framework defines 68 resource units across eight categories. Each unit has a reference ceiling and thresholds for Functional (≥46) and Comprehensive (≥71) status. LRI = average of all 68 scores.

IDResourceScaleCeilingFunctional ≥Comprehensive ≥
Lexical Resources (8)
LEX-01General dictionaryLog50,000 entries1,8009,500
LEX-02Bilingual dictionaryLog30,000 pairs1,2006,200
LEX-03ThesaurusLog60 domains1328
LEX-04Specialised glossariesLog5,000 terms2801,050
LEX-05WordlistsLog3,000 total words175660
LEX-06EtymologyLog10,000 words5802,900
LEX-07Dialectal vocabularyLog5,000 items2801,050
LEX-08Neologism registryLog2,000 terms140530
Grammatical Resources (8)
GRM-01Reference grammarLinear100%46%71%
GRM-02Pedagogical grammarLinear100%46%71%
GRM-03PhonologyLinear100%46%71%
GRM-04MorphologyLinear100%46%71%
GRM-05SyntaxLinear100%46%71%
GRM-06Tone / prosodyLinear100%46%71%
GRM-07Interlinear textsLog5,000 sentences2801,050
GRM-08Comparative grammarLog20 (L×D)611
Text Corpora (8)
COR-01General corpusLog5,000,000 words18,000175,000
COR-02Parallel corpusLog500,000 pairs2,80024,000
COR-03Annotated corpusLog500,000 tokens2,80024,000
COR-04News corpusLog10,000 articles5802,900
COR-05Religious corpusLog500,000 words2,80024,000
COR-06Literary corpusLog1,000,000 words4,80044,000
COR-07Synthetic corpusLog2,000,000 sentences8,50080,000
COR-08Historical corpusLog200,000 words1,30011,000
Audio and Visual (8)
AUD-01Elicited recordingsLog20,000 items1,1006,800
AUD-02Spontaneous speechLog200 hours5 hrs28 hrs
AUD-03Oral literatureLog1,000 items140530
AUD-04Transcribed audioLog500 hours2 hrs14 hrs
AUD-05Subtitled videoLog200 hours5 hrs28 hrs
AUD-06Oral historiesLog500 interviews95410
AUD-07Pronunciation referenceLog300 environments75230
AUD-08Dialect recordingsLog2,000 (V×I)175660
Educational (8)
EDU-01Literacy primerLinear100%46%71%
EDU-02Children's booksLog500 books55200
EDU-03TextbooksLog200 units38125
EDU-04Teacher trainingLinear100%46%71%
EDU-05Learning appLog10,000 items5802,900
EDU-06Flashcard setsLog5,000 items2801,050
EDU-07Audio learningLog100 hours2.5 hrs13 hrs
EDU-08Curriculum guideLinear100%46%71%
Digital and NLP (10)
NLP-01KeyboardLinear100%46%71%
NLP-02Spell checkerCompositeScore 46Score 71
NLP-03TokenizerLinear100%92% acc.97% acc.
NLP-04POS taggerLinear100%85% acc.93% acc.
NLP-05Morphological analyserLinear100%75% cov.90% cov.
NLP-06TTS systemLinearMOS 5.0MOS 3.5MOS 4.2
NLP-07ASR systemLinear (inv.)WER 0%WER < 50%WER < 15%
NLP-08Machine translationLogBLEU 40BLEU 15BLEU 28
NLP-09Language modelLog5,000,000 tokens18,000175,000
NLP-10Benchmark datasetsLog10 tasks58
Standardisation (8)
STD-01Official orthographyLinear100%46%71%
STD-02Tone markingLinear100%46%71%
STD-03Spelling standardLog30,000 entries1,2006,200
STD-04Style guideLinear100%46%71%
STD-05Terminology standardLog5,000 terms2801,050
STD-06FontLinear100%46%71%
STD-07Unicode encodingLinear100%46%71%
STD-08ISO 639 codeMilestone4 stepsStep 3Step 4
Cultural and Ethnographic (10)
ETH-01ProverbsLog2,000175660
ETH-02FolktalesLog50095410
ETH-03SongsLog50095410
ETH-04RiddlesLog1,000140530
ETH-05Kinship terminologyLinear100%46%71%
ETH-06Place namesLog3,000230880
ETH-07EthnobotanyLog2,000175660
ETH-08Ritual languageLog50 contexts1328
ETH-09Oral historiesLog300 accounts75230
ETH-10Traditional knowledgeLog20 domains611

Source: LaSAn Numerical Reference v1.0 — Wekify LLC, March 2026. An interactive view of all 68 units with per-unit contribution to LRI is available in the Simulator.

Status Bands

All resource unit scores and the LRI map to five status bands:

0Non-existent
1–20InitiatedCritical
21–45PartialWeak / Struggling
46–70FunctionalRecovering
71–100ComprehensiveHealthy / Thriving

Language Health Figure (7-band granularity): The LRI can also be interpreted on a finer scale for visualisation:

0–10Critical
11–20Severe
21–35Weak
36–50Struggling
51–65Recovering
66–80Healthy
81–100Thriving

Kifuliiru Lab Research Formulas

Simulations in LaSAn are grounded in Kifuliiru Lab research formulas—published, peer-reviewed equations that model language vitality, corpus growth, transmission, and community dynamics. Each formula links to its full documentation on kifuliiru.org:

Additional human-dimension formulas (Conflict Attrition, Speaker Attrition, Diaspora Drift, etc.) are in development and used in the Risks tab.

View full formula catalog on kifuliiru.org

Risk & Human-Dimension Formulas

Beyond LaSAn resource scoring, we use supplementary formulas for human-dimension risks (Kifuliiru Lab research). These capture factors that affect transmission and community vitality beyond raw corpus metrics:

  • Conflict Attrition Rate — documentation output lost to instability; scales with (1 − stability index)
  • Speaker Attrition Index — speaker loss over time, influenced by migration and aging
  • Diaspora Drift — generational fluency decline away from homeland; Gen 2 fluency typically lower than homeland
  • Community Scatter Index — fragmentation of speech communities across regions
  • Transmission Break Point (TBP) — critical threshold (e.g. 0.5) above which intergenerational transmission is at risk
  • Lab Ratio — continuity / shine ratio: balance between preservation (verified + audio) and growth (sentences)

These formulas are applied in the Simulation platform under the Risks tab (Human Dimension).

Data Sources & Updates

Baseline data (words, audio, verified, sentences, translations, numbers, word problems, math expressions) is sourced from the Kifuliiru corpus and contributor dashboard. When available, live metrics are fetched via the LaSAn API.

Projections assume per-contributor monthly rates that you can adjust in the Simulator. Four effort levels—Minimal, Sustained, Accelerated, Community Sprint—map to different growth assumptions (contributors, words/month, audio/month, etc.).

Updates: Baseline and formula parameters are refreshed as corpus data evolves. Methodology changes are documented in release notes.

References

  • LaSAn Numerical Reference v1.0 — Scale, ceilings, and breakpoints for all 68 resource units
  • LaSAn Platform Specification v1.0 — System architecture and data flows
  • LaSAn Status Criteria v1.0 — Status band definitions and decision rules

Wekify LLC, March 2026. Companion to LaSAn.