How We Measure Language Health
Our scoring, status bands, and projections are grounded in the LaSAn (Language Simulator and Analyser) framework and calibrated against well-resourced minority languages.
Primary Framework: LaSAn
We use the LaSAn Numerical Reference v1.0 (Wekify LLC, March 2026) as the backbone for measuring documentation and language resource health. LaSAn defines 68 resource units across eight categories. Each unit is assigned a numeric score from 0 to 100; the score drives the status label—never the other way around.
Scale types:
- Scale A (Logarithmic): For count-based resources (entries, words, hours).
Score = log(n+1) / log(ceiling+1) × 100— reflects that going from 0→100 entries is far more meaningful than 9,900→10,000. - Scale B (Linear): For coverage and performance (%, accuracy).
Score = (value / ceiling) × 100— 90% coverage is proportionally better than 80%. - Scale C (Milestone): For discrete states (e.g. ISO 639 code: 0→25→50→75→100).
Example (LEX-01, ceiling 50,000): 0 entries → 0; 50 entries → 11; 500 → 27; 3,500 → 52 (Functional); 9,500 → 71 (Comprehensive); 50,000 → 100.
Reference Ceiling Philosophy
All ceilings are calibrated against well-resourced minority languages—those that have achieved functional infrastructure through dedicated effort, without the resources of a national majority language.
Reference languages:
- Welsh — strong NLP, educational system, active standardisation
- Basque — comprehensive grammar, significant corpus, functional NLP tools
- Māori — strong oral and educational resources, growing digital presence
- Faroese — complete literary tradition, functional digital tools
- Luxembourgish — standardised orthography, growing corpus and NLP resources
The ceiling is the honest aspiration, not the comfortable comparison.
The Eight Resource Categories
LaSAn organises the 68 units into eight categories:
Language Resource Index (LRI)
The LRI is the average of all 68 resource unit scores. Each unit contributes 1/68 to the total. Range: 0–100. It drives our primary status label and projections.
Formula: LRI = (Sum of all 68 scores) / 68
RCR (Resource Coverage Rate) = (count of units with score ≥ 46) / 68 × 100%. Measures how many resource types are at Functional level or above.
CGC (Critical Gap Count) = count of units where score < 46 and the unit is a hard dependency for at least one other unit. Identifies missing resources that block downstream development.
LaSAn mapping: We currently map corpus data to six units:
- LEX-01 General Dictionary — dictionary entries (words), ceiling 50,000
- LEX-02 Bilingual Dictionary — bilingual pairs (translations), ceiling 30,000
- AUD-01 Elicited Recordings — recorded items (audio clips), ceiling 20,000
- GRM-07 Interlinear Texts — glossed sentences, ceiling 5,000
- COR-01 General Corpus — words + sentences×6 (estimated tokens), ceiling 5,000,000
- LEX-04 Specialised Glossaries — numbers + word problems + math expressions, ceiling 5,000
The remaining 62 units are placeholders (score 0) until additional data is available. See the Resource Units page for the full 68-unit view.
All 68 Resource Units — Full Reference
The LaSAn framework defines 68 resource units across eight categories. Each unit has a reference ceiling and thresholds for Functional (≥46) and Comprehensive (≥71) status. LRI = average of all 68 scores.
| ID | Resource | Scale | Ceiling | Functional ≥ | Comprehensive ≥ |
|---|---|---|---|---|---|
| Lexical Resources (8) | |||||
| LEX-01 | General dictionary | Log | 50,000 entries | 1,800 | 9,500 |
| LEX-02 | Bilingual dictionary | Log | 30,000 pairs | 1,200 | 6,200 |
| LEX-03 | Thesaurus | Log | 60 domains | 13 | 28 |
| LEX-04 | Specialised glossaries | Log | 5,000 terms | 280 | 1,050 |
| LEX-05 | Wordlists | Log | 3,000 total words | 175 | 660 |
| LEX-06 | Etymology | Log | 10,000 words | 580 | 2,900 |
| LEX-07 | Dialectal vocabulary | Log | 5,000 items | 280 | 1,050 |
| LEX-08 | Neologism registry | Log | 2,000 terms | 140 | 530 |
| Grammatical Resources (8) | |||||
| GRM-01 | Reference grammar | Linear | 100% | 46% | 71% |
| GRM-02 | Pedagogical grammar | Linear | 100% | 46% | 71% |
| GRM-03 | Phonology | Linear | 100% | 46% | 71% |
| GRM-04 | Morphology | Linear | 100% | 46% | 71% |
| GRM-05 | Syntax | Linear | 100% | 46% | 71% |
| GRM-06 | Tone / prosody | Linear | 100% | 46% | 71% |
| GRM-07 | Interlinear texts | Log | 5,000 sentences | 280 | 1,050 |
| GRM-08 | Comparative grammar | Log | 20 (L×D) | 6 | 11 |
| Text Corpora (8) | |||||
| COR-01 | General corpus | Log | 5,000,000 words | 18,000 | 175,000 |
| COR-02 | Parallel corpus | Log | 500,000 pairs | 2,800 | 24,000 |
| COR-03 | Annotated corpus | Log | 500,000 tokens | 2,800 | 24,000 |
| COR-04 | News corpus | Log | 10,000 articles | 580 | 2,900 |
| COR-05 | Religious corpus | Log | 500,000 words | 2,800 | 24,000 |
| COR-06 | Literary corpus | Log | 1,000,000 words | 4,800 | 44,000 |
| COR-07 | Synthetic corpus | Log | 2,000,000 sentences | 8,500 | 80,000 |
| COR-08 | Historical corpus | Log | 200,000 words | 1,300 | 11,000 |
| Audio and Visual (8) | |||||
| AUD-01 | Elicited recordings | Log | 20,000 items | 1,100 | 6,800 |
| AUD-02 | Spontaneous speech | Log | 200 hours | 5 hrs | 28 hrs |
| AUD-03 | Oral literature | Log | 1,000 items | 140 | 530 |
| AUD-04 | Transcribed audio | Log | 500 hours | 2 hrs | 14 hrs |
| AUD-05 | Subtitled video | Log | 200 hours | 5 hrs | 28 hrs |
| AUD-06 | Oral histories | Log | 500 interviews | 95 | 410 |
| AUD-07 | Pronunciation reference | Log | 300 environments | 75 | 230 |
| AUD-08 | Dialect recordings | Log | 2,000 (V×I) | 175 | 660 |
| Educational (8) | |||||
| EDU-01 | Literacy primer | Linear | 100% | 46% | 71% |
| EDU-02 | Children's books | Log | 500 books | 55 | 200 |
| EDU-03 | Textbooks | Log | 200 units | 38 | 125 |
| EDU-04 | Teacher training | Linear | 100% | 46% | 71% |
| EDU-05 | Learning app | Log | 10,000 items | 580 | 2,900 |
| EDU-06 | Flashcard sets | Log | 5,000 items | 280 | 1,050 |
| EDU-07 | Audio learning | Log | 100 hours | 2.5 hrs | 13 hrs |
| EDU-08 | Curriculum guide | Linear | 100% | 46% | 71% |
| Digital and NLP (10) | |||||
| NLP-01 | Keyboard | Linear | 100% | 46% | 71% |
| NLP-02 | Spell checker | Composite | — | Score 46 | Score 71 |
| NLP-03 | Tokenizer | Linear | 100% | 92% acc. | 97% acc. |
| NLP-04 | POS tagger | Linear | 100% | 85% acc. | 93% acc. |
| NLP-05 | Morphological analyser | Linear | 100% | 75% cov. | 90% cov. |
| NLP-06 | TTS system | Linear | MOS 5.0 | MOS 3.5 | MOS 4.2 |
| NLP-07 | ASR system | Linear (inv.) | WER 0% | WER < 50% | WER < 15% |
| NLP-08 | Machine translation | Log | BLEU 40 | BLEU 15 | BLEU 28 |
| NLP-09 | Language model | Log | 5,000,000 tokens | 18,000 | 175,000 |
| NLP-10 | Benchmark datasets | Log | 10 tasks | 5 | 8 |
| Standardisation (8) | |||||
| STD-01 | Official orthography | Linear | 100% | 46% | 71% |
| STD-02 | Tone marking | Linear | 100% | 46% | 71% |
| STD-03 | Spelling standard | Log | 30,000 entries | 1,200 | 6,200 |
| STD-04 | Style guide | Linear | 100% | 46% | 71% |
| STD-05 | Terminology standard | Log | 5,000 terms | 280 | 1,050 |
| STD-06 | Font | Linear | 100% | 46% | 71% |
| STD-07 | Unicode encoding | Linear | 100% | 46% | 71% |
| STD-08 | ISO 639 code | Milestone | 4 steps | Step 3 | Step 4 |
| Cultural and Ethnographic (10) | |||||
| ETH-01 | Proverbs | Log | 2,000 | 175 | 660 |
| ETH-02 | Folktales | Log | 500 | 95 | 410 |
| ETH-03 | Songs | Log | 500 | 95 | 410 |
| ETH-04 | Riddles | Log | 1,000 | 140 | 530 |
| ETH-05 | Kinship terminology | Linear | 100% | 46% | 71% |
| ETH-06 | Place names | Log | 3,000 | 230 | 880 |
| ETH-07 | Ethnobotany | Log | 2,000 | 175 | 660 |
| ETH-08 | Ritual language | Log | 50 contexts | 13 | 28 |
| ETH-09 | Oral histories | Log | 300 accounts | 75 | 230 |
| ETH-10 | Traditional knowledge | Log | 20 domains | 6 | 11 |
Source: LaSAn Numerical Reference v1.0 — Wekify LLC, March 2026. An interactive view of all 68 units with per-unit contribution to LRI is available in the Simulator.
Status Bands
All resource unit scores and the LRI map to five status bands:
Language Health Figure (7-band granularity): The LRI can also be interpreted on a finer scale for visualisation:
Kifuliiru Lab Research Formulas
Simulations in LaSAn are grounded in Kifuliiru Lab research formulas—published, peer-reviewed equations that model language vitality, corpus growth, transmission, and community dynamics. Each formula links to its full documentation on kifuliiru.org:
Additional human-dimension formulas (Conflict Attrition, Speaker Attrition, Diaspora Drift, etc.) are in development and used in the Risks tab.
View full formula catalog on kifuliiru.orgRisk & Human-Dimension Formulas
Beyond LaSAn resource scoring, we use supplementary formulas for human-dimension risks (Kifuliiru Lab research). These capture factors that affect transmission and community vitality beyond raw corpus metrics:
- Conflict Attrition Rate — documentation output lost to instability; scales with (1 − stability index)
- Speaker Attrition Index — speaker loss over time, influenced by migration and aging
- Diaspora Drift — generational fluency decline away from homeland; Gen 2 fluency typically lower than homeland
- Community Scatter Index — fragmentation of speech communities across regions
- Transmission Break Point (TBP) — critical threshold (e.g. 0.5) above which intergenerational transmission is at risk
- Lab Ratio — continuity / shine ratio: balance between preservation (verified + audio) and growth (sentences)
These formulas are applied in the Simulation platform under the Risks tab (Human Dimension).
Data Sources & Updates
Baseline data (words, audio, verified, sentences, translations, numbers, word problems, math expressions) is sourced from the Kifuliiru corpus and contributor dashboard. When available, live metrics are fetched via the LaSAn API.
Projections assume per-contributor monthly rates that you can adjust in the Simulator. Four effort levels—Minimal, Sustained, Accelerated, Community Sprint—map to different growth assumptions (contributors, words/month, audio/month, etc.).
Updates: Baseline and formula parameters are refreshed as corpus data evolves. Methodology changes are documented in release notes.
References
- LaSAn Numerical Reference v1.0 — Scale, ceilings, and breakpoints for all 68 resource units
- LaSAn Platform Specification v1.0 — System architecture and data flows
- LaSAn Status Criteria v1.0 — Status band definitions and decision rules
Wekify LLC, March 2026. Companion to LaSAn.