TuvaluGPT Research overview
Language technology
for Tuvaluan.
TuvaluGPT is a Tuvaluan–English language model project from Chorus Language Lab. Our work brings together parallel text, translation training and instruction tuning to study what specialized models can offer a language with limited digital resources.
TuvaluGPT and TuvaLLM were developed together in collaboration. TuvaLLM speech recognition is integrated into TuvaluGPT's voice features. Collaboration and integration.
- Languages
- Tuvaluan · English
- Model family
- Qwen3-30B-A3B
- Parameters
- 30B total · 3B active
- Reported evaluation
- March 2026
01 / Evaluation
Results in context.
In the reported shared comparison, the Stage B model scored highest overall and led six of seven task slices. This is a small, task-specific result: the comparison contains 28 examples and does not establish general model superiority.
| Model | Score |
|---|---|
| TuvaluGPT · Stage B | 41.8 |
| GPT-5.4 | 36.1 |
| Claude Sonnet 4.6 | 34.2 |
| Qwen3-30B · base | 13.7 |
chrF++ measures overlap with reference text; it does not directly measure factual accuracy or conversational usefulness. Model versions and the evaluation sample are specific to the reported runs.
Inspect task results and methodology02 / Method
Translation to instruction following.
Stage A
Translation training
A translation adapter is trained on aligned Tuvaluan–English text. This model then translates English instruction data into Tuvaluan to expand the available training material.
Stage B
Instruction tuning
A fresh adapter on the Qwen3 chat model learns from English, synthetic Tuvaluan, cross-lingual and translation examples. Stage A contributes data; its adapter weights are not carried forward.
03 / Limitations
What remains unresolved.
The corpus has substantial religious-domain coverage, and synthetic instructions can reproduce translation errors. Performance varies by direction and task. On a separate 12-paragraph literary evaluation, Stage B scored 47.1 chrF++ for English to Tuvaluan and 42.4 for Tuvaluan to English; GPT-5.4 scored 45.5 and 51.5 respectively.
Broader evaluation requires independent text, native-speaker review and checks for factual fidelity. The technical report documents the training settings, failure examples and gaps in reproducibility.
Read the full analysis04 / Resources