Skip to content
Chorus Language LabTuvaluGPT Research

TuvaluGPT Research overview

Language technology
for Tuvaluan.

TuvaluGPT is a Tuvaluan–English language model project from Chorus Language Lab. Our work brings together parallel text, translation training and instruction tuning to study what specialized models can offer a language with limited digital resources.

TuvaluGPT and TuvaLLM were developed together in collaboration. TuvaLLM speech recognition is integrated into TuvaluGPT's voice features. Collaboration and integration.

Languages
Tuvaluan · English
Model family
Qwen3-30B-A3B
Parameters
30B total · 3B active
Reported evaluation
March 2026

In the reported shared comparison, the Stage B model scored highest overall and led six of seven task slices. This is a small, task-specific result: the comparison contains 28 examples and does not establish general model superiority.

Shared Tuvaluan benchmark · March 2026 · chrF++ ↑
ModelScore
TuvaluGPT · Stage B41.8
GPT-5.436.1
Claude Sonnet 4.634.2
Qwen3-30B · base13.7

chrF++ measures overlap with reference text; it does not directly measure factual accuracy or conversational usefulness. Model versions and the evaluation sample are specific to the reported runs.

Inspect task results and methodology

Stage A

Translation training

A translation adapter is trained on aligned Tuvaluan–English text. This model then translates English instruction data into Tuvaluan to expand the available training material.

Stage B

Instruction tuning

A fresh adapter on the Qwen3 chat model learns from English, synthetic Tuvaluan, cross-lingual and translation examples. Stage A contributes data; its adapter weights are not carried forward.

The corpus has substantial religious-domain coverage, and synthetic instructions can reproduce translation errors. Performance varies by direction and task. On a separate 12-paragraph literary evaluation, Stage B scored 47.1 chrF++ for English to Tuvaluan and 42.4 for Tuvaluan to English; GPT-5.4 scored 45.5 and 51.5 respectively.

Broader evaluation requires independent text, native-speaker review and checks for factual fidelity. The technical report documents the training settings, failure examples and gaps in reproducibility.

Read the full analysis