Speech / ASR collection
Contributors read scripted prompts and record spontaneous narration in their language. Tagged by dialect, age band, state, and gender for stratified training splits. This is the voice-read annotation pipeline the panel runs today.
16-kHz mono .wav + JSONL manifest (Common Voice-compatible)
See the Pidgin sample