Skip to main content
Hausa · Yoruba · Pidgin datasets

Datasets for the Nigerian languages that break generic models.

Commission clean, consented Hausa, Yoruba, and Nigerian Pidgin data for model training and evaluation: translation, speech, preference, safety, and cultural-quality tasks.

Human contact: Chat on WhatsApp · [email protected]

Pidgin
Naturalness and culture judgement
Yoruba
Fresh train/held-out translation
Hausa
Regional coverage and speech

Translation pairs

Native translations that preserve meaning and local phrasing, with train/held-out split at collection time.

  • English ↔ Hausa
  • English ↔ Yoruba
  • English ↔ Nigerian Pidgin

Speech datasets

ASR and TTS collection for Nigerian language speech, with recording protocol and consent.

  • Prompted voice reads
  • Transcript review
  • Speaker consistency options for TTS

RLHF preferences

A/B judgement tasks where Nigerian speakers choose the more natural, helpful response.

  • NGPT-compatible export
  • Majority voting
  • Gold-check quality scoring

Cultural quality ratings

Human scores for whether text sounds like a Nigerian wrote it, not just whether BLEU or chrF looks good.

  • Pidgin naturalness
  • Politeness and formality
  • Local context fit

Commercial-safe collection

Consent and provenance are built into the export path instead of added later.

  • AI-data reuse consent
  • Anonymised contributor ids
  • Gold probes excluded from corpus

Delivery formats

Ship data in the format your training pipeline expects: JSONL, CSV, TMX-like translation tables, or custom manifests.

  • Dataset card
  • Language and split fields
  • Provenance metadata

Own the Nigerian-language data loop.

Start with one language and one task type, then expand once quality and cost are proven.