Skip to main content
Data-collection programs

Nigerian-language datasets, collected properly

One of our two core engines is collecting and labelling the data that trains AI to understand Nigerians. Below is every data-collection program our real, naira-paid Nigerian panel runs — across ten Nigerian languages, in speech, text, and preference signals.

Language coverage

Five launch languages, ~420M speakers

Nigerian Pidgin
West Africa's lingua franca · 75M+ speakers
Hausa
Northern Nigeria + Niger · 70M+ speakers
Yoruba
South-West Nigeria · 45M+ speakers
Igbo
South-East Nigeria · 30M+ speakers
Nigerian English
Accented English · 200M+ speakers

Tier-2 languages collected on demand: Fulfulde · Kanuri · Tiv · Edo · Efik — almost no open data exists for these, which is exactly why we collect them.

The programs

What we collect, and how to take part

“Live” programs route to contributors today. “On brief” programs run once a buyer commits a brief — collection is scoped to the contract so contributors are only paid for real, approved work.

Live

Speech / ASR collection

Contributors read scripted prompts and record spontaneous narration in their language. Tagged by dialect, age band, state, and gender for stratified training splits. This is the voice-read annotation pipeline the panel runs today.

16-kHz mono .wav + JSONL manifest (Common Voice-compatible)

See the Pidgin sample
On brief

TTS voice modelling

A single speaker records 100+ utterances, voice identity-locked for consistency. For labs training a single-speaker text-to-speech voice in a Nigerian language. Studio-grade re-takes included.

44.1-kHz mono .wav + per-utterance metadata

Brief a TTS speaker
On brief

RLHF / preference data

Bilingual Nigerian annotators rate model outputs across Pidgin/English code-switching, cultural appropriateness, and factual grounding, with free-text rationale per ranking. For aligning English-trained foundation models to Nigerian users.

Bradley-Terry-compatible preference exports + rationale (JSONL)

Scope an RLHF pilot
On brief

Translation pairs

Native translators (not back-translation from English MT) produce sentence-aligned English ↔ Pidgin/Yoruba/Igbo/Hausa pairs, tagged with regional variant. For MT bootstrapping and benchmark sets.

Sentence-aligned pairs in TMX, CSV, or JSONL

Brief a translation set
On brief

Adversarial red-team / safety

Trained Nigerian red-teamers stress-test a model in Pidgin, Yoruba, Igbo, and Hausa — low-resource languages are a known jailbreak vector. Structured adversarial prompts, severity ratings, and reproduction steps.

Severity-rated finding log + aggregated report (under NDA)

Scope a red-team
Free sample

Inspect the data before you commit

The Nigerian Pidgin ASR pilot has a public transcript preview — the pilot targets real speakers across all six Nigerian zones, with the JSONL manifest schema already documented. No email required to read it.

Explore the Pidgin sample
Why this panel

Real Nigerians. Consent before the first clip.

Real Nigerians, real phones

Real people on real Nigerian phones and networks — no emulators, no VPNs, no synthetic personas. Native Pidgin, Yoruba, Igbo, Hausa, and NG English, narrated by the region you're shipping to.

Consent + provenance spine

Every clip we licence is traceable to an append-only ai_data_reuse consent grant, taken before the contributor records anything. Work collected without a grant stays out of the licensed set. The data ships under a standard licence with a contributor-rights schedule.

Fast naira withdrawals

Contributors are paid in naira to verified Nigerian bank accounts — withdrawals settle in 3 to 5 business days. Transparent platform spread funds QA + AI-summarised review.

Full security & compliance posture on /trust; the sub-processor list and data-residency stance on /trust/data-residency. How testers earn — including from data work — is on /testers.

Commission a dataset or join the panel

Labs: the full rate card and a structured request path live on /ai-data. Contributors: sign up, add your Nigerian bank account for payouts, and you can take paid voice and language tasks as they're published.