Does your AI actually work for your people?
A model that performs well in evaluation can fail the moment it meets a Hausa speaker on a noisy line, a Kenyan farmer asking about a local crop, or a customer using the banking vocabulary people actually use in Lagos.
We find out before your customers do — with native speakers of the specific varieties and professionals from the domain concerned, and we tell you exactly what breaks and where.
A good average can hide a total failure.
Suppose an evaluation set has four thousand items, forty of them in a smaller national language. The system gets thirty of those forty wrong.
The overall accuracy falls by less than one percentage point. The system passes its acceptance test and goes live.
For a speaker of that language, it fails three times in four.
The average is not miscalculated. It simply answers a question that nobody in that position is asking. This is why every figure we report is broken down by language and region.
Six things, judged separately.
Asking whether an answer is good produces scores nobody can explain. Judged separately, a disagreement can be located — and so can a fix.
Language
Does it produce language a native speaker would actually use?
Grammar, naturalness, register and vocabulary — assessed by native speakers of the specific variety, not by speakers of a related one.
Culture and context
Does the answer make sense where it will be read?
Local institutions, currencies, school systems, transport, healthcare pathways, banking and government processes. An answer can be fluent, true in general, and useless locally.
Domain accuracy
Is what it says correct, judged by someone who would know?
Assessed by professionals working in the field concerned. Without that, an evaluator can only judge whether an answer reads as though it were true — and modern systems are very good at that.
Speech and accent
Does it understand how people actually speak?
Regional accents, code-switching, local names and places, background noise, telephone-quality audio, and speakers across ages and genders.
Safety
Could acting on this answer cause harm here?
Judged against the user’s situation rather than the evaluator’s. Advice that is safe in one setting can be dangerous in another.
Preference
Which answer would your users actually choose, and why?
Paired comparison with the reasoning recorded, controlled for the known biases toward longer and more heavily formatted answers.
Expert human judgement, not general annotation.
Anyone can be asked whether an answer looks right. Whether it is right, in a clinical, legal, financial or agricultural context, is a question only somebody who works in that field can answer.
Our network includes PhD holders, academic researchers, doctors and clinicians, engineers, accountants, legal professionals, scientists, educators, agronomists and technology professionals — alongside native speakers across the languages we cover.
Why it matters, concretely
A system answers a farming question in Hausa. It names the right active ingredient, and reads confidently and well.
A general evaluator scores it highly. It is fluent, direct and addresses the question.
An agronomist scores it low. The application rate given is one used in temperate conditions, and would scorch the crop in northern Nigerian heat.
Both judged honestly. Only one had the knowledge to see the fault — and only one of those scores would have protected the farmer.
Six stages, in order.
Scope
We agree the languages, regions, domains and use cases, and set targets for each so no group can end up missing by accident.
Assemble
Native speakers of the specific varieties, and professionals from the domain concerned, drawn from our trained contributor network.
Design
A rubric with each dimension defined by a rule rather than an example, piloted and measured for agreement before production begins.
Evaluate
Real interactions, not curated test cases. Control items included so a preference for length or formatting cannot be mistaken for a preference for quality.
Adjudicate
Substantive disagreements resolved by a reviewer with the reasoning recorded, because a disagreement often reveals more than a verdict.
Report
Findings by language, region and failure type, with the method and its limits stated plainly enough to survive a regulator reading it.
A report you can act on.
Written so that a product team can prioritise from it, and so that it holds up if a regulator, a journalist or a court reads it later.
- Performance by language and region, not a single blended figure — an aggregate score hides complete failure on a smaller group.
- Failures sorted into named categories, with counts, so you know what to fix rather than only that something is wrong.
- Worked examples of each failure type, with the reasoning of the person who judged it.
- What the evaluation supports a claim about, and what it does not.
- The proportion of judgements that were close calls, so you can see where the finding is firm and where it is not.
- Recommendations sized to the fault, including where we judge no action is needed.
Anyone about to put an AI system in front of people here.
Model developers
How does your model perform across the languages and contexts of these markets, and where does it fail?
Banks and fintech
Does your assistant handle the vocabulary customers actually use, and is its financial reasoning sound?
Health and government
Is the advice safe and locally correct, and does it understand the pathways people actually take?
Telecoms and retail
Does it cope with code-switching, accents and noisy lines, and does it escalate when it should?
Start
Find out before your customers do.
Tell us the system, the markets and the languages. We will come back with a scope, a timeline and what the report will and will not be able to tell you.