What we have shipped, and what we are building.
Structured accounts of real engagements — the problem, the approach, what the evaluation found and what changed as a result.
We are an early-stage company. There are two accounts here rather than twenty, and each states honestly what stage it is at. No case study describes work we have not done.
Property document verification for a market where forgery is routine.
Problem
Property and land title fraud is a major source of financial loss in Nigeria. Forged documents pass manual review at real estate platforms, law firms and among individual buyers, and generic document-verification tooling has never seen the formats involved.
Approach
We built a document intelligence pipeline around Nigerian land titles, NIN slips and CAC certificates specifically, using a labelled dataset of genuine and forged examples rather than adapting a model trained on other jurisdictions.
Evaluation
Tested against real and synthetic documents across the formats in active use, with review by people familiar with what a legitimate document looks like in practice.
Finding
Forgery signals in these documents are format-specific. Detection depends on knowing the conventions of each issuing body, which a general-purpose model does not encode.
Outcome
PlotYGuard runs in production today at plotyguard.com, verifying documents for buyers, lawyers and platforms.
Building a labelled Igbo and Nigerian Pidgin speech dataset.
Problem
Speech models trained largely on Western languages handle natural Nigerian speech poorly, particularly the constant code-switching between Igbo, Pidgin and English. The data needed to address that has not been collected at usable quality.
Approach
Through Lean Lab we collect and label native Igbo and Nigerian Pidgin speech across dialects and regions, using a structured annotation schema covering transcript accuracy, code-switching level, dialect and tone.
Evaluation
Quality is measured continuously through gold-standard items and reviewer agreement rather than sampled at the end of collection.
Finding
Dialect and code-switching are not edge cases in this data — they are the majority condition, which changes how a schema and a review process have to be designed.
Outcome
Collection is ongoing through Lean Lab and forms the foundation for Lean Language. Volumes and model results are not yet published; when they are, they will appear with their methodology.
Early partnerships and pilots