Kai Mensah
QA lead who moved into AI evaluation. Breaks things for a living.
Verified skills
Each of these was earned on a timed challenge and scored against a rubric that was published before the challenge started. The work is attached — you can read it.
Builds the evaluation infrastructure a team ships against, and connects offline metrics to observed production outcomes.
Level 4 on Debug the RAG Pipeline — insisted on independent retrieval measurement and per-change attribution.
Measures retrieval quality separately from answer quality, and uses hybrid or re-ranked retrieval where the data demands it.
Correctly separated retrieval from generation failure with evidence.
Diagnoses failures where the error is misleading, tests the cheapest hypothesis first, and verifies the root cause before fixing.
Mechanism-level explanation, cheapest-hypothesis-first ordering.
Claimed, not yet verified
Self-reported — the same standing as anything on a CV. Shown here for completeness and clearly separated, because a verified section only means something if the unverified one is honest.
Work samples
Unedited submissions, with the rubric score attached.
Closest role fits
Scored against published role specs. Every number is itemised on the role page.
This profile is evidence, not a CV. Anything marked verified can be checked by reading the work.