Skip to content
Provenly
Proof profile

Kai Mensah

QA lead who moved into AI evaluation. Breaks things for a living.

LondonCV title: QA EngineerJoined 7/31/2026
3
Verified proofs
2
Unverified claims
1
Work samples
33
Best role fit

Verified skills

Challenge-verified ×1.0 Self-claimed ×0.3 No evidence ×0

Each of these was earned on a timed challenge and scored against a rubric that was published before the challenge started. The work is attached — you can read it.

AI Evaluation
Challenge-verified7/31/2026
Level 4Leading

Builds the evaluation infrastructure a team ships against, and connects offline metrics to observed production outcomes.

Read the work →

Level 4 on Debug the RAG Pipeline — insisted on independent retrieval measurement and per-change attribution.

Retrieval & RAG Systems
Challenge-verified7/31/2026
Level 3Independent

Measures retrieval quality separately from answer quality, and uses hybrid or re-ranked retrieval where the data demands it.

Read the work →

Correctly separated retrieval from generation failure with evidence.

Debugging & Diagnosis
Challenge-verified7/31/2026
Level 3Independent

Diagnoses failures where the error is misleading, tests the cheapest hypothesis first, and verifies the root cause before fixing.

Read the work →

Mechanism-level explanation, cheapest-hypothesis-first ordering.

Claimed, not yet verified

Self-reported — the same standing as anything on a CV. Shown here for completeness and clearly separated, because a verified section only means something if the unverified one is honest.

Six years in QA
Data AnalysisL2 claimed
I built the regression suite

Work samples

Unedited submissions, with the rubric score attached.

Closest role fits

Scored against published role specs. Every number is itemised on the role page.

This profile is evidence, not a CV. Anything marked verified can be checked by reading the work.

Build your own