Lena Fischer
AI Engineer — production LLM systems at scale
Verified skills
Each of these was earned on a timed challenge and scored against a rubric that was published before the challenge started. The work is attached — you can read it.
Measures retrieval quality separately from answer quality, and uses hybrid or re-ranked retrieval where the data demands it.
Verified on Debug the RAG Pipeline.
Designs evals that target real failure modes, chooses metrics that resist gaming, and knows the limits of what the eval measures.
Proposed separate retrieval measurement with a ship threshold.
Diagnoses failures where the error is misleading, tests the cheapest hypothesis first, and verifies the root cause before fixing.
Correct causal diagnosis, though thinly evidenced.
Claimed, not yet verified
Self-reported — the same standing as anything on a CV. Shown here for completeness and clearly separated, because a verified section only means something if the unverified one is honest.
Work samples
Unedited submissions, with the rubric score attached.
Closest role fits
Scored against published role specs. Every number is itemised on the role page.
This profile is evidence, not a CV. Anything marked verified can be checked by reading the work.