Stop screening for proxies you already know are wrong.
Define the role as skills at required proficiency levels instead of years and keywords. Send one challenge link. Get a ranked shortlist where every score is itemised and the actual submitted work is one click away.
1 of your best AI Builder candidates would fail a keyword screen.
These people score 83+ on the AI Builder spec. Their job titles do not contain the words a filter looks for.
How it works
- 01Write a skill spec, not a job ad
Pick four to six skills and set a required proficiency level on each. There is no years field, no degree field, and no 'must have worked at' field — because those are the proxies that were failing you.
- 02Send one link
We match or generate a challenge for the spec. Candidates get a 60–90 minute task with a published rubric. You do nothing while they work.
- 03Read work, not documents
A ranked shortlist, blind by default — no name, photo, school, employer, or years. Every score itemised by skill, with the submission one click away.
- 04Interview about what you read
Point at the trade-off they made and ask why. Much more informative than 'tell me about a time', and much harder to rehearse.
The one substitution that changes everything
“5+ years experience with LLMs. BSc in Computer Science or equivalent. Experience at a high-growth startup preferred.”
Excludes the person with three shipped agents and no title. Includes the person who wrote “AI” in status reports for five years.
Seniority expressed as a proficiency level with a behavioural definition, so “Level 3” means something specific you can argue about productively.
Start from a role archetype
Ships end-to-end AI features: designs the system, writes the prompts, builds the evals, and owns whether it works in production.
Owns the production AI system: architecture, reliability, cost, latency, and what happens when the model is wrong.
Owns what the AI product does, what 'good enough' means, and the trade-off between capability, cost, and risk.
Owns the problem being solved, the sequence of work, and the trade-offs nobody else wants to make.
Builds the first version, decides the architecture, and figures out the requirements while shipping.
Finds the constraint, sizes the opportunity, designs the test, and reads the result honestly.
Questions people ask
- How is this different from a coding test?
- Coding tests measure execution on a puzzle. These challenges present a realistic, messy situation with genuine trade-offs and score judgment — what the candidate refuses to automate, what they cut, and how they would verify they were right.
- Won't candidates just use AI to answer?
- They will, and they should — using AI well is part of the job. Rubrics target the judgment an assistant does not supply. In practice, unedited assistant output is easy to spot: comprehensive, balanced, committing to nothing. That pattern scores in the middle, which is correct.
- How long does it take to set up a role?
- Under three minutes. Pick a role archetype, adjust the required proficiency level per skill, and send candidates one link. There is no years-of-experience field to fill in because we do not have one.
- Can I still interview candidates?
- Yes, and the interview gets better. Instead of asking someone to describe how they prioritise, you point at the cut they made in the challenge and ask why. That is much harder to rehearse.
- How do I defend a hiring decision made this way?
- Every score is itemised by skill, with the submitted work one click away. That is considerably more defensible than 'their CV looked like a fit', and it holds up better in a review.
- Does blind screening actually reduce bias?
- It removes some of the strongest proxies for class and background — name, school, employer, career gaps. It is not automatic: an assessment can encode bias too, which is why rubrics are public, exclude language fluency, and are sample-audited.