Why Mirror
A puzzle score doesn't show who can do the job.
Candidates memorize solutions, pass the test, then stall on a real codebase. Mirror screens on work like yours, so the people who pass are ready for day one.
Illustrative example: the same four candidates, screened two ways.
From the screen to the first month
Where a puzzle score stops predicting the job.
| Puzzle test | Mirror | |
|---|---|---|
| The screen | Recall of practised problems | Reasoning about a system like yours |
| Who passes | Whoever drilled the most puzzles | Whoever can do the role |
| Day one | Stalls on a real codebase | Has already worked problems shaped like yours |
| First months | Seniors spend weeks hand-holding | Less hand-holding, shipping sooner |
| Long term | Mis-hires and repeat searches | Fewer mis-hires, less time spent training |
Where the cost shows up
A bad screen is cheap on the day and expensive for months.
Senior time
Every week spent hand-holding a new hire is a week your seniors don't ship.
Mis-hires
A wrong hire costs the search, the onboarding, then the search again.
Ramp-up
Hires screened on your kind of work start closer to it.
Interview hours
Weak fits drop out at the screen, before your engineers meet them.
Screen for the job, not the puzzle.
Try it on a sample repo or your own.