Hiring intelligence that compounds.
Domain experts and an AI-native system — so every hire sharpens the next.
A closer look at how it works.
The best person for the role isn't reading your job post.
That's not bad luck. Every standard tool selects for who's available.
- Applications reach whoever happened to be looking, in the two weeks the post was up.
- A CV is the candidate's own marketing. It tells you what they claim, not how they work.
- Cold outreach lands next to a dozen others that week, and gets the same reply: none.
So we do the slow thing. We judge people on work we've actually seen, get introduced by the ones we've already placed, and keep talking for years with no role attached. By the time you have one, the conversation is years old.
Somewhere in a thousand résumés, there's a pattern.
That's not carelessness. Every hiring process runs out of attention before it runs out of candidates.
- A thousand applications means six seconds each. That's triage, not judgment.
- Keyword filters reject people who did the work but wrote about it differently. You never see them.
- Nobody checks which signals actually predicted a good hire, so the bar never improves.
So we let software do the reading. It holds the same bar at candidate one and candidate twelve hundred, sorts on the work rather than the words, and gets sharper with every hire we make. Your team only meets the ones worth meeting.
You cannot assess work you have never done.
That's not the recruiter's fault. Every standard check measures something other than the work.
- A recruiter has read the job description, not done the job. They can check keywords, not depth.
- Interviews reward people who interview well. That's a different skill from the one you're hiring for.
- References are chosen by the candidate. They were never going to say anything else.
So before anyone reaches you, they have been through someone who has done the work. They go at the real problems, where polish stops helping, and put their name on what comes back. Software narrows the field, an expert vouches for what is left, and the hire is still your call.
The AI org chart didn’t exist five years ago.
Someone builds the model. Someone teaches it what good looks like. Someone measures whether it works. Someone puts it to work inside a real business. Someone makes sure it does no harm on the way. Three of the five are new enough that there is no playbook for hiring them and no résumé pattern to match against — so most of it is guesswork.
Pretraining
Architecture, distributed training, kernel and GPU work, and the platform that keeps a run alive for weeks. The most crowded of the five on paper, and the easiest to get wrong — plenty of people have fine-tuned a model, far fewer have owned a training run at scale.
Post-training
SFT, RLHF, preference data, and the domain experts who show a model what good looks like. Barely a job title five years ago, so there's no clean résumé signal for it. You find these people by knowing what good judgment about data looks like.
Evaluation
Benchmarks, capability measurement, regression tracking, and the eval harnesses every other decision leans on. The rarest of the five by some distance, and the only reason you can trust what eventually ships.
Application
Forward-deployed engineers, solutions architects, and the product people who put a model to work inside a real business. The integration is rarely the hard part — knowing where the model will be confidently wrong, and designing around it, is.
Safety
Red-teaming, adversarial testing, alignment research, and the policy work that decides what ships at all. Adversarial instinct doesn't show up on a résumé — the people who have it usually found it somewhere other than a safety team.
Recent essays.
Hiring should not be guesswork.
Tell us the role you’re hiring for — pretraining, post-training, evaluation, application or safety. Chances are we already know the people you need.













