We are building benchmarks with vertical AI companies. This role would mostly involve building evals. You will spend most of your time deciding whether an AI system actually did the job, then building the datasets, graders, and checks that proves it. We are a small YC team. You work with the founders. If you want a 9-to-5 or a remote-first job, skip this. If you want to own how we measure things, keep reading.
**What you will do**
* Turn real failures into tasks we can score
* Build datasets, gold labels, graders, and human review loops
* Find ways to prevent reward hacking
* Find LLM incapabilities in production systems
**You will like this if**
* You have shipped agents.
* You have built evals, graders, or datasets for LLMs or agents
* You can sit with a lot of traces and find the few that matter
Visa sponsorship is offered for this role.