We companies build and launch public benchmarks that measure what AI can do in real-world domains. Our work includes [ITSMBench](https://newmeasure.ai/benchmarks), which evaluates whether AI models can autonomously resolve IT tickets. We’re looking for a intern to help build domain-specific benchmarks, develop tools that make benchmarking easier, and research ways to create evaluations that are fair, challenging, and useful.
**What you’ll work on**
* Help experts build realistic tasks and environments for domain-specific AI benchmarks.
* Develop tooling for running evaluations and investigating model failures.
* Research and test ideas for constructing fair and difficult benchmarks.
**What we’re looking for**
* Strong programming fundamentals and experience building software through coursework, personal projects, open-source contributions, or prior work.
* Interest in evals research, and understanding where models succeed or fail.
Experience with evaluation systems or research projects is a plus, but isn’t required.
There is a potential opportunity to join full time after three months.
Visa sponsorship is offered for this role.