As a Member of Technical Staff on Evals, you'll own how we measure AI Employee performance on the job.
**Areas you may work in:**
* Evaluation systems: datasets, offline replay, scorers, and regression alerts
* Feedback loops for models and agents
* Autonomous failure analysis
* Defining what good, degraded, and failed runs look like, and alerting on them
**You may be a good fit if you:**
* Have built evaluation or measurement systems, such as AI evals, experimentation, ranking, or search quality
* Can turn ambiguous quality questions into concrete metrics and decisions
* Have strong software engineering fundamentals and ship production systems
* Are comfortable where there are few established patterns
* Want to work in person in San Francisco
**Even better:**
* You've shipped LLM systems to production and learned where they break
* You have strong data instincts and work well with researchers
* You contribute to open source projects
Visa sponsorship is not available for this role.