Measuring the frontier
The old tests are getting too easy. Tejal Patwardhan, who leads OpenAI's frontier evals team, talks with host Andrew Mayne about why evals matter for research, how benchmarks break or get gamed, and what models should be judged on next as they keep getting more capable.
The old tests are getting too easy. Tejal Patwardhan, who leads OpenAI's frontier evals team, talks with host Andrew Mayne about why evals matter for research, how benchmarks break or get gamed, and what models should be judged on next as they keep getting more capable.