Stanford publishes a benchmark measuring AI agents on scientific workflows
The first release of Terminal-Bench-Science contains 70 tasks drawn from work researchers actually performed. In the first results, the best agent completes 30 percent of them.
A Stanford-led effort has published Terminal-Bench-Science, a benchmark that measures AI agents on scientific research workflows rather than software development.
The first release contains 70 tasks, each drawn from work practising researchers actually performed. In the first results, the best agent solves 30 percent of them.
For the task list, the method and the full results table: the Terminal-Bench announcement and the GitHub repository.