Stanford
Stanford publishes a benchmark measuring AI agents on scientific workflows
The first release of Terminal-Bench-Science contains 70 tasks drawn from work researchers actually performed. In the first results, the best agent completes 30 percent of them.