> ## Content Index
> Fetch the complete content index at: https://globalfeed.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Stanford publishes a benchmark measuring AI agents on scientific workflows
- URL: https://globalfeed.ai/en/stanford-publishes-a-benchmark-measuring-ai-agents-on-scientific-workflows/
- Published: 2026-08-28T03:37:57.000Z
- Updated: 2026-08-30T15:27:38.000Z
- Description: The first release of Terminal-Bench-Science contains 70 tasks drawn from work researchers actually performed. In the first results, the best agent completes 30 percent of them.
- Author: GlobalFeed Editor
- Tags: Stanford, benchmarks, agents, science, x-StanfordAILab, dil-en, elle

A Stanford-led effort has published **Terminal-Bench-Science**, a benchmark that measures AI agents on scientific research workflows rather than software development.

The first release contains **70 tasks**, each drawn from work practising researchers actually performed. In the first results, the best agent solves **30 percent** of them.

For the task list, the method and the full results table: [the Terminal-Bench announcement](https://www.tbench.ai/news/tb-science-announcement?ref=globalfeed.ai) and [the GitHub repository](https://github.com/harbor-framework/terminal-bench-science?ref=globalfeed.ai).