Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
Terminal-Bench-Science introduces a benchmark of 70 verifiable scientific workflows across five disciplines; the leading agent resolves only 30% of tasks. HN discusses whether its tests capture correctness and instruction-following, and the risks and promise of AI-assisted research.