Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or data vendors, set the bar for scientific capability in AI.
Read full article →