New AI Agent Evaluation Benchmarks Reveal Where Agents Still Fail
TECH

New AI Agent Evaluation Benchmarks Reveal Where Agents Still Fail

32+
Signals

Strategic Overview

  • 01.
    A wave of benchmarks released within days of each other in late September and early October 2026 moves AI agent evaluation beyond generic coding tasks into specialized domains: computational materials science, long-horizon scientific discovery, enterprise knowledge retrieval, cybersecurity command-line tool use, and binary reverse engineering.
  • 02.
    Across all of these benchmarks, agents handle narrow tasks under full guidance reasonably well, but performance drops sharply as tasks get longer, guidance is withdrawn, or the reasoning required becomes more domain-specific - CompMat-Bench pass rates fall as workflows lengthen, KaliBench caps open-weight exact-command accuracy at 42%, and SRE-Bench full-solve rates stay in the 27-32% range for frontier models.
  • 03.
    A separate paper argues that because an agent's behavior depends on configuration choices such as task information, self-verification tools, and time budget - not just the backbone model - roughly 54% of the variance in outcomes on identical repeated runs comes from randomness alone, not from what is being tested.

Deep Analysis

A Benchmark for Every Domain

In late September and early October 2026, four new agent benchmarks and one evaluation-methodology paper arrived within days of each other, joining SRE-Bench - an earlier 2026 reverse-engineering benchmark still being actively tested and discussed - to form a cluster of narrow, domain-specific agent evaluations. CompMat-Bench tests whether agents can execute real steps from published computational materials science studies without needing to rerun expensive simulations [1]. EurekaBench pushes agents through 26 long-horizon scientific-discovery tasks spanning neuroscience, geophysics, astrophysics, computer science, plasma physics, and chemistry, tied to 306 target insights [2]. Company Knowledge Bench, built by Kapa.ai from 1,000 real enterprise retrieval cases, measures how well agents find exactly the right internal documentation without drowning the context window in noise [3]. KaliBench grades 8,504 natural-language-to-command translations across 1,642 Kali Linux tools and 23 capability dimensions [4], with code and data published openly [5]. And SRE-Bench forces agents to reverse-engineer compiled binaries from 19 private programs built from scratch specifically to avoid training-data contamination [6]. The common thread is a shift away from generic coding and chat leaderboards toward narrow, expert-curated domains where correctness is harder to fake.

Guided Tasks Succeed, Autonomy Fails - Everywhere

Every benchmark in this cluster reproduces the same shape of result. On CompMat-Bench, agents score a respectable 66.0-90.4% pass rate on single tasks under full methodological guidance, but performance drops as workflows lengthen and guidance is withdrawn, with most failures traced to scientific reasoning errors rather than software mistakes [1]. On EurekaBench, agents already exceed human scientists at predictive accuracy but fall substantially short at deriving the scientific insight the task is built around [2]. On Company Knowledge Bench, a naive 'agentic grep' baseline pulls roughly four times the tokens of a fixed retrieval pipeline while only needing about one in every twenty-five chunks it returns, showing that being agentic alone does not guarantee efficiency; the benchmark's best-performing system was instead a separately engineered agentic retriever, which achieved the top accuracy score [3]. On KaliBench, no open-weight model clears 42% exact-command accuracy once tool hints are removed [4]. And on SRE-Bench, even frontier models that excel at source-level vulnerability research - the skill reverse engineering is supposed to build on - post full-solve rates of only 26.9-31.5% [6]. Across five very different domains, the pattern holds: narrow, well-scaffolded steps are solvable; open-ended, reduced-guidance reasoning is not.

Saturated on Paper, Shaky in Practice

SRE-Bench's creators built it specifically to be contamination-free, using over 5,000 expert-hours of clean-room program development so no binary could already be baked into a model's training data [6]. Despite that design, a related leaderboard reported a frontier model reaching a near-perfect pass@4 score on the benchmark with no budget or safety constraints [7], even though the paper's own controlled evaluation put the best full-solve rate at just 31.5% [6]. Community discussion elsewhere noted that the same top-scoring model separately performed far worse when asked to reproduce a program from documentation rather than explain what an existing binary does, prompting some observers to argue that knowing what code does and reproducing it look like almost separate skills, and that a single headline score under loosened conditions can overstate general capability. That split mirrors a broader skepticism visible across discussion of agent benchmarks generally: practitioners point to prior cases where simplistic or idle baseline agents have outscored genuinely capable models on tool-use leaderboards, and where fixing grading bugs in established coding benchmarks reshuffled a meaningful share of the rankings - evidence that a single aggregate score can reflect benchmark artifacts as much as real capability.

Why 'Agents Are Systems, Not Models' Matters

The most consequential paper in this batch may not be a benchmark at all. 'Agents Are Systems, Not Models' argues that scoring an LLM-based agent as if it were a single fixed model obscures what is actually driving performance: the task information it receives, whether it has a dedicated self-verification tool versus just being prompted to check its own work, the time budget it is given, and the backbone model itself [8]. The paper's starkest finding is that about 54% of outcome variance in its released 18,000-plus agent trajectories comes from simply re-running the identical configuration, not from changing the model or the task [8]. It also shows these levers interact rather than add up: extra time budget only helps an agent that already has adequate task information or a sufficiently capable backbone model [8]. That framing lines up with what practitioners trying to benchmark coding agents have been converging on independently - freezing the harness and swapping only the model, logging every tool call, and scoring against an evaluator the agent never saw - precisely so that harness and configuration differences stop masquerading as model capability differences.

Historical Context

2026
Built from more than 5,000 expert-hours of clean-room program development specifically to avoid the training-data contamination that undermined earlier reverse-engineering evaluations.
2026-09-30
Submitted to arXiv, introducing a 94-task benchmark for agents working through computational materials science research steps.
2026-09-30
Submitted to arXiv, introducing a 26-task, six-domain benchmark for agentic scientific discovery.
2026-10-01
Submitted to arXiv, proposing that agents be evaluated as configurable systems rather than fixed models, releasing over 18,000 agent trajectories.
2026-10-02
Published commentary noting that KaliBench caps open-weight models at 42% exact-command accuracy.

Power Map

Key Players
Subject

New AI Agent Evaluation Benchmarks Reveal Where Agents Still Fail

CH

Chenmu Zhang and co-authors, including Boris Yakobson

Creators of CompMat-Bench, the computational materials science agent benchmark

JI

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen and co-authors

Creators of EurekaBench, the scientific-discovery agent benchmark

KA

Kapa.ai

Builder of Company Knowledge Bench and vendor of agentic enterprise retrieval products evaluated by it

RI

RISys Lab (Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer)

Creators of KaliBench, the cybersecurity CLI-command benchmark

JE

Jeremy Spence, Nicholas Assaderaghi and co-authors

Creators of SRE-Bench, the contamination-free binary reverse-engineering benchmark; Vals AI separately hosts the public leaderboard built on it

LU

Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata

Authors of 'Agents Are Systems, Not Models', proposing agents be evaluated as configurable systems rather than fixed models

Fact Check

8 cited
  1. [1] CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
  2. [2] EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
  3. [3] Benchmarking retrieval for agents on messy real-world company knowledge
  4. [4] KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
  5. [5] RISys-Lab/KaliBench (GitHub repository)
  6. [6] The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark
  7. [7] SRE-Bench Leaderboard
  8. [8] Agents Are Systems, Not Models: Rethinking Agentic Evaluation

Source Articles

Top 5

THE SIGNAL.

Analysts

“Open-weight security copilots currently fail the exact-command standards real operators require, but KaliBench's verifiable-reward training recipe lets a small 8B model match a 685B mixture-of-experts model, giving security tool vendors a cheaper alternative to frontier API dependence.”

AI Weekly editorial commentary (Alexis Dufresne)
Commentary on KaliBench results

“Verification behavior changes substantially depending on whether an agent is given a dedicated verification tool versus relying on prompting alone, and some desired agent behaviors are more reliably achieved by changing the system than by changing the prompt.”

Luis Wiedmann and co-authors ('Agents Are Systems, Not Models')
Researchers proposing an agents-as-systems evaluation framing

“Models that perform well at source-level vulnerability discovery and patching do not automatically transfer that competence to reverse engineering a compiled binary and explaining its behavior.”

Jeremy Spence, Nicholas Assaderaghi and co-authors (SRE-Bench)
Researchers who built SRE-Bench

“Current agent benchmarks are easy to reward-hack: a do-nothing agent has outscored strong models on tool-use leaderboards, a large share of benchmark-graded 'correct' outputs have turned out to be wrong on inspection, and fixing grading issues on established coding benchmarks has reshuffled a meaningful share of leaderboard rankings - a problem he argues becomes dangerous once such scores feed into policy decisions.”

Daniel Kang (FAR.AI)
Speaker at the FAR.AI Alignment Workshop on benchmark validity
The Crowd

“New cybersecurity benchmark: SRE-Bench. Models are starting to become competent at finding and patching security vulnerabilities in source code, but can they reverse engineer a binary and understand its behavior? Many existing cybersecurity benchmarks test model capabilities...”

@@ValsAI226

“OpenAI's GPT 6 Astra has effectively saturated SRE-Bench, a cybersecurity benchmark testing whether models can reverse engineer binaries. We're sharing more information, along with a call for contributions to extend the benchmark:”

@@ValsAI160

“AI scientist agents are great at optimizing and finding solutions through endless trial and error. But is that really what science is about? As Terrance Tao eloquently put, the role of math and science should be more than that - they are "lighthouses" that guide and inspire...”

@@JiayiiGeng120

“Coding benchmarks that are quickly showcasing deep capability”

@u/Informal-Trouble218381
Broadcast
Agent Evaluation & Benchmarks - Agentic AI MOOC 2025 Lecture 4 Summary

Agent Evaluation & Benchmarks - Agentic AI MOOC 2025 Lecture 4 Summary

Daniel Kang - AI Agent Benchmarks Are Broken [Alignment Workshop]

Daniel Kang - AI Agent Benchmarks Are Broken [Alignment Workshop]

AI Agent Benchmarks Are Finally Getting Real

AI Agent Benchmarks Are Finally Getting Real

New AI Agent Evaluation Benchmarks Reveal Where Agents Still Fail — AI News | Agentic Brew