A Benchmark for Every Domain
In late September and early October 2026, four new agent benchmarks and one evaluation-methodology paper arrived within days of each other, joining SRE-Bench - an earlier 2026 reverse-engineering benchmark still being actively tested and discussed - to form a cluster of narrow, domain-specific agent evaluations. CompMat-Bench tests whether agents can execute real steps from published computational materials science studies without needing to rerun expensive simulations [1]. EurekaBench pushes agents through 26 long-horizon scientific-discovery tasks spanning neuroscience, geophysics, astrophysics, computer science, plasma physics, and chemistry, tied to 306 target insights [2]. Company Knowledge Bench, built by Kapa.ai from 1,000 real enterprise retrieval cases, measures how well agents find exactly the right internal documentation without drowning the context window in noise [3]. KaliBench grades 8,504 natural-language-to-command translations across 1,642 Kali Linux tools and 23 capability dimensions [4], with code and data published openly [5]. And SRE-Bench forces agents to reverse-engineer compiled binaries from 19 private programs built from scratch specifically to avoid training-data contamination [6]. The common thread is a shift away from generic coding and chat leaderboards toward narrow, expert-curated domains where correctness is harder to fake.

![Daniel Kang - AI Agent Benchmarks Are Broken [Alignment Workshop]](https://img.youtube.com/vi/4iyMb0ARiao/mqdefault.jpg)
