
Terminal-Bench
Benchmark for AI agents that actually operate a terminal
💸 No earnings reported yet
What it is
Terminal-Bench evaluates how well AI agents complete real end-to-end tasks inside a sandboxed shell — installing dependencies, running builds, debugging failures, and driving multi-step command-line workflows. It has become one of the reference scores for agentic coding ability, cited alongside SWE-bench when comparing frontier and open-weight models, and ships a harness so teams can run their own agents against the task suite.
How AI plugs in
Alternatives & related tools

Artificial Analysis
Independent AI model benchmarking

Epoch AI
Research and benchmarks tracking AI progress
Apollo Research
AI deception and scheming evaluation lab
LMArena
Crowdsourced human-preference model leaderboard
METR
Independent frontier-model dangerous-capability evaluator
MLCommons
Open engineering consortium behind MLPerf and AI safety benchmarks
★ Reviews
No reviews yet — be the first.Your rating
