AERIOXFLUX
Frontier Labs
benchmarks safetyNew

Terminal-Bench

Benchmark for AI agents that actually operate a terminal

weight 0.0Open SourceLaunched 2026-08-10

💸 No earnings reported yet

What it is

Terminal-Bench evaluates how well AI agents complete real end-to-end tasks inside a sandboxed shell — installing dependencies, running builds, debugging failures, and driving multi-step command-line workflows. It has become one of the reference scores for agentic coding ability, cited alongside SWE-bench when comparing frontier and open-weight models, and ships a harness so teams can run their own agents against the task suite.

How AI plugs in

Alternatives & related tools

★ Reviews

No reviews yet — be the first.

Your rating

Discussion (0)

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.