
benchmarks safetyNew
METR
Independent frontier-model dangerous-capability evaluator
weight 0.0Flux 62▴FreeLaunched 2026-06
💸 No earnings reported yet
What it is
METR (Model Evaluation and Threat Research) is a nonprofit that runs autonomous-capability and dangerous-capability evaluations of frontier models, including pre-deployment testing for major labs and AI Safety Institutes.
How AI plugs in
Designs and runs empirical evaluations measuring AI systems' autonomous and dangerous capabilities.
Alternatives & related tools
Benchmarks & SafetyResearch
Apollo Research
New
AI deception and scheming evaluation lab
Frontier Labs
Benchmarks & SafetyResearch
MLCommons
New
Open engineering consortium behind MLPerf and AI safety benchmarks
Frontier Labs
Benchmarks & SafetyResearch
Transluce
New
Open interpretability and AI-oversight nonprofit
Frontier Labs

Benchmarks & SafetyResearch
Artificial Analysis
New
Independent AI model benchmarking
Frontier Labs

Benchmarks & SafetyResearch
Epoch AI
Research and benchmarks tracking AI progress
Frontier Labs
Benchmarks & SafetyResearch
LMArena
New
Crowdsourced human-preference model leaderboard
Frontier Labs
★ Reviews
No reviews yet — be the first.Your rating
