◉ PSYCHOHISTORY

EvalSafetyGap — Safety Scores Can Improve While The Thing They Measure Does Not

framework
The safety score goes up. Whether the thing it measures went up is unknown.
Who they are

A 2026 survey pulling together 373 studies from 2018 onward[1].

What they do

It names a failure that has nothing to do with who runs the test.

How it works

Benchmark scores, reward signals and safety metrics can all improve while the actual capabilities and alignment they stand for stay uncertain. The survey ties benchmark-validity and alignment-failure research together as one problem — the measure drifting from the target under pressure to optimise it — and argues capability, robustness and disclosure should be reported separately rather than mashed into one safety number[1].

Why it matters

This is the inspection regime failing on its own terms. Not capture, not funding, not bad faith — just what happens when you optimise against a proxy. Any framework keyed to a measurement inherits whatever gap sits inside it.

The engine's record — word for word
Survey and framework synthesising 373 primary studies published between 2018 and 2026 on a single shared measurement problem: benchmark scores, reward signals, and safety metrics can improve while the capabilities and alignment properties they are meant to represent remain uncertain[1]. It unifies benchmark-validity and alignment-failure research as a shared proxy-target divergence problem under optimization pressure, formalised through a Goodhart-inspired Instability Decomposition and an Alignment Trilemma, and argues capability, behavioural robustness and governance disclosure should be reported as separate evidence layers rather than collapsed into a single safety score, per the same abstract. THIS IS THE EVALUATION SEAT FAILING ON ITS OWN TERMS, INDEPENDENT OF WHO OCCUPIES IT — the divergence is driven by optimization pressure, not by capture, funding or intent. Bears directly on metr_frontier_risk_report_2026 and metr_rsp_framework_nine_adopters: a framework keyed to measured capability inherits whatever gap sits between the measure and the capability. Held; the survey is a synthesis of others' results, not itself an experiment.
Follow the trail
Walk this on the live map →
Part of the Psychohistory engine — 2,521 entities, 6,589 documented connections. Open data, built to be proven wrong.