◉ PSYCHOHISTORY

TLO Benchmark (The Last Ones)

Idea
A 32-step simulated corporate network attack built by the UK's AI Security Institute — from initial scouting to full network takeover, estimated at about 20 hours of professional human hacker labor. An unreleased AI model became the first in history to complete it start to finish (3 of 10 runs, averaging 22 of 32 steps); the significance, per the entry, is that 'too dangerous to release' now rests on a specific measurable test the public version cannot pass.
The engine's definition — word for word
UK AI Security Institute custom-built 32-step corporate network attack simulation range. Spans initial reconnaissance to full network takeover. Estimated to require ~20 hours of human cybersecurity-professional labor. Claude Mythos Preview became the first model in history to solve start-to-finish (3/10 full completion, avg 22/32 steps). Structural significance: TLO is the mathematically quantified capability delta that justifies the gating. 'Too dangerous to release' is operationalized via a specific benchmark the public tier cannot match. Watch for similar custom state-level benchmarks emerging for adversary-side capability verification.
Walk this on the live map →
Part of the Psychohistory engine — 2,750 entities, 6,993 documented connections. Open data, built to be proven wrong.