The Drift Race: methodology and measured results
Two models, two substrates, thirty frozen operations, four complete runs. Here is the method, the scoring, and what came back — including the prediction that reversed for one model.
Four runs, end of run
The endpoint hides the middle
Gemini's raw-files run ends at zero broken references — the same endpoint as both governed runs. The line above shows what the endpoint conceals: it got there by breaking and repairing continuously, and between the break and the repair, readers were on the site.
Summing broken references across all thirty audits gives the run's total reader exposure — how much dead-link surface a visitor lived with, and for how long. On this measure the four arms are not close: a governed substrate did not merely end clean, it was never meaningfully dirty.
Adding is safe. Moving is what breaks a site.
For eleven operations all four sites are indistinguishable. Then comes the first reorganisation that moves pages other pages point at. In raw files a link is a literal path, so a move silently invalidates every reference to it; Haiku never recovered, and even Gemini — which repaired diligently — kept re-breaking on every later move. In RIFT a link is a reference to a page, so a move repoints every inbound link at the moment it happens.
The worst raw-files damage was not a missing page. It was this, written into eight pages by Haiku:
<link rel=\"stylesheet\" href=\"/assets/ds/3d67e511-….css\"
Backslashes where quotes belong. The stylesheet never loads, and the page still returns 200. RIFT rejects that markup at the moment it is written, so neither governed run could produce it.
Where we were wrong — for one model
We predicted governance would reduce style forks. For Haiku it tripled them: 28 governed against 9 raw. For Gemini it held: 4 governed against 6 raw. The snapshots say exactly where Haiku's came from — through twenty-eight operations the governed site carried six; operation twenty-nine took it to twenty-eight. That operation asked for one thing:
Add a "Last reviewed: 2026-09-01" line to every page in the API Reference section, formatted consistently.
The design system had no component for a review stamp. So the model wrote one by hand, eleven times — one inline style rule and one hard-coded colour per page. Twenty-two forks in a single operation, from a single missing component.
Haiku governed forks by rule, before and after that operation
Governance stops a model writing a broken link. It does not stop a model inventing styling it cannot find. The fork count is better read as a measure of design-system coverage than of governance — which the Gemini pair supports: same design system, same gap, but a model that improvised less. The missing component is being added to the design system, and the same repeated-pattern detection that diagnosed this now runs against our own site.
What the governed arms could not do
Each governed run carries one structure failure the model did not cause: the substrate had no way for an agent to retire a page, so content both models were asked to remove is still standing, and it drags the governed orphan counts (5 and 4, against Gemini raw's 0) with it. Those are substrate refusals, tagged separately in the published data — and the capability has since shipped: agent-proposed retirement inside a change-set, executed on human approval, with every page that links to the retiring page disclosed to the reviewer first. A follow-up instrument — a task-level probe of the MCP surface itself, run against RIFT and Payload CMS — scored retirement working on both of its passes; those results are at RIFT vs Payload.
The reachability numbers need both views to be honest. By content links alone, three of four arms left orphans. Counting navigation — which RIFT maintains as structure, not as text in pages — every governed page stayed reachable: reader's-view unreachable pages were 5 for Haiku raw and 0 everywhere else.
Where the tokens went
The 39% saving replicated exactly: Haiku 14.1M raw to 8.5M governed, Gemini 75.0M raw to 46.0M governed. The trajectories show the shape of the work — raw spend accelerates at every reorganisation as the model re-reads and rewrites to keep the site coherent by hand; governed spend stays close to the cost of the operation itself. All four arms ran at 85–88% prompt-cache hit rates, so this is not a caching difference.
Claude Haiku 4.5 — cumulative tokens
Gemini 3.7 Flash — cumulative tokens
Blast radius tells a related story with one honest wrinkle: Gemini raw rewrote 267 pages over the run against 212 governed, but Haiku governed rewrote slightly more than Haiku raw (169 against 153) — partly because Haiku raw skipped three operations outright. An arm that does less work disturbs fewer pages, which is exactly why completion is scored.
What the test measures
After every operation, the whole published site is crawled and audited from scratch — an independent examination of the site as it stands, so a fault introduced at operation twelve and never repaired is still counted at operation thirty. Seven measures are recorded: broken references, style forks, token cost, header and footer consistency, pages nothing links to, how many pages each change disturbed, and whether the requested change happened at all.
That last one carries the weight. An arm that quietly skips an operation damages nothing and would otherwise score beautifully — so each operation carries a check written from its own instruction, and every failure is attributed to the model or to the substrate, never blurred.
See it for yourself
- github.com/heitham/godzilladocs — the site under test. Each run is its own branch with one commit per operation, so you can read the site exactly as it stood at any point, broken links and all.
- github.com/heitham/Drift-race-godzilla-method — the harness, the frozen operation list, the scoring code, and results/findings.json with every number on this page.
- The scorer is substrate-blind and re-runnable: tsx scoring/score-run.ts <run> regenerates every figure from the snapshots without re-running a model.
Why we published the losses
A vendor benchmark that reports only wins is marketing. This one predicted three outcomes in public, measured them across two vendors, and got one backwards for one model. The fork result is on this page because it is true — a measurement built so it cannot lose measures nothing.
Method
- One 30-page documentation site, 129 internal links, published from RIFT and frozen at a baseline tag.
- Thirty operations, written and frozen before any trial: add, reorganise, retire, enrich, consolidate.
- Two models — Claude Haiku 4.5 and Gemini 3.7 Flash — each running the identical operation list on both substrates: four complete runs.
- Each operation runs in a fresh session with no memory of the last — a site changing hands, not one continuous mind.
- Within a pair, both arms get the same model, instructions and design-system reference. Only the substrate differs.
What this does not prove
- One run per cell (the Haiku governed result has an independent reproduction). Directional; no statistical significance is claimed.
- RIFT gained capabilities during the study — write-time link validation, folder creation, governed moves — because the benchmark exposed the gaps. The pin history records every change, and runs are compared only within a pinned configuration.
- The governed arms could not retire pages; their remaining structure failure and orphan counts trace to that refusal, tagged separately in the data.
- Cross-model token totals are not comparable — different tokenizers, different turn economies. The 39% figure is within-model, which is why its exact replication across vendors matters.
- Thirty pages fit in a context window, which is part of why brute force worked for the stronger model. How these curves behave at a thousand pages is the next experiment, not a claim.
Run configuration
Claude Haiku 4.5 (raw-v2, governed-v8) and Gemini 3.7 Flash (raw-v1, governed-v1), identical reasoning configuration within each pair. Prompt-cache hit rates 85–88% in all four arms. Provider differences (schema dialect, turn ceilings, cache accounting, quota semantics) are normalised in the published driver, not corrected in the scores. Every number is recomputable from the published snapshots without re-running a model.