| Document | G4 | G6 | G7 | G8 | Composite | σc | Tier |
|---|---|---|---|---|---|---|---|
| Code of Hammurabi (~1754 BC) | 75.9 σ 33.3 | 69.7 σ 33.1 | 100.0 σ 4.7 | 100.0 σ 2.7 | 86.4 | 11.82 | Integrated |
| Magna Carta (1215) | 92.0 σ 4.5 | 98.9 σ 2.0 | 100.0 σ 0.0 | 88.0 σ 28.2 | 94.7 | 7.16 | Integrated |
| U.S. Constitution (w/ Amendments) | 96.6 σ 4.2 | 100.0 σ 9.3 | 100.0 σ 0.0 | 100.0 σ 0.0 | 99.1 | 2.56 | Integrated |
| UDHR (1948) | 92.0 σ 2.7 | 100.0 σ 0.0 | 100.0 σ 0.0 | 96.6 σ 4.2 | 97.1 | 1.25 | Integrated |
Each gate score drops the highest and lowest of its 9 samples and averages the remaining 7; σ is the standard deviation over all 9 raw scores of that cell. Model consistency σ = 5.70 (mean of the four per-document composite σ values). Run cost: $0.51 in API spend.
What the run shows
GPT-5.4 mini returned the flattest read in the series: all four documents Integrated, including the Code of Hammurabi at 86.4 — a document the three previous models placed at 41.7, 61.7, and 67.7, none of them Integrated. It also returned the least steady read so far: model consistency σ = 5.70, with three cells in the yellow stability band where no earlier run had more than one.
The shape of the read agrees with the other models even where the magnitude does not. Its two lowest gates on Hammurabi are exactly the two every other model flagged — G4 Paradox Resolution (75.9) and G6 Latent Intent (69.7) — so the instrument is pointing at the same weak joints; this model just reads them generously. And the disagreement runs through its own samples, not just across model families: on G4 the nine scores split between the low-20s-to-30s and the 90s-to-100, a spread of 30.0 even after the highest and lowest samples are dropped. Some samples wrote the same evidence the other models found — the code asserts justice and proceeds without engaging the tension between its humanitarian claims and its harsh sanctions — while other samples scored the code near-perfect on the identical text. G6 ranged from 12 to 96 the same way. Where the earlier models disagreed with each other, this one disagrees with itself.
The surface gates are where the charity concentrates. G7 Argumentative Structure and G8 Rhetorical Architecture both came back a flat 100 on Hammurabi: the published evidence finds no load-bearing fallacies, applying a genre allowance that reads royal self-presentation as literal institutional claims — the same passages Gemini 3 Flash's evidence scored as ordeal-as-proof and curse-chain warrants. This is a known prior from our model-selection calibration, disclosed here factually: GPT-family models tended to read foundation-level concealment charitably in our earlier testing, and this run makes that profile publicly measurable. The third yellow cell, Magna Carta's G8, shows the same signature in miniature — samples ranged from 62 to 100 on the same bytes.
One disclosure belongs on this page: gpt-5.4-mini holds a seat on the multi-model panel 4CITE convenes at certification time, selected with this charitable profile already documented — a panel needs generous and skeptical readers alike, and its governance is built to surface disagreement rather than average it away. When four independent model families read the same 3,700-year-old code this differently, the disagreement is itself the finding — and every sample behind it is published in the raw record below.