The self-score, opened up (§XVII)
How the live score is built
The model scores itself against its own twelve weighted criteria — in public, forever. Each criterion starts on the design-stage self-assessment prior (≈88; an independent proxy panel scored the same text ≈74, and both are published) and only moves as real operating evidence accrues. Criteria with no evidence yet stay honestly on the prior.
Live self-score
89/100
Weighted across the twelve criteria.
Weighted mean
8.86/10
Still on the paper prior
11/12
Criteria awaiting real operating evidence before they can move.
Per criterion
| Criterion | Prior | Live | Evidence | Confidence | Source |
|---|---|---|---|---|---|
| Legitimacy & consent | 9.0 | 9.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
| Outcome quality | 8.0 | 8.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
| Rights protection | 9.0 | 9.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
| Accountability | 9.0 | 9.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
| Capture & corruption resistance | 9.0 | 9.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
| Transparency & verifiability | 9.0 | 9.4 | measured | 40% | audit ledger verified intact |
| Representation & proportionality | 9.0 | 9.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
| Resilience | 9.0 | 9.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
| Adaptability & self-correction | 9.0 | 9.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
| Intergenerational fairness | 9.0 | 9.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
| Simplicity & usability | 7.0 | 7.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
| Inclusiveness | 9.0 | 9.0 | paper prior | 0% | design prior (no operating evidence yet, §XVII.3) |
An honest 88 is the point: the model’s own §0.6.5 forbids declaring a paper 10/10. Outcome quality (Criterion 2) stays a hypothesis until real decisions are measured against their predictions.