The method — four claims
Four claims about how to build a game that measures something: score calibration not knowledge, put a ceiling in the set, derive levels from a graph, and treat the mandate as a by-product.
Source: https://games.sgit.ai/method/index.html↗ · site v0.5.0 · this file is generated from the same content
as the page, so the two cannot drift. Every page on this site has a .md twin; internal links
below point at them.
Four claims
Each of these is a decision that could have gone the other way, and each names what would show it was wrong. They are the transferable part: nothing here is specific to agents, and all four apply to any game meant to measure a belief.
1 — Score calibration, not knowledge — A proper scoring rule, with a wrong answer costing more than a right one earns, and don't know always available and always worth zero. The argument↗
2 — Put a ceiling in the question set — If every question has a real answer, the only error the game can measure is underestimating. Two fifths of these are things nothing can do. The argument↗
3 — Derive levels from the graph, not from vibes — A level is the number of hops to a one-way consequence, computed from the mesh — and it is distance, never a danger rating. The argument↗
4 — Make the mandate a by-product — Nobody fills in a mandate form honestly. Ask do you want it to? forty times and one falls out of the play. The argument↗
What they have in common
All four are ways of stopping a game from measuring the wrong thing while looking like it is working. A quiz with no ceiling, no scoring rule and hand-assigned levels still produces a number, still feels informative, and tells you nothing you did not put in. The four claims are what make the number arguable — and each one has a self-test in the vault that fails the build if it stops holding.