That softmax entropy and participation ratio decrease with inverse
temperature is classical. We contribute the certificate, not the phenomenon
— decide a rational-kernel substitute in exact arithmetic, enclose the softmax
story with sound interval exp/log, keep consecutive boxes
separated, and keep planted mutants red — on one frozen attention row.
One causal attention row from a tiny GPT
(layer 0, head 0, seed 0), frozen once. Teal is the rational kernel
w∝(1+β·s)2 — the decidable substitute.
Slate is softmax. Mutant chips must destroy the decrease; if they do not, the
instrument is theatre.
Loading instrument…
Live float view for intuition. Download the verifier for the exact rational decision; softmax enclosure ships with the page certificates (Node + eqcert).
Softmax entropy and participation ratio decrease with
inverse temperature is classical — majorization ordering attributed by Mattei &
Loureiro (arXiv:2602.14862) to Marshall &
Olkin, with dH/dβ=−β Varp(s) their Prop. 2
(also Dabah & Tirer, ICML 2025). We contribute the certificate, not the
phenomenon: on frozen GPT-tiny scores we
decide PR↓ for
w∝(1+β·s)p in exact rationals, and
enclose softmax H/PR on the committed grid
(and continuous-β sign via those identities) with sound interval
exp/log, consecutive boxes
hi(βi+1)<lo(βi), and planted mutants
that go red.
Scores s are the last-query causal
attention logits from a tiny GPT at seed 0 (31 positions). Grid
β ∈ {0.25, 0.5, 1, 1.5, 2, 3, 4, 6, 8}. Nothing here is trained; the fixture is
committed and sha256-pinned in each certificate.
dH/dβ=−β Varp(s). Cite and enclose
— do not re-derive.A — Softmax (enclosed).
Bounds that contain the truth, not a float that looks tidy.
Shannon entropy and participation ratio of softmax(β·s) decrease on the
grid after sound interval enclosure; consecutive boxes separate. Continuous-β sign on
(0,∞) for this non-flat fixture follows the published identities
dH/dβ=−β Varp(s) (Mattei & Loureiro Prop. 2)
and d(Σp²)/dβ=2 Covp(p,s). Certs:
CERT-ml-beta-softmax.json, CERT-ml-beta-continuous.json.
B — Rational kernel (decided).
When intervals hesitate, rationals finish the argument.
On the substitute
wi ∝ (1+β·si)2
(same spirit as
rational attention),
participation ratio is strictly decreasing — decided by exact BigInt comparison.
Sibling p=1 and dual Σw²↑ also decide. Cert:
CERT-ml-beta-pr.json.
Falsifiers (must stay red).
A claim that cannot fail is not a claim.
Flatten s; replace β·s by (β−3)²·s;
control that the true curve is not increasing.
The mechanism — softmax concentration under temperature — is occupied. The wedge this note claims is the machine-checkable decision with falsifiers on a committed fixture.
| source | relation | here |
|---|---|---|
| Mattei & Loureiro, arXiv:2602.14862 | Prop. 2: dH/dβ=−β Var; Thm 3.1 majorization |
Cite; enclose; do not derive |
| Dabah & Tirer, ICML 2025 (arXiv:2402.05806) | Temperature / majorization for softmax | Cite as mechanism |
| ATMA (arXiv:2606.25156) | PR as a count channel in architectures | Same statistic, different job |
| Geshkovski et al., NeurIPS 2023 | Token particles cluster under attention | Sibling object; not this claim |
| Poly-attention / sinks literature | Efficiency, KV practice, training entropy | Crowded; out of scope |
When the arithmetic cannot represent the model
honestly, change the model until the claim is decidable — keep the phenomenon,
keep the falsifiers. Rational PR↓ is owned by a short Chebyshev-association identity on this fixture (exact rationals). That rule is the load-bearing transfer from
het-agent macro with only + − × ÷
and
rational attention.
Softmax enclosure uses the shared interval
exp/log in
interval transcendentals
(eqcert.
Scope. Continuous-β H/PR holds for this frozen non-flat score vector (Var + Cov identities). Not a claim about trained dynamics, token clustering, or attention efficiency. The three JSON certificates pin the fixture and verifier by sha256; re-run them to reproduce.
CERT-ml-beta-{pr,softmax,continuous}.json)