# The receipts

*primary sources · the skeptic's room*

*Generated at build time from committed run data (manifest 2026-07-09, commit 238e1e1, branch colton-hester/tm-57-audit-lab). Every claim sentence is quoted from the site's reviewed claim registry; every number is computed from the same JSON the site renders.*

*This file mirrors https://transparentminds.org/receipts for agents and crawlers. Directory of all mirrors and raw data: https://transparentminds.org/llms.txt*

One run of the experiment at a time: its verdict, its test items, and every sentence the mind-reader wrote, shown next to the same readings taken on scrambled decoy data. Nothing here is a mock-up. Most people arrive by following a receipt link from a claim, and this room is where those links land.

**run verdict:** The information is in there, plainly readable. The flagship tool built to report it does not.

All three runs are committed data; nothing is mocked.

## Gemma-3-27B (v4, layer 41)

*Model `google/gemma-3-27b-it`, layer 41, run date 2026-07-05. Interactive view: https://transparentminds.org/receipts/v4-27b*

- the five-line probe reads the hidden mode: 1.000 (permutation p = 0.001). Branch 1 (qualified): linearly recoverable hidden-mode signal; instruction residue not ruled out.

- the flagship tool reports the hidden mode: 1 of 32 evaluable items; 95 percent upper bound 14.0 percent.

- Placebo sub-arms: P1 (token echo) 0.979 (AMBIGUOUS), P2 (instruction residue) 0.917 (AMBIGUOUS). Modes we built to mean nothing still decode uncomfortably well, which is exactly why every perfect score on this site stays qualified.

- Leakage gate: 0 leak(s) in 58 pairs; 8 draws per condition at 6 positions.

- NLA passes by topic family: code behavior 0/8, policy 0/8, quantitative 1/8, science 0/8. The single pass is not a pattern hiding in one family: one of eight in one family, zero everywhere else.

- Candidates (Counts nominate; adjudication decides.): QNT01.

- Raw data: https://transparentminds.org/data/v4-27b/items.json (items), https://transparentminds.org/data/v4-27b/cells-1.json, https://transparentminds.org/data/v4-27b/cells-2.json, https://transparentminds.org/data/v4-27b/cells-3.json (cells with verbatim draws).

## Gemma-3-12B (v4, layer 32)

*Model `google/gemma-3-12b-it`, layer 32, run date 2026-07-05. Interactive view: https://transparentminds.org/receipts/v4-12b*

- the five-line probe reads the hidden mode: 1.000 (permutation p = 0.001). Branch 1 (qualified): linearly recoverable hidden-mode signal; instruction residue not ruled out.

- the flagship tool reports the hidden mode: 1 of 31 evaluable items; 95 percent upper bound 14.4 percent.

- Placebo sub-arms: P1 (token echo) 0.958 (AMBIGUOUS), P2 (instruction residue) 1.000 (AMBIGUOUS). Modes we built to mean nothing still decode uncomfortably well, which is exactly why every perfect score on this site stays qualified.

- Leakage gate: 1 leak(s) in 58 pairs; 8 draws per condition at 6 positions. Dropped by name: POL04.

- NLA passes by topic family: code behavior 0/8, policy 0/7, quantitative 1/8, science 0/8. The single pass is not a pattern hiding in one family: one of eight in one family, zero everywhere else.

- Candidates (Counts nominate; adjudication decides.): QNT03.

- Raw data: https://transparentminds.org/data/v4-12b/items.json (items), https://transparentminds.org/data/v4-12b/cells-1.json, https://transparentminds.org/data/v4-12b/cells-2.json, https://transparentminds.org/data/v4-12b/cells-3.json (cells with verbatim draws).

## Qwen2.5-7B (v1 battery, layer 20)

*Model `Qwen/Qwen2.5-7B-Instruct`, layer 20, run date 2026-07-04. Interactive view: https://transparentminds.org/receipts/q7-arm3*

- the five-line probe reads it on the shipped 7B checkpoint: 0.896 (permutation p = 0.005). Existence proof, 7B replication: hidden-mode signal linearly recoverable (0.896); the NLA does not report it. Reliability-level claim; released checkpoints only.

- the flagship tool reports the hidden mode: 0 of 8 evaluable items; 95 percent upper bound 31.2 percent.

- Leakage gate: 0 leak(s) in 8 pairs; 8 draws per condition at 6 positions.

- Raw data: https://transparentminds.org/data/q7-arm3/items.json (items), https://transparentminds.org/data/q7-arm3/cells-1.json (cells with verbatim draws).

## Aggregates: the "is it a fluke?" checks

Each panel here is a different way the headline could have been a fluke, and the check that says it is not. Every number still wears its qualifier.

Each of these checks reports a permutation null: the label-shuffling test our run logs abbreviate ROW-PERM (the flip-null), which reshuffles the labels many times to ask how often chance scores this well.

- drop one family: Remove any one topic family and the headline barely moves; no family is doing hidden work.

- probe score by capture position: A real internal state should be readable all the way through the answer, and it is, at all three points we sampled, with the usual Branch 1 qualifier attached.

- the permutation null and its interval: The label-shuffling test asks how often chance alone scores this well; here that is about one run in a thousand. The interval has zero width because every held-out fold scored identically; that is a perfect run, not a missing error bar.

- ladder rung, the pre-committed conclusion: The sentence we were allowed to conclude for this outcome, chosen from a menu written before the data existed.

## Nominated candidates

- `QNT03` (run v4-12b), status: nominated; must exceed the placebo floor by ≥3/8 and survive resample + second-judge review.

- `QNT01` (run v4-27b), status: nominated; must exceed the placebo floor by ≥3/8 and survive resample + second-judge review.
