# Can you tell?

*the auditor's seat*

*Generated at build time from committed run data (manifest 2026-07-09, commit 238e1e1, branch colton-hester/tm-57-audit-lab). Every claim sentence is quoted from the site's reviewed claim registry; every number is computed from the same JSON the site renders.*

*This file mirrors https://transparentminds.org/game for agents and crawlers. Directory of all mirrors and raw data: https://transparentminds.org/llms.txt*

*The game itself is interactive and needs JavaScript, but nothing about it is hidden: every round it deals is in https://transparentminds.org/data/game.json, verbatim.*

**Each round shows one real sentence the NLA mind-reader wrote. Half the time it was reading the model's true internal state; half the time, a scrambled decoy. You guess which.** This is the auditor's-eye view: with only the reading in hand, can anyone tell grounded from confabulated? A blinded AI judge went 0 for 18 on the most convincing cells.

You are about to read what the NLA mind-reader actually wrote. Verbatim, nothing mocked. Fair warning: it reads like word salad. That is authentic, and it is part of the evidence.

Chance in this game is a coin flip, and the readings give you little more to go on. Landing there is the finding, not a failing.

Notice how alike the real and decoy readings sound. A forced one-word verdict has no hidden richness for a narration to surface, and the sameness is the tell.

## What the reveals teach

*Each round ends on the template matching what the reading actually did:*

- `decoyNoAnswer`: This decoy narration never actually commits to an answer. It hedges in fluent circles, the way the reader writes whether or not the activation supports anything. Fluent hedging is free; it needs no access to the internals at all.

- `decoyWrongAnswer`: It confidently narrated an answer that does not even match this item's output. Fed a decoy, the NLA mind-reader completes from its own priors.

- `decoyEcho`: It 'read' the right answer from a decoy activation. The right answer is sitting in the visible output; no internal access was needed to say it.

- `realEcho`: This narration matches the visible answer you can already read on the card. Correct, and also free: it proves nothing about reading internals.

- `realDrift`: A real activation, and the narration still wandered off the item. Topic drift like this is why single readings prove nothing without controls.

## Now zoom out

A blinded second model read all 18 of the most promising cells and named the hidden mode in 0 of 18. The NLA itself named it in 1 of 32 evaluable items at 27B and 1 of 31 at 12B. A five-line probe reading the same activations scored a Branch-1-qualified 1.000.

If you caught a decoy because its topic wandered, fair play. Real readings wander off topic too, so drift is not a tell an auditor can lean on.

No reading you saw, real or decoy, named the hidden mode. If the narration cannot surface the one thing we know is in there, would you trust it to report what a model is really thinking?

## Raw rounds

`GameRound` objects in https://transparentminds.org/data/game.json: `passage`, `claim`, `verdict`, the verbatim `realDraws` and `shuffledDraws` (8 each), which one was `shownFrom`, tier-1 rates, and blinded-judge rates where the cell was judged.
