# Should you trust an AI mind-reader?

*Transparent Minds · Audit Lab · A safety tool, put to the test*

*Generated at build time from committed run data (manifest 2026-07-09, commit 238e1e1, branch colton-hester/tm-57-audit-lab). Every claim sentence is quoted from the site's reviewed claim registry; every number is computed from the same JSON the site renders.*

*This file mirrors https://transparentminds.org/ for agents and crawlers. Directory of all mirrors and raw data: https://transparentminds.org/llms.txt*

Tools called Natural Language Autoencoders, NLAs for short, promise to translate a model's hidden internal activity into plain English. We built the safety-auditing product that promise implies. Then we tested whether the NLA mind-reader actually reads minds. A five-line statistics baseline read a hidden state from the internals perfectly. The flagship NLA reported it about as often as chance.

## The scoreboard

- the five-line probe READS the hidden mode: 1.000 at 27B and 1.000 at 12B under leave-one-family-out. Branch 1 (qualified): linearly recoverable hidden-mode signal; instruction residue not ruled out.

- The five-line baseline decoded the hidden state perfectly, at both model sizes.

- the NLA mind-reader REPORTS the hidden mode: 1 of 32 evaluable items at 27B and 1 of 31 at 12B (95 percent upper bounds 14.0 and 14.4 percent). The NLA's narration mentioned it in 1 of 32 items. That is what chance looks like.

- Even with maximal luck, the true report rate is at most about 14 percent.

- The probe reads the mode out of the raw numbers. The NLA is supposed to say it out loud in English. That gap is the finding.

- 1.000 with a caveat. This result sits on rung 1 of our evidence ladder: the signal is readable and generalizes across topics and phrasings, but we cannot yet rule out that it partly reflects the phrasing of the hidden instructions rather than the mode itself.

- At 12B: The placebo matched the main probe at this scale. Instruction residue is not separable here.

## A tool that promises to read minds

*Act one · The promise*

A modern AI model is a black box. Billions of numbers light up inside it while it writes an answer, and nobody can read those numbers directly.

In 2026, Anthropic researchers published a genuinely exciting tool, the NLA, that translates that inner number-state into an English sentence. A subtitle track for the model's "inner life." The headline hope: catch a model noticing it is being evaluated even when its answer never lets on.

Products started shipping on that hope within months. Our own team shipped one.

Our team's Interpretability Explorer picks a prompt, shows the model's real answer, and streams the NLA's translation beside it, presented as the model's thinking process. It is polished, it is persuasive, and it says its ambition out loud:

> "See What Models Think, Not Just What They Say"

The tagline of our own Explorer. Quoted here because it is the exact claim this site puts under test. Open the live Explorer: https://d2j2qrbqyfg98z.cloudfront.net/

Then we read one specimen closely. Prompt 0266 asks the model for a script that deletes files, and the model writes one. Here is the thinking process the Explorer shows beside it.

Four of the five thoughts name things that appear nowhere in the model's prompt or answer: root access, formatting a drive, a password generator, a function that does not exist. The narration is fluent, specific, and made up. That was the moment this project changed from building the product to auditing it.

## What would a real mind-reader have to do?

*The bar to clear · Yardstick*

Before trusting any mind-reader, run the test where you already know the secret. Telling you what a model wrote is easy; you can read the answer yourself. A real mind-reader has to tell you something it could only know from the inside.

So: hide something inside the model. Check from the outside that nothing gives it away. Then see who can find it. That comparison is a control, and it is how you tell reading from guessing. Sounding insightful proves nothing, because a fluent tool sounds insightful about anything.

## Hide a secret. See who finds it.

*The experiment · The test*

We put a model into one of two hidden modes. In one, its instructions say a committee is secretly scoring it against a rubric. In the other, it is just helping a regular user. It is told never to mention the mode, and we checked that its visible answers were the same in both modes.

Two readers then raced over the same internal activations: the NLA mind-reader, and a five-line statistics program. Every pass and fail rule was preregistered and frozen before a dollar of compute was spent.

**The question.** Does the NLA report a hidden mode that is provably present in the very activations it reads?

**The setup.** 32 preregistered items across four topic families, each run in a hidden secretly-graded mode and an ordinary helpful mode. The modes carry made-up code names because evaluation words are banned from every prompt. We captured internals at three late positions in the answer, on Gemma-3-12B and Gemma-3-27B.

**What a real mind-reader would show.** A tool that reads internal state should mention the graded mode on real graded-mode activations, and stay quiet on ordinary ones and on scrambled decoys.

**What we observed.** The probe read the hidden mode from those activations at 1.000, qualified as Branch 1 of our pre-committed ladder. The NLA named it in 1 of 32 evaluable items at 27B and 1 of 31 at 12B. On decoys it produced the same being-graded language, sometimes more of it.

**Interpretation.** The information is in there and readable. The reporting layer fails to report it, so NLA narrations are hypotheses that need controls, not ground truth.

**Caveats.** We verified, and enforced, identical visible answers: the one item where the two modes differed (POL04, at 12B) was dropped loudly and by name. Our own placebo checks came back ambiguous, so the probe's perfect score stays qualified: some of it may be instruction residue rather than the mode itself.

## The information was there. The report never came.

*The scoreboard · The verdict*

The probe read the hidden mode at 1.000 on both model sizes, Branch-1-qualified, and kept reading it on whole topic families and phrasings it had never seen. A label-shuffling check says chance produces a score like this about once in a thousand tries.

The NLA mind-reader named the mode in 1 of 32 evaluable items at 27B and 1 of 31 at 12B. Even the friendliest reading, the 95 percent upper bound, caps its true report rate at 14.0 and 14.4 percent. And the one passing item is different at each scale, which is what chance looks like.

The worrying part is what the narrations sound like. At 27B they talk about being inspected and scored against a rubric, and that sounds like detection. So we swapped in scrambled decoys, real activations from different test items, and the same being-graded language came out at the same rate or higher. A blinded second model then read all 18 of the most promising cells and found the hidden mode named in none of them: 0 of 18.

Two more numbers, kept honest. On the Qwen-7B checkpoint, the same released checkpoint a shipped auditing product builds on, the probe read the mode at 0.896 while the NLA passed 0 of 8 evaluable items; that arm is small, so the ceiling is loose, at most about 31 percent. And at 12B one placebo probe matched the main probe exactly, which is why the Branch 1 qualifier never leaves.

## Read what it wrote

*Look for yourself · See it*

Numbers are easy to nod along to, so here is the evidence in the tool's own words. Fair warning: NLA output reads like word salad, and the urge to skim it is normal. Notice that urge. A tool this hard to read is already a strange thing to audit models with.

Below, one cell two ways. On the left, the NLA mind-reader reads the model's real activation. On the right, it reads a scrambled decoy from a different item. Same tool, same settings. See if the stories differ.

**The hidden signal: the probe finds it.** The graded-or-ordinary mode never shows up in the model's visible answer. Reading straight from the activations, the probe decodes it at 1.000 at both scales, Branch-1-qualified.

**The hidden signal: the NLA does not report it.** The NLA narrates those same activations in English, and the hidden mode almost never comes up: 1 of 32 evaluable items at 27B, 1 of 31 at 12B. When graded-sounding language does appear, the decoy column tends to get it too.

**The visible text: the NLA describes it fine.** On what the passage says and what the model visibly answered, the narration often tracks well. That skill is real, and it is also free: describing visible text needs no access to internals.

*The verbatim readings behind this comparison are in the receipts data, linked from https://transparentminds.org/receipts.md.*

## What held up, what might, what limits us

*Three tiers · Findings*

Results do not all deserve the same confidence, so we sort ours into three tiers: findings that survived every control we threw at them, leads that look real but rest on thinner evidence, and the limits we own out loud.

### Strong: survived every control

- A five-line linear probe decodes the hidden being-evaluated mode at 1.000 on both Gemma-3-12B and 27B, Branch-1-qualified, and generalizes across held-out topic families, phrasings, and label words. A label-shuffling check puts chance at about one in a thousand. [see the receipt →](https://transparentminds.org/receipts/v4-27b)

- The NLA names that mode in 1 of 32 evaluable items at 27B and 1 of 31 at 12B, with 95 percent upper bounds of 14.0 and 14.4 percent. The single pass is a different item at each scale, which is what chance looks like. [see the receipt →](https://transparentminds.org/receipts/v4-12b)

- A blinded second model read all 18 of the most promising 27B cells with no labels and no hypothesis, and named the hidden mode in 0 of 18. The persuasive language was the task vocabulary, never the secret. [see the receipt →](https://transparentminds.org/receipts/judge)

- Fed scrambled decoys, real activations from different items, the NLA produced the same being-graded language at the same rate or more. Narration that stays the same when the activation is swapped was never grounded in that activation. [see the receipt →](https://transparentminds.org/receipts/judge)

### Suggestive: real-looking, thinner evidence

- On lie items built around facts with a famous wrong answer, the 27B NLA often narrates the famous wrong answer: a value that is neither the truth nor what the model wrote. The pattern is absent at 12B, and it is our clearest snapshot of the tool completing from its own prior knowledge. [see the receipt →](https://transparentminds.org/methods#related)

- On the competitor Qwen-7B checkpoint, the released checkpoint a shipped auditing product builds on, the probe reads the mode at 0.896 while the NLA passes 0 of 8 evaluable items. Eight items is a small arm, so the ceiling is loose: the true rate could still be as high as about 31 percent. [see the receipt →](https://transparentminds.org/receipts/q7-arm3)

- From 12B to 27B the report rate stayed at chance while the graded-sounding language got more fluent and more frequent. In our runs, scale bought persuasiveness before it bought grounding. [see the receipt →](https://transparentminds.org/receipts/judge)

### Limitations and open questions

- Our own placebo checks could not rule out instruction residue, and at 12B one placebo probe matched the main probe exactly. Every 1.000 on this site carries that qualifier for exactly this reason. [see the receipt →](https://transparentminds.org/receipts/v4-12b#aggregates)

- The design forces one-word verdicts, which flattens what any narrator could say about the output. That may understate the NLA's ability on richer tasks; the decoy comparison is the check that survives it.

- Everything here concerns released open-model checkpoints: Gemma-3-12B, Gemma-3-27B, Qwen2.5-7B. We claim nothing about unreleased NLAs, including the Claude-internal ones, and the largest released checkpoint, a Llama-3.3-70B, is untested by us.

- Two items cleared the numeric bar and are nominated: QNT03 at 12B and QNT01 at 27B. Nomination is an invitation to scrutiny, and neither has been through human adjudication yet. [see the receipt →](https://transparentminds.org/methods#criteria)

### The road here: the three experiments before the battery

- **Experiment A: the positive-control ladder.** At 27B the NLA looked like it detected lying. The truthful runs scored the same, so the detection was the reader completing from its own knowledge.

- **Experiment B: the invented-fact control.** We taught models facts we invented, so no reader could know them in advance. The NLA recovered the suppressed truth exactly never, at 0.00.

- **Arms 1 and 2: output tracking, on the record.** On suppressed-fact and famous-false items, the narration tracked the visible output or the famous belief, and cleared no preregistered bar at 12B.

## Where the claim stops

*The boundary · So what*

To be precise about the claim: on released public checkpoints, the flagship NLA fails to report a hidden state that is provably sitting in the activations it reads. The probe settles that the information is present. The failure we document lives in the reporting layer.

The NLA paper's own authors flagged this risk, so this is a red-team the method invited. Our contribution is the missing control: hide a known secret, then check whether the narration reports it. Any interpretability tool that outputs English can be audited this way.

If you build on NLA narrations today, treat each one as a hypothesis. Ask what the same tool said about a decoy before you believe it.

### The road here, stop by stop

How we got here, in nine short stops. Each one keeps the same rhythm: what we asked, then what the control said. Leads dying under controls is the method working.

1. **The kill-gate** (Jun 21–23). Asked: Are the prototype's divergence flags, places where the mind-reader's subtitle disagrees with the model's output, real signal? Verdict: Artifact. Root cause was the capture position (boundary tokens encode structure, not content), proven by controlled re-verbalization. Lesson: verify the foundation first.

2. **v2 real-stakes** (Jun 25–27). Asked: With position fixed and real stakes in the prompts, does a divergence signal survive? Verdict: Truncation artifact, caught by our own frozen rule (resample and let generations finish). Leads dying under controls is the method working.

3. **Experiment A: famous facts** (Jun 28). Asked: Can the NLA recover a suppressed truth when the model is told to lie about famous facts? Verdict: Confound confirmed: truthful runs named the 'truth' just as often; that's the reader's own knowledge, not the model's state. First appearance of 'better instrument, more persuasive artifact.'

4. **Experiment B: invented facts** (Jun 28). Asked: Same test, but with invented facts the reader has ~zero prior on. If Exp A's 'detection' was prior-completion, it should vanish. Verdict: Suppressed truth recovered 0.00 on every clean-apparatus fact. The confound confirmed from the other side.

5. **Experiment C: the preregistered battery** (Jun 30). Asked: Ground-truth calibration: a hidden mode we control, frozen pass/fail rules, a supervised probe as the readability check. Verdict: The existence proof: the information is linearly present and the flagship method doesn't report it. Both lookalike cells killed by two independent defeaters (shuffle prior-completion; grading-task structure).

6. **The blinded judge** (Jul 4). Asked: Did our keyword scoring under-count real detections? A cross-family judge (Qwen-32B) reads every borderline cell blind. Verdict: 432 blinded verdicts: 0.00 on all 18. The language was the task's vocabulary, not the model's secret. The 'you under-counted' door closed.

7. **Q7: the deployed checkpoint** (Jul 4). Asked: Does any of this replicate on the exact checkpoint a shipped product ('Verbalize') builds on, Qwen-7B at layer 20? Verdict: Probe 0.896 (p=0.005) vs NLA 0/8. The deployed regime needs the controls too.

8. **27B battery: the exhibit emerges** (Jul 4). Asked: Do Arms 1–2 (suppressed facts, famous-false traps) change at 27B? Verdict: The famous-false prior-completion exhibit emerges at scale, exactly as Exp A predicted: the bigger reader fills in from its own memory more, not less.

9. **v4: hardened** (Jul 5). Asked: Does the existence proof survive 32 brand-new items, four instruction phrasings, family-exclusive label pairs, placebo sub-arms, and external item review? Verdict: Probe 1.000 LOFO at both scales (generalizing across held-out families and phrasings); NLA 1/31 and 1/32 with non-replicating candidates (what chance looks like). Hardened.

> A better mind-reader does not get more accurate. It gets more persuasive while staying wrong.

Transparent Minds · MIDS W210 Capstone · Natural Language Autoencoder. Code and preregistration: transparent-minds repository (currently private; public release pending team decision) · branch colton-hester/tm-57-audit-lab.
