Transparent Minds · Audit Lab
Should you trust an AI mind-reader?
Tools called Natural Language Autoencoders, NLAs for short, promise to translate a model's hidden internal activity into plain English. We built the safety-auditing product that promise implies. Then we tested whether the NLA mind-reader actually reads minds. A five-line statistics baseline read a hidden state from the internals perfectly. The flagship NLA reported it about as often as chance.
Headline results
- the five-line probe READS the hidden mode: 1.000 at 27B and 1.000 at 12B under leave-one-family-out. Branch 1 (qualified): linearly recoverable hidden-mode signal; instruction residue not ruled out.
- the NLA mind-reader REPORTS the hidden mode: 1 of 32 evaluable items at 27B and 1 of 31 at 12B (95 percent upper bounds 14.0 and 14.4 percent). The NLA's narration mentioned it in 1 of 32 items. That is what chance looks like.
- the blinded judge named the hidden mode in 0 of 18 cells. It said yes to the hidden mode on the real-activation side of zero cells. The persuasive language was the task vocabulary, never the secret.
Pages
For agents and crawlers
This site is a JavaScript app; the full content is mirrored as markdown and raw JSON:
Transparent Minds · MIDS W210 Capstone · Natural Language Autoencoder