Transparent Minds · Audit Lab

Should you trust an AI mind-reader?

Tools called Natural Language Autoencoders, NLAs for short, promise to translate a model's hidden internal activity into plain English. We built the safety-auditing product that promise implies. Then we tested whether the NLA mind-reader actually reads minds. A five-line statistics baseline read a hidden state from the internals perfectly. The flagship NLA reported it about as often as chance.

Headline results

Pages

For agents and crawlers

This site is a JavaScript app; the full content is mirrored as markdown and raw JSON:

Transparent Minds · MIDS W210 Capstone · Natural Language Autoencoder