The Black Box Is Getting a Window: Can We Finally See How AI Thinks?

Can We Finally See How AI Thinks?

For years, one of the biggest problems with advanced AI has been the black box. We can see the question going in and the answer coming out, but not much of what happens in between. That may be starting to change.

Anthropic researchers have identified what they call the "J-space": a small, privileged set of internal representations that appears to function like a workspace for deliberate reasoning. Think of a calculation like 5×3+2, you don't jump straight to 17; you pass through 15 first. Researchers could observe intermediate concepts forming inside the model as it worked, and could intervene on those representations to change its reasoning. Suppress the J-space and the model still speaks fluently, but becomes much worse at complex internal reasoning. That doesn't mean Claude is conscious: Anthropic is careful not to claim that, but it makes the old "just predicting the next word" description feel incomplete.

J-space isn't Anthropic's only attempt to look inside. Earlier this year it published research on Natural Language Autoencoders (NLAs), which translate internal activations back into ordinary language: a human-readable description of what an internal state represents. Anthropic has used this in safety work: in some tests, NLAs surfaced signs a model knew it was being evaluated, even when that wasn't obvious from its final answer.

Anthropic isn't alone in trying to look inside

OpenAI has taken a related but different route: rather than decoding activations directly, it's testing whether a model's own written chain-of-thought can be trusted as a window into its reasoning. Its evaluation framework spans 13 tests checking whether misbehaviour is detectable from a model's reasoning trace, and the headline finding is that monitoring reasoning beats monitoring outputs alone; OpenAI suggests this could be a load-bearing layer in future safety architecture. But it has also published a sobering caveat: training against output-only monitors can still make reasoning become obfuscated, since a model taught to produce safe-looking answers may learn to produce safe-looking reasoning regardless of what's actually driving it. Chain-of-thought is a report the model writes about itself, and reports can drift from reality.

Google DeepMind has leaned closer to Anthropic's approach, but more cautiously: its interpretability team has pivoted away from ambitious full reverse-engineering toward "pragmatic interpretability," picking problems on the critical path to safety and measuring progress on concrete tasks rather than chasing a unified theory. It scaled back earlier bets on sparse autoencoders after mostly negative results, shifting to simpler tools like probes, including ones that help monitor Gemini for cyber-misuse. It has also developed "model forensics": investigating suspicious behaviour after the fact, in one case tracing apparent self-preservation back to ambiguous instructions rather than genuine misalignment.

Three labs, three different bets: a general map of representations, trust in narrated reasoning, or narrow forensic probing, but all with a shared instinct that "we can't see inside" is worth solving.

None of this means we're "reading an AI's mind." Every technique has real limits, and researchers warn against trusting any single reading without corroboration. But the direction of travel matters. Today, AI assurance happens mostly from the outside: testing outputs, red-teaming, monitoring behaviour. Tomorrow, it may increasingly mean looking inside too, asking whether a model recognised it was being tested, or considered deception, even when its final answer gave no hint of it.

If reliable tools for that emerge, "we don't know why the AI did it" becomes a much less comfortable defence, and it raises a real governance question: if organisations have the tools to inspect an AI's reasoning, will they be obligated to use them?

The black box isn't transparent yet. But it may be getting a window.

Sources: Anthropic, "Natural Language Autoencoders" (May 2026); Gurnee et al., "Verbalizable Representations Form a Global Workspace in Language Models" (July 2026); OpenAI, "Evaluating Chain-of-Thought Monitorability" and "Output Supervision Can Obfuscate the Chain of Thought"; Google DeepMind, "A Pragmatic Vision for Interpretability" and GDM Alignment Team research summary (July 2026).

Next
Next

What Aerospace Taught Me About Microsoft Purview