Particle.news
Download on the App Store

Anthropic Finds Hidden ‘J‑Space’ That Reveals Near‑Future Tokens Inside Claude Opus 4.6

The Jacobian‑based J‑lens gives researchers a partial way to read and intervene on internal token concepts, raising new audit and safety questions for high‑risk AI uses.

Overview

  • Anthropic published a paper and demos on Thursday showing it used a new Jacobian lens, the J‑lens, to expose a representational subspace called J‑space inside Claude Opus 4.6 that encodes words the model is likely to produce soon.
  • The J‑lens links internal activations to future token probabilities by computing Jacobian vectors so researchers can map easily verbalizable concepts inside the model rather than only the next-word logits.
  • Lab examples reported by Anthropic show J‑space can hold intermediate reasoning traces such as math steps, protein and fluorescent labels, and ASCII‑face mappings that do not always appear in the model’s final output.
  • In experiments on an earlier Opus 4.5 model, researchers swapped and altered J‑space activations and broke multi‑step reasoning about 70% of the time while the model’s surface fluency and grammar stayed intact.
  • Anthropic open‑sourced the J‑lens and released a Neuronpedia demo to invite scrutiny, while warning that the tool offers partial glimpses not full guarantees and that the findings heighten regulatory and safety concerns for finance and other high‑risk deployments.