Truth Is Not a Direction: Why LLM Probes Cannot Capture Truth

Truth is not a direction: a Tarski attack on LLM probes

Truth Is Not a Direction: Why LLM Probes Cannot Capture Truth

I explore why no probe on a language model's embedding space can definitively pin down truth. Drawing on Tarski's undefinability theorem, I demonstrate that any system expressive enough to describe its own truth probe creates a paradox. While simple truth probes work surprisingly well on basic cases, they cannot function as universal oracles because self-reference inevitably leads to logical contradictions.

No definable probe over a model’s representation space can exactly capture truth for any language rich enough to describe that probe and its outputs.
  1. kstenerud

    > It might seem absurd to you to even suggest superhuman AIs could function as a truth-oracle (it certainly does to me), but there are two reasons to take it seriously. First, it is how these things will be used practically by the vast majority of people. They are already replacing standard Google search results, and I’ve had many discussions end with people delegating final authority on the truth to an AI.

    There are already organizations taking advantage of this, actively producing "AI propaganda", meaning propaganda aimed at the LLMs themselves in order to influence their understanding of what is truthful and bend it towards powerful actors' agendas. They're not even hiding the fact that they're doing this.

  2. Legend2440

    I think this article pushes the premise farther than is reasonable.

    The best anyone expects from an LLM "truth vector" is that it would encode the model's belief about whether the statement is true. Of course a perfect truth oracle is impossible.

  3. baq

    Title is a bit clickbaitish, but the content is well worth reading - came in with my pitchfork ready and left agreeing with basically all of it, with questions like ‘what if the probe could return 3 dimensions: truthfulness, knowledge confidence and decidability?’

    Also the observation that people treat LLMs like oracles when they’re everything but is spot on, something I’ve also been thinking about and it’s quite a bit scary.

  4. aesthesia

    Fun, though as hinted at the end, the point of LLM "truth" probes is to measure the model's internal judgment of truthfulness. There's no reason this judgment, even if measured with 100% accuracy, couldn't be mistaken or logically inconsistent.

  5. ziofill

    A direction that is 99.99% accurate survives this argument completely. For all practical purposes one does not need totality.

More from this day

2026-07-29