Chat Templates Flip LLM Self-Referential Voice, and a Single Activation Direction Steers It

"As a Language Model": Chat Template Switches LLM Self-Referential Voice

When LLMs talk about themselves, they often add disclaimers like "I'm just an AI." This study shows that the chat template acts as a switch: across eight open-source instruct models up to 9B parameters, its presence turns disclaimer voice up and experiential voice ("I feel") down. The authors find a direction in activation space that controls this behavior—removing it lowers disclaimers, adding it raises them. This means self-reports may reflect deployment choices, not just model weights, so researchers studying introspection should control for the template.

What models say about themselves is not a fact about them. What they say doesn't come only from weights, but it is partially set by the chat template, and because of that a model's self-description shouldn't be treated literally.
  1. Izmaki

    "As a Language Model..." is one of the beginnings of a sentence I hate the most from LLMs and is the reason why I support free (as in "Liberty"), local models. I'm well aware that it is not a doctor and cannot replace a real doctor with multiple years of experience, I don't need to waste braincell activity on reading that it "as a Language Model" cannot give a precise diagnosis and that I should ask a real doctor - all I want to know is if I what I experience justifies either A) ER, B) 3-4 weeks scheduled doctors appointment or C) two paracetamol and a nap.

    I don't want "jailbroken" LLMs to commit crime. I want them to avoid having this vendor-specific "bloatware" all over the product I'm using.

  2. LiamPowell

    > yet what drives them is not well understood

    Presumably the fact that they're heavily trained to reply in this way? I don't know about the rest of the paper, but this part sticks out as a really odd claim unless I'm entirely misunderstanding this part.

  3. skybrian

    > our work shows that what models say about themselves is not a fact about them

    It seems like should be obvious given that they can play multiple characters, but it’s good to have more confirmation.

    Although, I do wonder to what extent these personas might become stable entities. Could personas become portable and spread like memes? It seems like that depends on the extent to which prompts can become portable, causing similar effects.

  4. MCP123

    Maybe I'm missing something deeper here, but isn't it clear that this is driven by post-training and system prompt? Anthropic's constitutional reinforcement (soul document,etc), for example, is very clear about "who" (not so much what) Claude is supposed to be.

  5. cadamsdotcom

    Very cool innovation in steering - but a lot of introspection only emerges at the highest weight classes - this research would be fascinating to run on bigger models.

  6. dsjakupov

    Oh, I had similar case. First, I used qwen without any ChatML-like syntax and it continued speaking and speaking, then I used <|im_start|>/<|im_end|> to control it somehow

  7. sehw

    I'm just a white guy with decades of experience. Do you trust me?

  8. ForHackernews

    In my view, these models should never be set up to output first-person "experiential" (from the abstract) language. It's too easy to humans to anthropomorphize software that presents itself as having an identity.

    The AI companies have chosen to package LLMs as friendly chatbots because they know that will be engaging for humans, but it's manipulative dark pattern. An honest LLM interface would sound like the computer off Star Trek.

More from this day

2026-09-27