Why Pangram's AI Detection Is Brittle, Not Broken

I wouldn't say Pangram is broken, but I would say that it's brittle

Why Pangram's AI Detection Is Brittle, Not Broken

After being accused of using AI, I tested Pangram and found it produces contradictory results depending on context. A section flagged as 100% AI becomes human when analyzed alone, yet the full essay is deemed human. The tool's percentage meter seems broken, often forcing binary outcomes instead of reflecting the actual mix of human and LLM text. This unreliability makes it dangerous for professional consequences.

Pieces that Pangram call 100% human written are part of a larger piece that Pangram calls 100% AI written which in turn is part of an essay that Pangram calls 100% human written, a Russian doll of contradictory results.
  1. timpera

    I recently tried Pangram, created an account, wrote a few lines about how my day went, and it was flagged as likely AI-assisted. It clearly doesn't work.

    It's especially bad that they keep insisting that it works very well, because thousands of people will probably end up falsely accused of AI usage as a result.

  2. btown

    It frustrates me that AI detection for student essays effectively works as a protection racket: if a student wants to write an essay without AI assistance and reliably get credit for doing so, they still need to pay a company like Pangram for an individual account to ensure their own work isn't accidentally flagged (especially if checking multiple drafts a day).

    And even this isn't perfect, nor is it guaranteed that enterprise and individual accounts are tuned the same way. So students also need to proactively use audit/keystroke logging systems to protect themselves against accusations, which creates a type of "panopticon" on one's early/ephemeral drafts, including language of frustration (who among us hasn't typed curses into an unsaved draft at some point?), that can massively stifle creative thought. And if an institution provides such a tool, their centralized access simply worsens the "panopticon" characteristics.

    There's no easy solution, here, sadly.

  3. runako

    This product does not even have a plausible theory of how it could work.

    LLM-generated text does not carry a watermark or other identifying marks. The "theory" is that an LLM trained on human writing, to mimic human writing, can be distinguished from actual human writing in under 100 words.

    Notably the first diagram on the research overview page (https://www.pangram.com/research/how-it-works) shows feedback for "misclassified human examples." This is a category error; Pangram will not find out when it has misclassified text in the wild, except in rare cases. Only the "licensed human-written text" in its training data can be used as feedback.

    Scams like Pangram also cause real harms, mostly because laypeople do not understand that what is being offered is not possible. Pangram advertises 99.98% accuracy, and they pitch it as a tool for teachers and universities. Translated: if a college like University of Alabama rolled this out, you could expect ~40 students to have their lives upended by this snake oil, every year. (And how can one even prove that an allegation is false, that they did write a given text?) And this is the best case, using the number on Pangram's homepage.

  4. Catloafdev

    I'm not sure I'd use the term 'brittle' to describe snake oil.

  5. achileas

    Can something be broken that never actually worked?

More from this day

2026-07-26