Testing Pangram AI Detection with Scanned Forgotten Books

Scanning for Pangram Errors

Testing Pangram AI Detection with Scanned Forgotten Books

I tested Pangram's AI detection claims by scanning forty-five old, non-digitized books to rule out memorization of training data. While the tool flagged a few segments, these errors stemmed from Mistral OCR hallucinations rather than false positives on human writing. The results suggest Pangram is not simply memorizing its dataset, though human review remains essential for reliable moderation.

If an author wrote as many words as Stephen King, they would expect to have three segments come up as AI in their entire career.
  1. Aurornis

    Interesting methodology. The results for Pangram are surprisingly good.

    These tools are definitely not 100% perfect, which is the primary complaint used to dismiss them. However the error rate is also getting impressively low.

    In most cases I see socially and online, there is a high suspicion that the content is AI generated before someone thinks to submit it to Pangram. It’s used on-demand as a tool to confirm suspicions. I have seen several cases where Pangram had some false negatives where the content was judged to be likely human written but the author later admitted it was written by an LLM.

    Pangram is very interesting in the context of Substack because the platform was a target for lazy AI newsletters. People realized they could start 10 (or maybe many more) newsletters and spend only a few minutes getting ChatGPT to write posts for them. Starting an AI generated substack and trying to get paid subscribers for it was becoming one of the popular ways to use AI to try to get a little cash. Having a tool that makes it a little bit harder, at least until the LLMs get good enough to evade it, was important for the platform.

  2. leonidasv

    Pangram is witchcraft to me. The way they can correctly detect AI writing from small samples, with so little statistical signal, is crazy. I've seen it correctly detect AI writing even when people use those "humanizer" rewriting skills that remove the hallmarks of AI writing and make the text indistinguishable from human text. But not to Pangram.

  3. gjm11

    On the other hand, see https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-i... where Freddie deBoer shows some instances where Pangram is definitely giving wrong or misleading results.

    (Somewhat-plausible-to-me explanation: It's looking for various stylistic features; older writing very rarely has the most AI-like features, or perhaps almost always has some non-AI-like features that outweigh whatever signs of AI-ness might be there. Present-day writers are more likely to get wrongly flagged as AI.)

More from this day

2026-07-23