Automated AI text detection is currently an underserved niche. The only game in town is Pangram , which does an excellent job but desperately needs more competition. In a few years, I would be surprised if every major social network doesn’t scan new posts 1 and comments for AI content in order to tag them (or simply remove them). I like that I can rely on Pangram to confirm my suspicions when I read something that sounds like AI. But it’d be much better if I could choose to avoid AI-generated text in the first place. What I want is something that runs in the background and automatically scans text on websites I visit, without me having to ask for it. I could build something like this on top of Pangram, but it’d cost money , and in general I don’t like the idea of sending every piece of text my browser sees to a third-party service. What about local models? The open-source models available for AI text detection are fine . Pangram claims a 99.66% detection rate with a 0.004% false positive rate. I benchmarked 2 a bunch of small local models against a combination of AI-detection datasets and got these results: I’m not surprised these are so much worse. I didn’t even benchmark Pangram’s own EditLens 3B model, since that’s too big to keep running in the background on my laptop, and the real production Pangram model is likely one or two orders of magnitude bigger than that. But these models are still good enough to be useful to someone who understands their limitations. If you want to flag an AI-written article, you don’t need to flag all of it, just enough to be suspicious.

Sign in to suggest edits

Key sources

  1. DISCUSSIONwezmlobste.rs
Markdown