I wonder about this technique vs simple SVM classifiers: https://x.com/rosmine/status/2056406399471558872?s=20
The link you sent is for generating text which attempts to defeat those classifiers.
21–30 of 184 posts
I wonder about this technique vs simple SVM classifiers: https://x.com/rosmine/status/2056406399471558872?s=20
The link you sent is for generating text which attempts to defeat those classifiers.
Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Images, absolutely, there are…
Largely AI generated books are a vastly different situation than a one paragraph homework assignment. But multiple rounds of homework assignments would change the accuracy.
Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Images, absolutely, there are…
> but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Hard disagree. LLMs (especially base ones, that only received pre-training) can produce output that is undistinguishable from human writing (because that's what they were trained to do). But commercial chat models are specifically tuned in a way that maximizes user engagement. It's that specific tuning that is very easy…
Did you actually try them? I did.They generated even more "slopey" text than instruction-tuned ones.
The problems are simply too great if an LLM detector has any false positives at all. Imagine how soul-crushing writing an entire dissertation by hand and having it rejected because some “good enough” LLM detector decides you write too much like an AI.
I think the fundamental problem is that training current SOTA AI models is very expensive. If a simple "classical" model can detect them, presumably at much lower algorithmic cost, then why wouldn't the model trainers use these same tools to feed back into their models to improve them at low cost to make them better? It's an arms race. Any cheap pattern can and presumably will be used to retrain if it becomes and eff…
Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Images, absolutely, there are…
> but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Hard disagree. LLMs (especially base ones, that only received pre-training) can produce output that is undistinguishable from human writing (because that's what they were trained to do). But commercial chat models are specifically tuned in a way that maximizes user engagement. It's that specific tuning that is very easy…
There are two problems with this.
The first is that it would still misclassify human-authored text written under the same incentive, and most people have various incentives to "maximize engagement".
And the second is that then people would just make other models that are tuned for defeating that sort of classifier, which would be used whenever the classifier is being used.