Earlier quoted context omitted.
It's almost like its a systemic failure that is artificially created so that people wont think critically... hmmm
> is artificially created You imply that thousands of year ago everybody was thinking critically? Thinking critically is hard, stressful and might take some joy from your life.
Why wordfreq will not be updated
381–390 of 542 posts
Re: Why wordfreq will not be updated
#382Earlier quoted context omitted.
Ok, but what i said is true regardless of SEO, and that SEO has also fed back into english before LLMs were a thing. If you only train on those subsets you'll also end up with a chatbot that doesn't speak in a way we'll identify as natural english.
Yet. Give it time. The LLMs will train our future children.
Re: Why wordfreq will not be updated
#383Earlier quoted context omitted.
> I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. Blog spam was generally written by humans. While it sucked for other reasons, it seemed fine for measuring basic word frequencies in human-written tex…
Isn't it the other way around? SEO text carefully tuned to tf-idf metrics and keyword stuffed to them empirically determined threshold Google just allows should have unnatural word frequencies. LLM content should just enhance and cement the status quo word frequencies. Outliers like the word "delve" could just be sentinels, carefully placed like trap streets on a map.
So it's classic positive feedback. LLM uses delve more, delve appears in training data more, LLM uses delve more...
Who knows what other semantic quirks are being amplified like this. It could be something much more subtle, like cadence or sentence structure. I already notice that GPT has a "tone" and Claude has a "tone" and they're all sort of "GPT-like." I've read comments online that stop and make me question whether they're coming from a bot, just because their word choice and structure echoes GPT. It will sink into human writing too, since everyone is learning in high school and college that the way you write is by asking GPT for a first draft and then tweaking it (or not).
Unfortunately, I think human and machine generated text are entirely miscible. There is no "baseline" outside the machines, other than from pre-2022 text. Like pre-atomic steel.
Re: Why wordfreq will not be updated
#384I'm going to call it: The Web is dead. Thanks to "AI" I spend more time now digging through searches trying to find something useful than I did back in 2005. And the sites you do find are largely garbage. As a random example: just trying to find a particular popular set of wireless earbuds takes me at least 10 minutes, when I already know the company, the company's website, other vendors that sell the company's goods…
This is going to be the thing that makes me quit Amazon. If I'm missing something and there's still a way to to a direct search, please tell me.
Re: Why wordfreq will not be updated
#385Earlier quoted context omitted.
Isn't it the other way around? SEO text carefully tuned to tf-idf metrics and keyword stuffed to them empirically determined threshold Google just allows should have unnatural word frequencies. LLM content should just enhance and cement the status quo word frequencies. Outliers like the word "delve" could just be sentinels, carefully placed like trap streets on a map.
But you can already see it with Delve. Mistral uses "delve" more than baseline, because it was trained on GPT. So it's classic positive feedback. LLM uses delve more, delve appears in training data more, LLM uses delve more... Who knows what other semantic quirks are being amplified like this. It could be something much more subtle, like cadence or sentence structure. I already notice that GPT has a "tone" and Claude…
Some day we may view this as the beginnings of machine culture.
Re: Why wordfreq will not be updated
#386Earlier quoted context omitted.
>AI didn't just occur in 2021. Nobody knows how much text was machine generated prior to 2021 But we do know that now it's a lot more, with a big LOT.
I assume you are correct but how can we know rather than assume? I am not sure we can, so why get worked up about "internet died in 2021" when many would claim with similar conviction that it's been dead since 2012, or 2007, or ...
Re: Why wordfreq will not be updated
#387I created https://lowbackgroundsteel.ai/ in 2023 as a place to gather references to unpolluted datasets. I'll add wordfreq. Please submit stuff to the Tumblr.
:'( I thought I was clever for realising this parallel myself! Guess it's more obvious than I thought. Another example is how data on humans after 2020 or so can't be separated by sex because gender activists fought to stop recording sex in statistics on crime, medicine, etc.
Re: Why wordfreq will not be updated
#388Earlier quoted context omitted.
But you can already see it with Delve. Mistral uses "delve" more than baseline, because it was trained on GPT. So it's classic positive feedback. LLM uses delve more, delve appears in training data more, LLM uses delve more... Who knows what other semantic quirks are being amplified like this. It could be something much more subtle, like cadence or sentence structure. I already notice that GPT has a "tone" and Claude…
> LLM uses delve more, delve appears in training data more, LLM uses delve more... Some day we may view this as the beginnings of machine culture.
Have you ever seen someone use their smartphone? They're not "here," they are "there." Forming themselves in cyberspace -- or being formed, by the machine.
Re: Why wordfreq will not be updated
#389Earlier quoted context omitted.
This feels like a second, magnitudes larger Eternal September. I wonder how much more of this the Internet can take before everyone just abandons it entirely. My usage is notably lower than it was in even 2018, it's so goddamn hard to find anything worth reading anymore (which is why I spend so much damn time here, tbh).
I hope this trend accelerates to force us all into grass-touching and book-reading. The sooner, the better.
I already find myself mentally filtering out audible releases after a certain date unless they're from an author I recognize.
Re: Why wordfreq will not be updated
#390Earlier quoted context omitted.
Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.
> The training data is weighted by a quality metric At least in Googles case, they're having so much difficulty keeping AI slop out of their search results that I don't have much faith in their ability to give it an appropriately low training weight. They're not even filtering the comically low-hanging fruit like those YouTube channels which post a new "product review" every 10 minutes, with an AI generated thumbnail…