Live data from Hacker News

Why wordfreq will not be updated

github.com

441–450 of 542 posts

Re: Why wordfreq will not be updated

#441
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

> I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. Blog spam was generally written by humans. While it sucked for other reasons, it seemed fine for measuring basic word frequencies in human-written tex…

serpent eating its own tail

GOGI.

Re: Why wordfreq will not be updated

#442
post #234

Earlier quoted context omitted.

It's high quality when the content is within HN's bubble. Anything related to health, politics, or Microsoft is full of misinformation, ignorance, and garbage like any other site. The Microsoft discussions in particular are extremely low quality.

I disagree. Even politics spurs intelligent, nuanced discussion here on HN. And to hold up discussions about MS as an example of 'extremely' low quality discussion is, ah, interesting. Do you have any recent examples of such discussions?

> spurs intelligent, nuanced discussion here on HN

relative to what? reddit?

also there's a trade off between entropy and "quality". too much "quality" and everyone gets bored and goes somewhere more entertaining

Re: Why wordfreq will not be updated

#443
post #351

I understand the frustration shared in this post but I wholeheartedly disagree with the overall sentiment that comes with it. The web isn't dead, (Gen)AI, SEO, spam and pollution didn't kill anything. The world is chaotic and net entropy (degree of disorder) of any isolated or closed system will always increase. Same goes for the web. We just have to embrace it and overcome the challenges that come with it.

T H A N K S

Re: Why wordfreq will not be updated

#444
post #265

Earlier quoted context omitted.

IMO HN actually scores quite highly in terms of health/politics and so forth content because the both mainstream and fringe ideas get both shown and pushback. A vaping discussion brought up glycerin used was safe and the same thing used in smoke machines and someone else brought up a study showing that smoke machines are an occasional safety issue. Nowhere near every discussion goes that well but stick around and you…

> IMO HN actually scores quite highly in terms of health/politics and so forth content because the both mainstream and fringe ideas get both shown and pushback. As someone with domain expertise here, I wholeheartedly disagree. HN is very bad at percolating accurate information about topics outside its wheelhouse, like clinical medicine, public health, or the natural sciences. It is also, simultaneously, extremely pro…

people don't normally talk about healthcare on here so I'm not really sure what you're referring to or what your specialty is

Re: Why wordfreq will not be updated

#445
post #331

Earlier quoted context omitted.

Isn't it the other way around? SEO text carefully tuned to tf-idf metrics and keyword stuffed to them empirically determined threshold Google just allows should have unnatural word frequencies. LLM content should just enhance and cement the status quo word frequencies. Outliers like the word "delve" could just be sentinels, carefully placed like trap streets on a map.

1. People don't generally use the (big, whole-web-corpus-trained) general-purpose LLM base-models to generate bot slop for the web. Paying per API call to generate that kind of stuff would be far too expensive; it'd be like paying for eStamps to send spam email. Spambot developers use smaller open-source models, trained on much smaller corpuses, sized and quantized to generate text that's "just good enough" to pass m…

On point 1, that’s surprising to me. A 2,000 word blog post would be 10 cents with GPT-4o. So you put out 1,000 of them, which is a lot, for $100.

Re: Why wordfreq will not be updated

#446

I'm going to call it: The Web is dead. Thanks to "AI" I spend more time now digging through searches trying to find something useful than I did back in 2005. And the sites you do find are largely garbage. As a random example: just trying to find a particular popular set of wireless earbuds takes me at least 10 minutes, when I already know the company, the company's website, other vendors that sell the company's goods…

> Their old website was plain and worked great, letting me quickly search through their products and quickly purchase them. Last night I literally struggled to add things to cart and check out; it was actually harrowing.

Hey, who cares about making services that work when we can give people a cool chatbot assistant and a 1800 number with no real-person alternative to the decision tree

Re: Why wordfreq will not be updated

#447
post #80

Earlier quoted context omitted.

Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.

> The training data is weighted by a quality metric At least in Googles case, they're having so much difficulty keeping AI slop out of their search results that I don't have much faith in their ability to give it an appropriately low training weight. They're not even filtering the comically low-hanging fruit like those YouTube channels which post a new "product review" every 10 minutes, with an AI generated thumbnail…

Reminds me of a Google search I did yesterday: “Hezbollah” yields a little info box with headings “Overview”, “History”, “Apps” and “Return policy”.

I’m guessing that the association between “pagers” and “Hezbollah” ended up creating the latter two tabs, but who knows. Maybe some AI video out there did a product review of Hezbollah.

Re: Why wordfreq will not be updated

#448

Earlier quoted context omitted.

On Amazon, you used to be able to search the reviews and Q&A section via a search box. This was immensely useful. Now, that search box first routes your search to an LLM, which makes you wait 10-15 seconds while it searches for you. Then it presents its unhelpful summary, saying "some reviews said such and such", and I can finally click the button to show me the actual reviews and questions with the term I searched.…

You can still get to product reviews directly and search them. Here's an example: Product page (copy the identifier at the end): https://www.amazon.com/Long-Thanks-Hitchhikers-Guide-Galaxy-... Review page (paste the identifier at the end): https://www.amazon.com/product-reviews/B001OF5F1E/ This seems to bypass all of the LLM stuff for now.

Pretty good! Unfortunately it does not include the Q&As, which are often just as useful as the reviews.

Re: Why wordfreq will not be updated

#449

Earlier quoted context omitted.

> I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. Blog spam was generally written by humans. While it sucked for other reasons, it seemed fine for measuring basic word frequencies in human-written tex…

serpent eating its own tail GOGI.

The Inhuman Centipede

Re: Why wordfreq will not be updated

#450
post #383

Earlier quoted context omitted.

Isn't it the other way around? SEO text carefully tuned to tf-idf metrics and keyword stuffed to them empirically determined threshold Google just allows should have unnatural word frequencies. LLM content should just enhance and cement the status quo word frequencies. Outliers like the word "delve" could just be sentinels, carefully placed like trap streets on a map.

But you can already see it with Delve. Mistral uses "delve" more than baseline, because it was trained on GPT. So it's classic positive feedback. LLM uses delve more, delve appears in training data more, LLM uses delve more... Who knows what other semantic quirks are being amplified like this. It could be something much more subtle, like cadence or sentence structure. I already notice that GPT has a "tone" and Claude…

is the use of miscible here a clue? Or just some workplace vocabulary you've adapted analogically?
Post reply on HN