Live data from Hacker News

Why wordfreq will not be updated

github.com

481–490 of 542 posts

Re: Why wordfreq will not be updated

#481
This has to be the most annoying hacker news comment section I've ever seen. It's just the same ~4 viewpoints rehashed again, and again, and again. Why don't folks just upvote other comments that say the same thing instead of repeating the same things?

And now a hopefully new comment: having a word frequency measure of the internet as we're going into AI being more used would be IMMENSELY useful specifically _because_ more of the internet is being AI generated! I could see such a dataset being immensely useful to researchers who are looking for the impacts of AI on language, and to test empirically a lot of claims the author has made in this very post! What a shame that they stopped measuring.

Also: as to the claims that AI will cause stagnation and a reduction of the variance of English vocabulary used, this is a trend in English that's been happening for over 100 years ( https://evoandproud.blogspot.com/2019/09/why-is-vocabulary-s... ). I believe the opposite will happen, AI will increase the average persons vocabulary, since chat AIs tends to be more professionally written than a lot of the internet. It's like being able to chat with someone that has an infinite vocabulary. It also makes it possible for people to read complicated documents well out of their domain, since they can ask not just for definitions but more in depth explanations of what words/sections mean.

Here's to a comment that will never be read because of all the noise in this thread :/

Re: Why wordfreq will not be updated

#482
post #464
post #458

Earlier quoted context omitted.

But then you'll be competing for clicks with others who put out 1,000,000 posts for less costs because they used a small, self hosted model.

if you are a sales & marketing intern, have a potato laptop and $100 budget to spend on seo, you aren't going to be self hosting anything even if you know what that means.

This is about high-volume blog/news-spam created specifically to serve ads and affiliate links, not about occasional content marketing for legitimate companies.

Re: Why wordfreq will not be updated

#483

Ok so post author is AI skeptic and this is his retaliation, likely because his work is affected. I believe governments should address the problem with welfare but being against technical advances is always being in the wrong side of history.

This is a tech site, where >50% of us are programmers who have achieved greater productivity thanks to LLM advances. And yet we're filled to the gills with Luddite sentiments and AI content fearmongering. Imagine the hysteria and the skull-vibrating noise of the non-HN rabble when they come to understand where all of this is going. They're going to do their darndest to stop us from achieving post-economy.

I think programmers are in the perfect profession to call LLMs out for just how bad they are. They are fancy auto-complete and I love them in my daily usage, but a big part of that is because I can tell when they are ridiculously wrong. Which is so often you really have to question how useful they would be for anything where they aren’t just fancy auto-complete.

Which isn’t AIs fault. I’m sure they can be great in cancer detection, unless they replace what we’re already doing because they are cheaper than doctors. In combination with an expert AI is great, but that’s not what’s happening is it?

Re: Why wordfreq will not be updated

#484
post #479
post #466

I agree with the general ethos of the piece (albeit a few of the details are puzzling and unnecessarily partisan - content on X isn't invariably worthless drivel, nor does what Reddit is doing make much intellectual as opposed to economic [IPO-influenced] sense - but this line: 'OpenAI and Google can collect their own damn data. I hope they have to pay a very high price for it, and I hope they're constantly cursing t…

> puzzling and unnecessarily partisan - content on X isn't invariably worthless drivel Maybe this is because I’m European, but what is partisan about calling X invariably worthless drivel? Seems a lot like facts to me considering what has been going on with the platform moderation since Elon Musk bought it. It’s so bad that the EU consider it a platform for misinformation these days.

Do you have a citation on that last claim?

Re: Why wordfreq will not be updated

#485

Earlier quoted context omitted.

It's high quality when the content is within HN's bubble. Anything related to health, politics, or Microsoft is full of misinformation, ignorance, and garbage like any other site. The Microsoft discussions in particular are extremely low quality.

When economics has come up I've been curious and asked my brother about some of the stuff in the more upvoted comments (he has his PhD in economics with a focus on labor specifically) his reaction has always been something like "that doesn't match my understanding of that" or "I think their analysis is a bit oversimplified". My experience here is that it's pretty good for things outside of tech (at least better than…

I don't have a PhD but I do have some background in economics, and economics is consistently one of the worst areas on HN. I think it's representative of society in general. There's something about economics that makes it feel like you can just reason through it with common sense, whereas that's rarely true in reality.

Re: Why wordfreq will not be updated

#486

Intuitively I feel like word frequency would be one of the things least impacted by LLM output, no?

Think of an LLM as a person on the internet. Just like everyone else, they have their own vocabulary and preferred way of talking which means they’ll use some words more than others. Now imagine we duplicate this hypothetical person an incredible amount of times and have their clones chatter on the internet frequently. ‘Certainly’ this would have an effect.

Yes but this person learned to mimic the internet at large. Theoretically its preferred way of talking would be the average of all training data, as mimicry is GPT's training objective, and would therefore have very similar word distributions. Only, this doesn't account for RLHF and prompts spreading memetically among users.

Re: Why wordfreq will not be updated

#487
post #481

This has to be the most annoying hacker news comment section I've ever seen. It's just the same ~4 viewpoints rehashed again, and again, and again. Why don't folks just upvote other comments that say the same thing instead of repeating the same things? And now a hopefully new comment: having a word frequency measure of the internet as we're going into AI being more used would be IMMENSELY useful specifically _because…

I read, but I can't say I like it. :-D People will ELI5 everything to understand it, no hard word understand necessary, up-goer-five-style, then "de-compress" it into floral (Amorphophallus Titanum scented) GPT speak when sending responses back out.

Re: Why wordfreq will not be updated

#488

Earlier quoted context omitted.

> I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. Blog spam was generally written by humans. While it sucked for other reasons, it seemed fine for measuring basic word frequencies in human-written tex…

Isn't it the other way around? SEO text carefully tuned to tf-idf metrics and keyword stuffed to them empirically determined threshold Google just allows should have unnatural word frequencies. LLM content should just enhance and cement the status quo word frequencies. Outliers like the word "delve" could just be sentinels, carefully placed like trap streets on a map.

  Too deep we delved, and awoke the ancient delves.

Re: Why wordfreq will not be updated

#489
post #343

Earlier quoted context omitted.

> about this thing that professionals dedicate their whole lives to mastering After doing some healthcare work I ended up understanding that some topics are not well known even by the professionals dedicating their whole lives to that because there are big gaps in the human knowledge on the topics. I agree that people that think they can reason in two minutes about anything are a problem, but it's not a healthcare on…

As to the uncertainty and mysteries, you are 100% correct. One of the big failure modes for engineers in dealing with human health is the assumption that things are as simple and logical as the stuff we build, when it's simply not at all like that. There are (1) big arguments over basic things like "why do SSRI's work?" Outside of LLM's I can't think of a thing in software where we are still arguing about why things…

Regarding the human genome project specifically it was research and no matter what was claimed (give us all of these medical breakthroughs) we (as the public) should understand there is no guarantee. Similarly to how most tech startups propose plans that lead to huge scales and ROI, but nobody is amazed when 3-4 years later they have a modest revenue (the lucky ones).

The benefits for understanding more about genomes are growing (ex: list of adverse effects based on genotype https://go.drugbank.com/pharmaco/genomics) but the field is/(was) so chaotic (just one example: there was not one standard about how to count: https://tidyomics.com/blog/2018/12/09/2018-12-09-the-devil-0...) and so lacking data that it will take many years to reap the benefits (ex: one of the largest study UK Bio bank gave access to researchers only in 2017 - https://en.wikipedia.org/wiki/UK_Biobank)

Re: Why wordfreq will not be updated

#490
post #481

This has to be the most annoying hacker news comment section I've ever seen. It's just the same ~4 viewpoints rehashed again, and again, and again. Why don't folks just upvote other comments that say the same thing instead of repeating the same things? And now a hopefully new comment: having a word frequency measure of the internet as we're going into AI being more used would be IMMENSELY useful specifically _because…

You haven't read the whole thing. It says that: or that could benefit generative AI.
Post reply on HN