Live data from Hacker News

Why wordfreq will not be updated

github.com

511–520 of 542 posts

Re: Why wordfreq will not be updated

#511
post #388
post #385

Earlier quoted context omitted.

> LLM uses delve more, delve appears in training data more, LLM uses delve more... Some day we may view this as the beginnings of machine culture.

Oh no, it's been here for quite a while. Our culture is already heavily glued to the machine. The way we express ourselves, the language we use, even our very self-conception originates increasingly in online spaces. Have you ever seen someone use their smartphone? They're not "here," they are "there." Forming themselves in cyberspace -- or being formed, by the machine.

chat is this real?

Re: Why wordfreq will not be updated

#512
post #502
post #494

Earlier quoted context omitted.

On a meta level I agree that having this kind of dataset with "before and after" would be pretty interesting. On an object level I do not predict that this would increase the overall diversity of language usage - and in fact it would be extremely surprising if this was even possible due to some general mathematical properties of neural networks - nor would "more professional writing," though I do agree with this char…

Haha I'm pleasantly surprised to see my comment at the top, I genuinely thought it would drown to the bottom! Not due to disagreement, just due to sheer volume and being posted rather late in this posts lifespan. Anyways my meta comment wasn't that I disagreed with all the other comments, I was just frustrated at how repetitive they were of one another. When I go to leave a comment, I do a pass reading through all or…

The most ridiculous aspects for me were the anthropomorphizing (Reminds me of that one Sam Altman interview a bit) and the use of "infinite", which both doesn't really work on vibes (as many have noted, while I'm sure chatGPT has been exposed to every word, its pattern of communication is very "regression to the mean" among them), but also is silly if taken literally, because unless we're counting like some quirky technically-grammatical combinatoric compounding that we in practice infer the meaning of from composition of what we identify as separate individual words (like just hyphenating a bunch of adjectives and a noun or something) there's not really an argument for there being "infinite vocabulary" in the same sense that there is for "infinite possible sentences" because being a valid word requires at least that someone can meaningfully comprehend what is meant by it, and coordination requirements of this nature tend to truncate infinities

The case for ChatGPT doing significant coinage that sticks isn't particularly strong either, partially from theory and partially because I'd think I'd've heard a lot of complaints about it by now, and the ones on hackernews would be repetitive to the point of seeming unavoidable (we agree on that for sure)

Anyway, re: the silliest hype I've heard all week, I'm mostly just trying to find humor in what has been a pretty bad hype wave for someone who's pathologically bad at sounding like the kind of nontechnical hype guys that pervade any tech hype wave but is nonetheless mostly seeking out jobs in this field because it's what my expertise is in. Incredibly awful job market for a lot of people I realize, but it feels like a special hell I get for getting into ML research before it was (quite so) cool. I'm trying to fight the negativity but I've gotten screwed over a lot lately, but I don't have anything against you personally for being silly on hn

Re: Why wordfreq will not be updated

#513

Earlier quoted context omitted.

> I hear this complaint often but in reality I have encountered fairly little content in my day to day that has felt fully AI generated? How confident are you in this assessment? > straight drivel We're past the point where what AI generates is "straight drivel"; every minute, it's harder to distinguish AI output from actual output unless you're approaching expertise in the subject being written about. > a team of co…

> every minute, it's harder to distinguish AI output from actual output unless you're approaching expertise in the subject being written about. So, then what really is the problem with just including LLM-generated text in wordfreq? If quirky word distributions will remain a "problem", then I'd bet that human distributions for those words will follow shortly after (people are very quick to change their speech based on…

> So, then what really is the problem with just including LLM-generated text in wordfreq?

> Why not just own the fact that LLMs are going to be affecting our speech?

The problem is that we cannot tell what's a result of LLMs affecting our speech, and what's just the output of LLMs.

If LLMs result in a 10% increase of the word "gimple" online, which then results in a 1% increase of humans using the word "gimple" online, how do we measure that? Simply continuing to use the web to update wordfreq would show a 10% increase, which is incorrect.

Re: Why wordfreq will not be updated

#514

Earlier quoted context omitted.

The earth will recover. We may not, but earth will.

And in a few million years, the next intelligent life form will examine remains of human texts, and wonder: with all the tools and knowledge they possessed, how could they not have prevented their demise? (Sorry for pessimism and offtopicism)

We are but puny agents of entropy.

Re: Why wordfreq will not be updated

#515
post #397

Earlier quoted context omitted.

Would you mind dropping the link talking about this point? (context: I'm a total outsider and have no idea what TFA is.)

TFA means "the featured article", so in this case the "Why wordfreq will not be updated" link we're talking about.

The Fucking Article, from RTFA - Read the Fucking Article - and RTFM - Read the Fucking Manual/Manpage

Re: Why wordfreq will not be updated

#516

I'm going to call it: The Web is dead. Thanks to "AI" I spend more time now digging through searches trying to find something useful than I did back in 2005. And the sites you do find are largely garbage. As a random example: just trying to find a particular popular set of wireless earbuds takes me at least 10 minutes, when I already know the company, the company's website, other vendors that sell the company's goods…

On Amazon, you used to be able to search the reviews and Q&A section via a search box. This was immensely useful. Now, that search box first routes your search to an LLM, which makes you wait 10-15 seconds while it searches for you. Then it presents its unhelpful summary, saying "some reviews said such and such", and I can finally click the button to show me the actual reviews and questions with the term I searched.…

Ran into this the other day. Amazon.ca still has the old version for now

Re: Why wordfreq will not be updated

#517

Earlier quoted context omitted.

Also it breaks the languagr barreer, you can now read the Chinese internet if you want, or chat transparently in Arabic. That's going to be interesting.

At the moment though (and ever since decent online translation services were a thing), it feels one-way, that is, people from that side of the internet coming to the anglosphere internet moreso than anglosphere people going internet-abroad. I may be wrong.

As a Frenchman, I learned very quickly that my language sphere market and resource pool is so much smaller than the English one that it's 10 times less effective to do anything in it.

I understand the position.

The only exception would be China, but the GFW is probably not helping.

LLV might lower the cost of that so much that it will become more interesting to do so.

Re: Why wordfreq will not be updated

#518

Earlier quoted context omitted.

There are smaller, gated communities that are still very valuable. You're posting in one. But yes, the open Internet is basically useless now, thanks ultimately to advertising as a business model.

this is not a gated community at all

True, that is maybe too strong a phrase, but I think it's close to accurate. I think the culture & medium provide kind of a self-selecting gate: it's just plain text and links to articles, with the discussion expected by culture to be fairly serious. I think that turns off enough people that it kind of forms its own gate shutting out the people that make "eternal Septembers" happen. But yeah, ultimately, you're right.

Re: Why wordfreq will not be updated

#519
post #396
post #322

Earlier quoted context omitted.

No disagreement for the most part. I used to be able to say search for Trek bike derailleur hanger and the first result would be what I wanted. Now I have to scroll past 5 ads to buy a new bike, one that's a broken link to a third party, and if I'm really lucky, at the bottom of page 1 will be the link to that part's page. The shitification of the web is real.

R.I.P. Sheldon Brown T_T (The Agner Fog of cycling?)

He was a legend.

Re: Why wordfreq will not be updated

#520
post #486

Earlier quoted context omitted.

Think of an LLM as a person on the internet. Just like everyone else, they have their own vocabulary and preferred way of talking which means they’ll use some words more than others. Now imagine we duplicate this hypothetical person an incredible amount of times and have their clones chatter on the internet frequently. ‘Certainly’ this would have an effect.

Yes but this person learned to mimic the internet at large. Theoretically its preferred way of talking would be the average of all training data, as mimicry is GPT's training objective, and would therefore have very similar word distributions. Only, this doesn't account for RLHF and prompts spreading memetically among users.

> Theoretically its preferred way of talking is would be the average of all the training data

This is incorrect. Furthermore, what the LLM says is also determined by what its user wants it to say, and how frequently the user wants the LLM to post on the internet. This will have a large effect on the internet’s word frequency distribution.

Post reply on HN