Earlier quoted context omitted.
> LLM uses delve more, delve appears in training data more, LLM uses delve more... Some day we may view this as the beginnings of machine culture.
Oh no, it's been here for quite a while. Our culture is already heavily glued to the machine. The way we express ourselves, the language we use, even our very self-conception originates increasingly in online spaces. Have you ever seen someone use their smartphone? They're not "here," they are "there." Forming themselves in cyberspace -- or being formed, by the machine.
Why wordfreq will not be updated
511–520 of 542 posts
Re: Why wordfreq will not be updated
#512Earlier quoted context omitted.
On a meta level I agree that having this kind of dataset with "before and after" would be pretty interesting. On an object level I do not predict that this would increase the overall diversity of language usage - and in fact it would be extremely surprising if this was even possible due to some general mathematical properties of neural networks - nor would "more professional writing," though I do agree with this char…
Haha I'm pleasantly surprised to see my comment at the top, I genuinely thought it would drown to the bottom! Not due to disagreement, just due to sheer volume and being posted rather late in this posts lifespan. Anyways my meta comment wasn't that I disagreed with all the other comments, I was just frustrated at how repetitive they were of one another. When I go to leave a comment, I do a pass reading through all or…
The case for ChatGPT doing significant coinage that sticks isn't particularly strong either, partially from theory and partially because I'd think I'd've heard a lot of complaints about it by now, and the ones on hackernews would be repetitive to the point of seeming unavoidable (we agree on that for sure)
Anyway, re: the silliest hype I've heard all week, I'm mostly just trying to find humor in what has been a pretty bad hype wave for someone who's pathologically bad at sounding like the kind of nontechnical hype guys that pervade any tech hype wave but is nonetheless mostly seeking out jobs in this field because it's what my expertise is in. Incredibly awful job market for a lot of people I realize, but it feels like a special hell I get for getting into ML research before it was (quite so) cool. I'm trying to fight the negativity but I've gotten screwed over a lot lately, but I don't have anything against you personally for being silly on hn
Re: Why wordfreq will not be updated
#513Earlier quoted context omitted.
> I hear this complaint often but in reality I have encountered fairly little content in my day to day that has felt fully AI generated? How confident are you in this assessment? > straight drivel We're past the point where what AI generates is "straight drivel"; every minute, it's harder to distinguish AI output from actual output unless you're approaching expertise in the subject being written about. > a team of co…
> every minute, it's harder to distinguish AI output from actual output unless you're approaching expertise in the subject being written about. So, then what really is the problem with just including LLM-generated text in wordfreq? If quirky word distributions will remain a "problem", then I'd bet that human distributions for those words will follow shortly after (people are very quick to change their speech based on…
> Why not just own the fact that LLMs are going to be affecting our speech?
The problem is that we cannot tell what's a result of LLMs affecting our speech, and what's just the output of LLMs.
If LLMs result in a 10% increase of the word "gimple" online, which then results in a 1% increase of humans using the word "gimple" online, how do we measure that? Simply continuing to use the web to update wordfreq would show a 10% increase, which is incorrect.
Re: Why wordfreq will not be updated
#514Earlier quoted context omitted.
The earth will recover. We may not, but earth will.
And in a few million years, the next intelligent life form will examine remains of human texts, and wonder: with all the tools and knowledge they possessed, how could they not have prevented their demise? (Sorry for pessimism and offtopicism)
Re: Why wordfreq will not be updated
#515Earlier quoted context omitted.
Would you mind dropping the link talking about this point? (context: I'm a total outsider and have no idea what TFA is.)
TFA means "the featured article", so in this case the "Why wordfreq will not be updated" link we're talking about.
Re: Why wordfreq will not be updated
#516I'm going to call it: The Web is dead. Thanks to "AI" I spend more time now digging through searches trying to find something useful than I did back in 2005. And the sites you do find are largely garbage. As a random example: just trying to find a particular popular set of wireless earbuds takes me at least 10 minutes, when I already know the company, the company's website, other vendors that sell the company's goods…
On Amazon, you used to be able to search the reviews and Q&A section via a search box. This was immensely useful. Now, that search box first routes your search to an LLM, which makes you wait 10-15 seconds while it searches for you. Then it presents its unhelpful summary, saying "some reviews said such and such", and I can finally click the button to show me the actual reviews and questions with the term I searched.…
Re: Why wordfreq will not be updated
#517Earlier quoted context omitted.
Also it breaks the languagr barreer, you can now read the Chinese internet if you want, or chat transparently in Arabic. That's going to be interesting.
At the moment though (and ever since decent online translation services were a thing), it feels one-way, that is, people from that side of the internet coming to the anglosphere internet moreso than anglosphere people going internet-abroad. I may be wrong.
I understand the position.
The only exception would be China, but the GFW is probably not helping.
LLV might lower the cost of that so much that it will become more interesting to do so.
Re: Why wordfreq will not be updated
#518Earlier quoted context omitted.
There are smaller, gated communities that are still very valuable. You're posting in one. But yes, the open Internet is basically useless now, thanks ultimately to advertising as a business model.
this is not a gated community at all
Re: Why wordfreq will not be updated
#519Earlier quoted context omitted.
No disagreement for the most part. I used to be able to say search for Trek bike derailleur hanger and the first result would be what I wanted. Now I have to scroll past 5 ads to buy a new bike, one that's a broken link to a third party, and if I'm really lucky, at the bottom of page 1 will be the link to that part's page. The shitification of the web is real.
R.I.P. Sheldon Brown T_T (The Agner Fog of cycling?)
Re: Why wordfreq will not be updated
#520Earlier quoted context omitted.
Think of an LLM as a person on the internet. Just like everyone else, they have their own vocabulary and preferred way of talking which means they’ll use some words more than others. Now imagine we duplicate this hypothetical person an incredible amount of times and have their clones chatter on the internet frequently. ‘Certainly’ this would have an effect.
Yes but this person learned to mimic the internet at large. Theoretically its preferred way of talking would be the average of all training data, as mimicry is GPT's training objective, and would therefore have very similar word distributions. Only, this doesn't account for RLHF and prompts spreading memetically among users.
This is incorrect. Furthermore, what the LLM says is also determined by what its user wants it to say, and how frequently the user wants the LLM to post on the internet. This will have a large effect on the internet’s word frequency distribution.