Live data from Hacker News

Why wordfreq will not be updated

github.com

291–300 of 542 posts

Re: Why wordfreq will not be updated

#291
post #238

Earlier quoted context omitted.

Have "good" small webs EVER prevailed? Every content system seems to get polluted by noise once it hits mainstream usage: IRC, Usenet, reddit, Facebook, geocities, Yahoo, webrings, etc. Once-small curated selections eventually grow big enough to become victims of their own successes and taken over by spam. It's always an arms race of quality vs quantity, and eventually the curators can't keep up with the sheer volume…

its so easy to solve this problem, not sure why anyone hasnt done it yet. 1. build a userbase, free product 2. once userbase get big enough, any new account requires a monthly fee, maybe $1 3. keep raising the fee higher and higher, until you get to the point that the userbase is manageable. no ads, simple.

Until N ad views are worth more than $X account creation fee. Then the spammers will just sell ad posts for $X*1.5.

I can’t find it, but there’s someone selling sock puppet posts on HN even.

Re: Why wordfreq will not be updated

#292
post #265

Earlier quoted context omitted.

It's high quality when the content is within HN's bubble. Anything related to health, politics, or Microsoft is full of misinformation, ignorance, and garbage like any other site. The Microsoft discussions in particular are extremely low quality.

IMO HN actually scores quite highly in terms of health/politics and so forth content because the both mainstream and fringe ideas get both shown and pushback. A vaping discussion brought up glycerin used was safe and the same thing used in smoke machines and someone else brought up a study showing that smoke machines are an occasional safety issue. Nowhere near every discussion goes that well but stick around and you…

As a software engineer married to a healthcare professional, I disagree strongly about the quality of the healthcare discussions here. A whole lot of the conversation is software engineers who think that they can reason from first principles in two minutes about this thing that professionals dedicate their whole lives to mastering, and who therefore don't understand the most basic concepts of the field.

Sometimes I try and engage, but honestly, mostly I think it's not worth it. Otherwise you end up doing this with your life: https://xkcd.com/386/

Re: Why wordfreq will not be updated

#293
post #85

Earlier quoted context omitted.

I never thought about that. Humans losing their ability to detect AI content from reality ? It's frightening.

It's worse because many humans don't know they are. I see a lot of outrage around fake posts already. People want to believe bad things from the other tribes. And we are going to feed them with it, endlessly.

Did you think the same thing when photoshop came out?

It's relatively trivial to photoshop misinformation in a really powerful and undetectable way- but I don't see (legitimate) instances of groundbreaking news over a fake photo of the president or a CEO etc doing something nefarious. Why is AI different just because it's audio/video?

Re: Why wordfreq will not be updated

#294

I feel so conflicted about this. On the one hand, I completely agree with Robyn Speer. The open web is dead, and the web is in a really sad state. The other day I decided to publish my personal blog on gopher. Just cause, there's a lot less crap on gopher (and no, gopher is not the answer). But... A couple of weeks ago, I had to send a video file to my wife's grandfather, who is 97, lives in another country, and does…

> Back in pre-LLM days, it's not like I would have hired a x264 expert to do this job for me. I would have either had to spend hours more on this task, or more likely, this 97 year old man would never have seen his great granddaughter's dance

Didn't most DVD burning software include video transcoding as a standard feature? Back in the day, you'd have used Nero Burning ROM, or Handbrake - granted, the quality may not have been optimized to your standards, but the result would have been a watchable video (especially to 97 year-old eyes)

Re: Why wordfreq will not be updated

#295

Earlier quoted context omitted.

I fail to see the difference. Actually, programming was one of the first field where LLMs shown proficiency. The helper nature of LLMs is true in all the fields so far, in the future this may change. I believe that for instance in the case or journalism the issue was already there: three euros per post written without clue by humans. Anyway in the long run AI will kill tons of jobs. Regardless of blog posts like that…

I don't know what difference you are referring to. I was agreeing with you. And also agreed: many trumpet the merits of "unassisted" human output. However, they're suffering from ancestor veneration: human writing has always been a vast mine of worthless rock (slop) with a few gems of high-IQ analysis hidden here and there. For instance, upon the invention of the printing press, it was immediately and predominantly u…

Sorry I didn't imply we didn't agree but that programmers were and are going to be impacted as much as writers for instance, yet I see an environment where AI is generally more accepted as a tool.

About your last point sometimes I think that in the future there will be models specifically distilling the climax of selected thinkers, so that not only their production will be preserved but maybe something more that is only implicitly contained in their output.

Re: Why wordfreq will not be updated

#296
post #193

Wow there is so much vitriol both in this post and in the comments here. I understand that there are many ethical and practical problems with generative AI, but when did we stop being hopeful and start seeing the darkest side of everything? Is it just that the average HN reader is now past the age where a new technological development is an exciting opportunity and on to the age where it is a threat? Remember, the Lu…

When?

For some of us, it was 1994, the eternal September.

For some of us, it was when Aaron Swartz left us.

For some of us, it was when Google killed Google Reader (in hindsight, the turning point of Google becoming evil).

For some others, like the author of this post, it's when twitter and reddit closed their previously open APIs.

Re: Why wordfreq will not be updated

#297
post #68

I hear this complaint often but in reality I have encountered fairly little content in my day to day that has felt fully AI generated? AI assisted sure, but is that a problem if a human is in the mix, curating? I certainly have not encountered enough straight drivel where I would think it would have a significant effect on overall word statistics. I suspect there may be some over-identification of AI content happenin…

> I hear this complaint often but in reality I have encountered fairly little content in my day to day that has felt fully AI generated? How confident are you in this assessment? > straight drivel We're past the point where what AI generates is "straight drivel"; every minute, it's harder to distinguish AI output from actual output unless you're approaching expertise in the subject being written about. > a team of co…

> every minute, it's harder to distinguish AI output from actual output unless you're approaching expertise in the subject being written about.

So, then what really is the problem with just including LLM-generated text in wordfreq?

If quirky word distributions will remain a "problem", then I'd bet that human distributions for those words will follow shortly after (people are very quick to change their speech based on their environment, it's why language can change so quickly).

Why not just own the fact that LLMs are going to be affecting our speech?

Re: Why wordfreq will not be updated

#298
post #265

Earlier quoted context omitted.

IMO HN actually scores quite highly in terms of health/politics and so forth content because the both mainstream and fringe ideas get both shown and pushback. A vaping discussion brought up glycerin used was safe and the same thing used in smoke machines and someone else brought up a study showing that smoke machines are an occasional safety issue. Nowhere near every discussion goes that well but stick around and you…

> IMO HN actually scores quite highly in terms of health/politics and so forth content because the both mainstream and fringe ideas get both shown and pushback. As someone with domain expertise here, I wholeheartedly disagree. HN is very bad at percolating accurate information about topics outside its wheelhouse, like clinical medicine, public health, or the natural sciences. It is also, simultaneously, extremely pro…

Obviously on an objective scale HN isn’t good, but nobody is doing a good job here.

I’ve worked on the government side of this stuff and find it disheartening.

Re: Why wordfreq will not be updated

#299
post #203

I regret the situation led to the OP feel discourage about the NLP community, wo which I belong, and I just want to say "we're not all like that", even though it is a trend and we're close to peak hype (slightly past even?). The complaint about pollution of the Web with artificial content is timely, and it's not even the first time due to spam farms intended to game PageRank, among other nonsense. This may just mean…

Have "good" small webs EVER prevailed? Every content system seems to get polluted by noise once it hits mainstream usage: IRC, Usenet, reddit, Facebook, geocities, Yahoo, webrings, etc. Once-small curated selections eventually grow big enough to become victims of their own successes and taken over by spam. It's always an arms race of quality vs quantity, and eventually the curators can't keep up with the sheer volume…

Any curation mechanism that depends on passion and/or the goodwill of volunteers is unsustainable.

Re: Why wordfreq will not be updated

#300
This is one of the vanguards warning of the changes coming in the post-AI world.

>> Generative AI has polluted the data

Just like low-background steel marks the break in history from before and after the nuclear age, these types of data mark the distinction from before and after AI.

Future models will begin to continue to amplify certain statistical properties from their training, that amplified data will continue to pollute the public space from which future training data is drawn. Meanwhile certain low-frequency data will be selected by these models less and less and will become suppressed and possibly eliminated. We know from classic NLP techniques that low frequency words are often among the highest in information content and descriptive power.

Bitrot will continue to act as the agent of Entropy further reducing pre-AI datasets.

These feedback loops will persist, language will be ground down, neologisms will be prevented and...society, no longer with the mental tools to describe changing circumstances; new thoughts unable to be realized, will cease to advance and then regress.

Soon there will be no new low frequency ideas being removed from the data, only old low frequency ideas. Language's descriptive power is further eliminated and only the AIs seem able to produce anything that might represent the shadow of novelty. But it ends when the machines can only produce unintelligible pages of particles and articles, language is lost, civilization is lost when we no longer know what to call its downfall.

The glimmer of hope is that humanity figured out how to rise from the dreamstate of the world of animals once. Future humans will be able to climb from the ashes again. There used to be a word, the name of a bird, that encoded this ability to die and return again, but that name is already lost to the machines that will take our tongues.

Post reply on HN