Live data from Hacker News

Why wordfreq will not be updated

github.com

221–230 of 542 posts

Re: Why wordfreq will not be updated

#221
post #116

Reading through this entire thread, I suspect that somehow generative AI actually became a political issue. Polarized politics is like a vortex sucking all kinds of unrelated things in. In case that doesn't get my comment completely buried, I will go ahead and say honestly that even though "AI slop" and paywalled content is a problem, I don't think that generative AI in itself is a negative at all. And I also think t…

The simplest example that comes to mind of something frequency analysis might be useful for would be if you had simple ciphertext where you knew that the characters probably 1:1 mapped, but you didn't know anything about how.

It could also be useful for guessing whether someone might have been trying to do some kind of steganographic or additional encoding in their work, by telling you how abnormal compared to how many people write it is that someone happened to choose a very unusual construction in their work, or whether it's unlikely that two people chose the same unusual construction by coincidence or plagiarism.

You might also find statistical models interesting for things like noticing patterns in people for whom English or others are not their first language, and when they choose different constructions more often than speakers for whom it was their first language.

I'm not saying you can't use an LLM to do some or all of these, but they also have something of a scalar attached to them of how unusual the conclusion is - e.g. "I have never seen this construction of words in 50 million lines of text" versus "Yes, that's natural.", which can be useful for trying to inform how close to the noise floor the answer is, even ignoring the prospect of hallucinations.

Re: Why wordfreq will not be updated

#222
post #217

Earlier quoted context omitted.

Have "good" small webs EVER prevailed? Every content system seems to get polluted by noise once it hits mainstream usage: IRC, Usenet, reddit, Facebook, geocities, Yahoo, webrings, etc. Once-small curated selections eventually grow big enough to become victims of their own successes and taken over by spam. It's always an arms race of quality vs quantity, and eventually the curators can't keep up with the sheer volume…

> Have "good" small webs EVER prevailed? You ask on HN, one of the highest quality sites I've ever visited in any age of the Internet. IRC is still alive and well among pretty much the same audience as always. I'm not sure it's fair to compare that with the others.

Well, niche forums are kinda different when they manage to stay small and niche. Not just HN but car forums, LED forums, etc.

But if they ever include other topics, they risk becoming more mainstream and noisy. Even within adjacent fields (like the various Stacks) it gets pretty bad.

Maybe the trick is to stay within a single small sphere then and not become a general purpose discussion site? And to have a low enough volume of submissions where good moderation is still possible? (Thank you dang and HN staff)

Re: Why wordfreq will not be updated

#223

Earlier quoted context omitted.

It did feel emotive but this wasn't the main point. Data is harder to get (or more expansive) and more polluted.

Felt super emotive to me, the problems the author is outlining, a) might not be an actual problems b) just require new thinking to solve

The problems are well-known and highly-documented. You should leave the determination of (b) up to those who know and understand (a), which includes the author.

Re: Why wordfreq will not be updated

#224

It could be used to spot LLM generated text. compare the frequency of words to those used in human natural writings and you spot the computer from the human.

Hardly. You are talking about a statistical test, which will have rather large errors (since it is based on word frequencies). Not to mention word frequencies will vary depending on the type of text (essay, description, advertisement, etc).

Re: Why wordfreq will not be updated

#225
One of the examples is the increased usage of "delve" which Google Trends confirms increased in usage since 2022 (initial ChatGPT release): https://trends.google.com/trends/explore?date=all&q=delve&hl...

It seems however it started increasing most in usage just these last few months, maybe people are talking more about "delve" specifically because of the increase in usage? A usage recursion of some sorts.

Re: Why wordfreq will not be updated

#226
post #98
post #76

Earlier quoted context omitted.

And they're all flooded with low effort trash and useless. The only remaining reliable source - now that many newspapers are axing the remaining staff in favour of LLMs - is pre-2020 print cookbooks. Anything online or printed later must be assumed to be tainted, full of untested sewage and potentially dangerous suggestions.

The wife and I use the internet for recipe ideas ... but we hardly ever follow them directly anymore. We're no formally-trained chefs but we've been home cooks for over 20 years now, and so many of them are self-evidently bad, or distinctly suboptimal. The internet chef's aversion to flavor is a meme with us now; "add one-sixty-fourth of a teaspoon of garlic powder to your gallon of soup, and mix in two crystals of t…

It's not just online recipes, but cookbooks written for the Better Home & Gardens crowd. The ones who write "curry powder" (and mean the yellow McCormick stuff which is so bland as to have almost no flavour) or call for one clove of garlic in their recipe.

I joke with folks that my assumption with "one clove of garlic" is that they really mean "one head of garlic" if you want any flavour. (And if the recipe title has "garlic" in it and you are using one clove, you’re lying.)

Re: Why wordfreq will not be updated

#227

Ok so post author is AI skeptic and this is his retaliation, likely because his work is affected. I believe governments should address the problem with welfare but being against technical advances is always being in the wrong side of history.

This is a tech site, where >50% of us are programmers who have achieved greater productivity thanks to LLM advances. And yet we're filled to the gills with Luddite sentiments and AI content fearmongering. Imagine the hysteria and the skull-vibrating noise of the non-HN rabble when they come to understand where all of this is going. They're going to do their darndest to stop us from achieving post-economy.

I fail to see the difference. Actually, programming was one of the first field where LLMs shown proficiency. The helper nature of LLMs is true in all the fields so far, in the future this may change. I believe that for instance in the case or journalism the issue was already there: three euros per post written without clue by humans.

Anyway in the long run AI will kill tons of jobs. Regardless of blog posts like that. The true key is governments assistance.

Re: Why wordfreq will not be updated

#228
post #193

Wow there is so much vitriol both in this post and in the comments here. I understand that there are many ethical and practical problems with generative AI, but when did we stop being hopeful and start seeing the darkest side of everything? Is it just that the average HN reader is now past the age where a new technological development is an exciting opportunity and on to the age where it is a threat? Remember, the Lu…

Give us examples of generative AI in challenging applications (biology, medicine, physical sciences), and you'll get a lot of optimism. The text LLM stuff is the brute force application of the same class of statistical modeling. It's commercial, and boring.

Re: Why wordfreq will not be updated

#229

"I don't think anyone has reliable information about post-2021 language usage by humans." We've been past the tipping point when it comes to text for some time, but for video I feel we are living through the watershed moment right now. Especially smaller children don't have a good intuition on what is real and what is not. When I get asked if the person in a video is real, I still feel pretty confident to answer but…

There are a series of challenges like: https://www.nytimes.com/interactive/2024/09/09/technology/ai... https://www.nytimes.com/interactive/2024/01/19/technology/ar... These are a little bit unfair, in that we're comparing handpicked examples, but I don't think many experts will pass a test like this. Technology only moves forward (and seemingly, at an accelerating pace). What's a little shocking to me is the speed of…

> One revolution I'm still coming to grips with is automated manufacturing. Going on aliexpress, so much stuff is basically free. I bought a 5-port 120W (total) charger for less than 2 minutes of my time. It literally took less time to find it than to earn the money to buy it.

Is there a big recent qualitative change here? Or is this a continuation of manufacturing trends (also shocking, not trying to minimize it all, just curious if there’s some new manufacturing tech I wasn’t aware of).

For some reason, your comment got me thinking of a fully automated system, like: you go to a website, pick and choose charger capabilities (ports, does it have a battery, that sort of stuff). Then an automated factor makes you a bespoke device (software picks an appropriate shell, regulators, etc). I bet we’ll see it in our lifetimes at least.

Re: Why wordfreq will not be updated

#230

> the Web at large is full of slop generated by large language models, written by no one to communicate nothing That’s neither fair nor accurate. That slop is ultimately generated by the humans who run those models; they are attempting (perhaps poorly) to communicate something . > two companies that I already despise Life’s too short to go through it hating others. > it's very likely because they are creating a plagi…

> It is not at all clear that a machine learning from text should be treated any differently from a human being learning from text

Given that LLMs and human creativity work on fundamentally different principles, there is every reason to believe there is a difference.

Post reply on HN