Live data from Hacker News

Why wordfreq will not be updated

github.com

451–460 of 542 posts

Re: Why wordfreq will not be updated

#451
post #208
post #184

Earlier quoted context omitted.

To whomever downvoted parent: please don't act against people brave enough to state that they don't know something. This is a desired quality, increasingly less present in IT work environments. People afraid of being shamed for stating knowledge gaps are not the folks you want to work with.

I feel like there's a minimum "due diligence" bar to meet though before asking, otherwise it comes across as "I'm too lazy to google the reference and connect the dots myself, but can someone just go ahead and distill a nice summary for me"

modern polite way of saying rtfm

Re: Why wordfreq will not be updated

#452
post #4

I created https://lowbackgroundsteel.ai/ in 2023 as a place to gather references to unpolluted datasets. I'll add wordfreq. Please submit stuff to the Tumblr.

FYI: My two datasets, DebateSum and OpenDebateEvidence/OpenCaseList in their current forms qualify for this, as they end at latest in 2022.

Re: Why wordfreq will not be updated

#453

I'm going to call it: The Web is dead. Thanks to "AI" I spend more time now digging through searches trying to find something useful than I did back in 2005. And the sites you do find are largely garbage. As a random example: just trying to find a particular popular set of wireless earbuds takes me at least 10 minutes, when I already know the company, the company's website, other vendors that sell the company's goods…

I've been slowly detaching myself from the web for the past 10 years. These days I mostly build offline apps using native technologies. Those capabilities are still around. They just receded for a while because they'd gotten so polluted with toolbars and malware. But now the malware is on the other side, and native apps are cool again. If you know where to look. Here's my shingle: https://akkartik.name/freewheeling-apps

On the other hand, what you call "The Web" seems to be just what you can get at through search engines. There's still the old web, the thing that's mediated by relationships and reputation rather than aggregation services with billions of users. Like the link I shared above. Or this heroically moderated site we're using right now.

Re: Why wordfreq will not be updated

#454
post #193

Wow there is so much vitriol both in this post and in the comments here. I understand that there are many ethical and practical problems with generative AI, but when did we stop being hopeful and start seeing the darkest side of everything? Is it just that the average HN reader is now past the age where a new technological development is an exciting opportunity and on to the age where it is a threat? Remember, the Lu…

When? For some of us, it was 1994, the eternal September. For some of us, it was when Aaron Swartz left us. For some of us, it was when Google killed Google Reader (in hindsight, the turning point of Google becoming evil). For some others, like the author of this post, it's when twitter and reddit closed their previously open APIs.

Aaron Swartz would have loved open source GenAI models.

Re: Why wordfreq will not be updated

#456

"I don't think anyone has reliable information about post-2021 language usage by humans." We've been past the tipping point when it comes to text for some time, but for video I feel we are living through the watershed moment right now. Especially smaller children don't have a good intuition on what is real and what is not. When I get asked if the person in a video is real, I still feel pretty confident to answer but…

There are a series of challenges like: https://www.nytimes.com/interactive/2024/09/09/technology/ai... https://www.nytimes.com/interactive/2024/01/19/technology/ar... These are a little bit unfair, in that we're comparing handpicked examples, but I don't think many experts will pass a test like this. Technology only moves forward (and seemingly, at an accelerating pace). What's a little shocking to me is the speed of…

Democracy (and Republics) are thousands of year old. Computation is also quite old though it only sky-rocketed with electricity and semiconductors. This is not the first time the global world created a potential for exponential growth (I'll consider the Pharaohs and Roman empires to be ones).

There is the very real possibility that everything just stalls and plateau where we are at. You know, like our population growth, it should have gone exponentially but it did not. Actually, quite the reverse.

Re: Why wordfreq will not be updated

#458
post #331

Earlier quoted context omitted.

1. People don't generally use the (big, whole-web-corpus-trained) general-purpose LLM base-models to generate bot slop for the web. Paying per API call to generate that kind of stuff would be far too expensive; it'd be like paying for eStamps to send spam email. Spambot developers use smaller open-source models, trained on much smaller corpuses, sized and quantized to generate text that's "just good enough" to pass m…

On point 1, that’s surprising to me. A 2,000 word blog post would be 10 cents with GPT-4o. So you put out 1,000 of them, which is a lot, for $100.

But then you'll be competing for clicks with others who put out 1,000,000 posts for less costs because they used a small, self hosted model.

Re: Why wordfreq will not be updated

#459

Earlier quoted context omitted.

reading stuff like this makes me so happy. no matter how fucked up something may be there is always a way to clean right up.

glances nervously at atmospheric CO2

The earth will recover. We may not, but earth will.

Re: Why wordfreq will not be updated

#460
post #383

Earlier quoted context omitted.

But you can already see it with Delve. Mistral uses "delve" more than baseline, because it was trained on GPT. So it's classic positive feedback. LLM uses delve more, delve appears in training data more, LLM uses delve more... Who knows what other semantic quirks are being amplified like this. It could be something much more subtle, like cadence or sentence structure. I already notice that GPT has a "tone" and Claude…

is the use of miscible here a clue? Or just some workplace vocabulary you've adapted analogically?

Human me just thought it was a good word for this. It implies some irreversible process of mixing, I think that characterizes this process really well.
Post reply on HN