Earlier quoted context omitted.
To whomever downvoted parent: please don't act against people brave enough to state that they don't know something. This is a desired quality, increasingly less present in IT work environments. People afraid of being shamed for stating knowledge gaps are not the folks you want to work with.
I feel like there's a minimum "due diligence" bar to meet though before asking, otherwise it comes across as "I'm too lazy to google the reference and connect the dots myself, but can someone just go ahead and distill a nice summary for me"
Why wordfreq will not be updated
451–460 of 542 posts
Re: Why wordfreq will not be updated
#452I created https://lowbackgroundsteel.ai/ in 2023 as a place to gather references to unpolluted datasets. I'll add wordfreq. Please submit stuff to the Tumblr.
Re: Why wordfreq will not be updated
#453I'm going to call it: The Web is dead. Thanks to "AI" I spend more time now digging through searches trying to find something useful than I did back in 2005. And the sites you do find are largely garbage. As a random example: just trying to find a particular popular set of wireless earbuds takes me at least 10 minutes, when I already know the company, the company's website, other vendors that sell the company's goods…
On the other hand, what you call "The Web" seems to be just what you can get at through search engines. There's still the old web, the thing that's mediated by relationships and reputation rather than aggregation services with billions of users. Like the link I shared above. Or this heroically moderated site we're using right now.
Re: Why wordfreq will not be updated
#454Wow there is so much vitriol both in this post and in the comments here. I understand that there are many ethical and practical problems with generative AI, but when did we stop being hopeful and start seeing the darkest side of everything? Is it just that the average HN reader is now past the age where a new technological development is an exciting opportunity and on to the age where it is a threat? Remember, the Lu…
When? For some of us, it was 1994, the eternal September. For some of us, it was when Aaron Swartz left us. For some of us, it was when Google killed Google Reader (in hindsight, the turning point of Google becoming evil). For some others, like the author of this post, it's when twitter and reddit closed their previously open APIs.
Re: Why wordfreq will not be updated
#455Re: Why wordfreq will not be updated
#456"I don't think anyone has reliable information about post-2021 language usage by humans." We've been past the tipping point when it comes to text for some time, but for video I feel we are living through the watershed moment right now. Especially smaller children don't have a good intuition on what is real and what is not. When I get asked if the person in a video is real, I still feel pretty confident to answer but…
There are a series of challenges like: https://www.nytimes.com/interactive/2024/09/09/technology/ai... https://www.nytimes.com/interactive/2024/01/19/technology/ar... These are a little bit unfair, in that we're comparing handpicked examples, but I don't think many experts will pass a test like this. Technology only moves forward (and seemingly, at an accelerating pace). What's a little shocking to me is the speed of…
There is the very real possibility that everything just stalls and plateau where we are at. You know, like our population growth, it should have gone exponentially but it did not. Actually, quite the reverse.
Re: Why wordfreq will not be updated
#457Re: Why wordfreq will not be updated
#458Earlier quoted context omitted.
1. People don't generally use the (big, whole-web-corpus-trained) general-purpose LLM base-models to generate bot slop for the web. Paying per API call to generate that kind of stuff would be far too expensive; it'd be like paying for eStamps to send spam email. Spambot developers use smaller open-source models, trained on much smaller corpuses, sized and quantized to generate text that's "just good enough" to pass m…
On point 1, that’s surprising to me. A 2,000 word blog post would be 10 cents with GPT-4o. So you put out 1,000 of them, which is a lot, for $100.
Re: Why wordfreq will not be updated
#459Re: Why wordfreq will not be updated
#460Earlier quoted context omitted.
But you can already see it with Delve. Mistral uses "delve" more than baseline, because it was trained on GPT. So it's classic positive feedback. LLM uses delve more, delve appears in training data more, LLM uses delve more... Who knows what other semantic quirks are being amplified like this. It could be something much more subtle, like cadence or sentence structure. I already notice that GPT has a "tone" and Claude…
is the use of miscible here a clue? Or just some workplace vocabulary you've adapted analogically?