Live data from Hacker News

As AI eats the web, the internet’s collective memory is disappearing

thewalrus.ca

241–250 of 1001 posts

Re: As AI eats the web, the internet’s collective memory is disappearing

#241

Earlier quoted context omitted.

On a whim I watched my favorite childhood movie last week, Hackers. It's goofy in some ways, but man it captures the "wild west" feeling of early and mid 90's internet. It was just you, a slow connection to anywhere, and open ports all over the place. Right after I watched it, I dusted off an old hub, connected a few external usb-to-ethernet adapters to my work PC VM's, and now run them through an OpenBSD packet filt…

That movie is ridiculous and still great. At least as a nostalgia pump. And the red box references and pots patching and such were true enough, even if everything else got hilarious hollywood treatment.

If I recall correctly, the film actually had Emmanuel Goldstein of 2600 and Kevin Mitnick as consultants, so despite the hollywood treatment you have moments of accuracy like https://www.youtube.com/watch?v=4U9MI0u2VIE. Probably the only time Compilers: Principles, Techniques and Tools made it to the big screen. At the yearly 2600 conference, they used to always do a big group rollerblade through NYC.

Re: As AI eats the web, the internet’s collective memory is disappearing

#242
post #154

Earlier quoted context omitted.

Sure - it sucks, unfortunately the alternative is the sites going away entirely. When the load from scraper bots is constantly knocking the site offline the choices are literally to allow it to remain inaccessible for much of the time, put up a layer of defenses with all the user-annoyance compromises that entails, or just give up and unpublish the site.

The alternative is simple.. Go dark. VPN tech is known from like 30 years. Pretty much everyone can use it (VPN providers). But instead using it to browse net, build VPN overlay networks of interest for people. Gaming networks, R&D networks, Retro Networks. People will peer to PoP and use resources. Bad actor? BAN it from network. You have control. This could be done in Internet, but big corpos and big money won the…

Continuing on your suggestion.

There could be open source tooling to create custom private "closednets", with

- trust ring mechanism to allow invitations, flagging, banning, and banning those that invite people who were banned

- the rules of the closednet

- search engine with opt-in scraping

- portal (remember the 80s?) with all the registered nodes, perhaps by service category such as public git repo hosts, web sites etc.

etc.

The first closednet could be Hacker News.

Re: As AI eats the web, the internet’s collective memory is disappearing

#243
post #225

Earlier quoted context omitted.

That's exactly why my previous public GitHub repo is now private.

So only GitHub Copilot can read it then? Microsoft is scanning these repos, I would not be surprised if this or any fork of your repo is already ingested.

They say they don't do that. But maybe I should be more skeptical.

Re: As AI eats the web, the internet’s collective memory is disappearing

#244
post #161
post #153

Funny, I was just thinking this morning that Google searches are absolutely horrible these days. It's like it has amnesia, a lot of recent history seems to be just gone. Especially on non US specific sites too.

They're so horrible that I've started defaulting to their AI summaries. And I hate those summaries. It's just that the regular results are so terrible now, and seemingly getting worse at a noticeable pace. I used to not worry. I was sure that a competitor would come along and fix search. But the longer that's not happening, the more nervous I'm getting that we'll actually lose search. If a few more years pass in the…

How is Kagi indirectly Russian?

Re: As AI eats the web, the internet’s collective memory is disappearing

#245
post #196

AI will kill the internet because it is killing the incentive to make it. It is an industrial-strength example of why we don’t allow stealing.

Eh, part of what made early-internet so good was exactly that it did allow copying by users; the DRM era was later. But it's a very good example of why not to allow for profit copying, because that absolutely will crowd out the original. Piracy has to exist at the margin. The zero piracy world would also eat its memories because none would leak into archives. Remember Qubi? It wasn't even popular enough for people to pirate.

Re: As AI eats the web, the internet’s collective memory is disappearing

#246
post #166

Earlier quoted context omitted.

Brave search works really well, I haven't switched to Google search for months.

If I recall correctly Brave scrapes the web via their users, cannot be individually disallowed in robots.txt and Brandon Eich is conversing in a pretty hostile manner in every thread about him or his company.

I'm not in the brave ecosystem otherwise nor a big fan. But the search engine was competitive with google when it was still cliqz, before it was shut down there and the leftovers bought by brave. And it still works really well.

Even if brave were problematic it would be the lesser evil to me.

Re: As AI eats the web, the internet’s collective memory is disappearing

#247
post #232

Earlier quoted context omitted.

I've run into major problems with LLMs as I maintain my home Linux systems. If I just copied commands they list, I would have a near 100% failure rate, as most of the information they have has been gleaned from forum posts that are years out of date. They're a good jumping off point, but I need to delve into the original sources just like I did when I used Google.

I'm not sure. I run a homelab tailscale/k3s setup with gitops, dozens of services, VMs, backups, etc, all vibed by claude, and works just fine. Didn't write a single line of code for this. I don't know kubernetes and never will.

> I don't know kubernetes and never will.

He said, proud of his own ignorance.

Re: As AI eats the web, the internet’s collective memory is disappearing

#248
post #153

Funny, I was just thinking this morning that Google searches are absolutely horrible these days. It's like it has amnesia, a lot of recent history seems to be just gone. Especially on non US specific sites too.

That's 100% correct. Since Google's "helpful content update" Webmasters get massive amounts of "Crawled, not indexed" reports for anything Google considers "more of the same" or "thin content". If you're not an authority on a subject, simply meaning: you already rank for similar content, or if you don't get links from more popular domains, your content is in the abyss. It's all under the guise of "We're fighting SPAM…

> We're watching the end of Google's hegemony for sure.

Who is standing by to replace them though? OpenAI and Anthropic certainly not, they are burning money in a fire pit to stay alive. There is no way in hell they can afford the compute necessary to replace Google.

Re: As AI eats the web, the internet’s collective memory is disappearing

#249
post #173

Earlier quoted context omitted.

The Internet has been shrinking massively. My earliest experiences with the Internet were discovering the world of hobby OS dev around the turn of the millennium, when I chanced upon someone’s personal website talking about their OS, with source code and screenshots and dedicated forum. My mind was blown. I spent two years finding hundreds of small websites dedicated to the topic, hung out on IRC communities with oth…

Well, just as the "internet" supplanted newspapers, magazines (gosh those classic gaming mags), and broadcast television for many people, and how the newspapers replaced the town criers before them, why shouldn't the "internet" be supplanted by a more accessible medium? Why should I have to suffer through Fandom raping me with screen-obscuring banners and "PLEASE ALLOW ADS" just to make some sense of fucking Warhamme…

This is better for you as an individual, but worse for society as a whole. As you've noticed, monetization is the weak spot. AI will not escape being ruined by monetization, but it might be harder to notice when it arrives.

Re: As AI eats the web, the internet’s collective memory is disappearing

#250
post #100

Earlier quoted context omitted.

The danger that's concerning people (rightly or wrongly) isn't that LLMs are going to be an intermediary to your website. It's that they'll be the only thing reading it. No one will ever read your post or know what you wrote. The only consumers will be LLMs, they'll train on a version that strips out you as the author (probably more due to expedience than any sort of malice; it's not like you're famous , are you?), a…

Yes, that would be fine. I write primarily communicate ideas, not for credit or fame. Empirically, however, LLMs don't strip out the author: the big models know a lot about what I've written even with search disabled. Ex: https://claude.ai/share/8cbcdf88-a360-421a-8c06-ae7b7992e866

Sure they know about you if you ask, but generally they won't credit you if they cite an idea from their latent space that came from you.
Post reply on HN