Live data from Hacker News

As AI eats the web, the internet’s collective memory is disappearing

thewalrus.ca

81–90 of 1001 posts

Re: As AI eats the web, the internet’s collective memory is disappearing

#81

Earlier quoted context omitted.

It took you 4 days because you used Gemini. Gemini is the worst AI model I ever used. It is way behind even open models. It looks like Google just reached its AOL moment.

Come now, copilot is worse in every way

Copilot can use any model like Sol or Opus and is just a harness so not sure why people say this.

Re: As AI eats the web, the internet’s collective memory is disappearing

#82
After publishers successfully sued the Internet Archive over its digital lending program, calling it unauthorized copying

No. The court specifically determined that the Internet Archive was guilty of unauthorized copying. It was not simply an unfounded or unproven allegation. The Authors Guild, the National Writers Union, the European Writers Council, and the Society of Authors in the UK all came out against the Internet Archive, and supported the suit.

Each new restriction limits the archive’s ability to act as a comprehensive backstop.

This self-inflicted damage to the wayback machine is the real tragedy of this entire affair. When IA was asked to stop CDL - many times - founder Brewster Kahle continued. The National Writers Union tried to open a dialogue as early as 2010 but was ignored:

The Internet Archive says it would rather talk with writers individually than talk to the NWU or other writers’ organizations. But requests by NWU members to talk to or meet with the Internet Archive have been ignored or rebuffed.

https://nwu.org/nwu-denounces-cdl/

When the requests to abandon CDL turned into demands, Kahle dug in his heels. When the inevitable lawsuits followed, and IA lost, he insisted that he was still in the right and plowed ahead with appeals. And here we are today.

Re: As AI eats the web, the internet’s collective memory is disappearing

#83
post #14

I feel like collecting, curating, and protecting high quality corpuses of "truth" is going to become increasingly important for high quality AI. There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate…

This already exists, there are archives of Reddit or other sites, and Anna's Archive for papers and books.

Re: As AI eats the web, the internet’s collective memory is disappearing

#84

Earlier quoted context omitted.

> The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid. It does, but do you think that people at that time thought anywhere near as much about preserving their scribbles as we do? I'd venture a guess that we've created more "content" since the advent of the internet than in all of human history prior, and most of it is stored on things that aren't even designed to last…

I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.

> I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.

Well as far as digital letters, photos, menus, etc. are concerned, there's nothing stopping someone with the means and motivation from investing in a doomsday storage facility specifically designed to store these things for posterity. If people can do this for crypto they can do it for digital content.

As for physical items, do you know how much junk Americans have in storage units? The US self-storage industry generates over $40 billion in annual revenue. We're probably keeping more "stuff" in storage units where it has a chance of surviving a zombie apocalypse than at any point in human history.

Re: As AI eats the web, the internet’s collective memory is disappearing

#85
post #25

Funny as I just cancelled my Kagi sub to get the Gemini ai pro sub. The deal was too good to pass up

The deal will be great for now, while they lure you in and get you to drop the competition -- then they will raise prices later.

Then you cancel and go to another provider, rinse and repeat. This is already what is happening with streaming platforms.

Re: As AI eats the web, the internet’s collective memory is disappearing

#86
The article touches on something that I've been thinking about with regards to Google's AI strategy; the automatically-generated AI search summaries are not great. They very frequently confidently misinterpret what the user is searching for and generate half a page of useless information that pushes actual results down the page, and they are occasionally hilariously incorrect, with hallucinated facts.

This is probably a difficult-to-solve problem; given that they generate billions of these a day, not even Google can afford to devote enough compute to each query to reliably generate quality results. You can see this by selecting the "AI mode" from the search interface after getting the mediocre summary - the results are much better and generally perfectly usable. Though even that is probably a special minimal-compute version of the lowest tier of Gemini, it's still maybe an order of magnitude more capable than whatever generates the search summaries.

The bigger problem is that these search summaries are the default and by far the most common interaction that the general public has with "AI", and because this experience sucks, they just assume that all LLMs are similarly stupid and mostly useless. In non-technical spaces I frequently see the argument that "AI" is not useful for anything, all it generates is garbage hallucinations, and almost invariably they cite some actual terrible experience with the Google AI search summary. I would argue that the strategy of adding LLM summaries to every search is the worst of both worlds - it makes classic search worse while poisoning users against the idea of actual LLM-assisted search.

Re: As AI eats the web, the internet’s collective memory is disappearing

#87
post #40

I spent the last three days (off and on) using Gemini to configure my edge router 4 with my iOS devices on a vpn and it's been awesome. In the past I'd do a google search and read a few sources of documentation, do another google search and read another set of documentation. Now, Gemini aggregates multiple pages together so all of the work of reading source docs from multiple locations is now n a single step. Oh, I s…

All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else. If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to re…

I've seen websites put up some draconian measures to try and get a grip on the scraping. So much for the sub-second loading experience when you have Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human. It's made the web browsing experience so much worse.

Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.

Re: As AI eats the web, the internet’s collective memory is disappearing

#88
post #67
post #40

Earlier quoted context omitted.

All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else. If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to re…

> If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.

The danger that's concerning people (rightly or wrongly) isn't that LLMs are going to be an intermediary to your website. It's that they'll be the only thing reading it. No one will ever read your post or know what you wrote. The only consumers will be LLMs, they'll train on a version that strips out you as the author (probably more due to expedience than any sort of malice; it's not like you're famous, are you?), and your idea might get embedded into a set of model weights somewhere. No human will see a byte of it.

Are you actually saying you'd be OK with that?

Re: As AI eats the web, the internet’s collective memory is disappearing

#89
post #5

I occasionally use Google Search when DuckDuckGo fails to give me relevant. Almost always, Google has better results. Though I can find its AI answers annoying aggressive. I'll look up like two search terms and the AI will bullshit multiple paragraphs out of despite having zero context of what I am looking for. DuckDuckGo seems to have detection of whether it should give an AI answer. And it allows you to have more g…

https://noai.duckduckgo.com is a thing fyi i don’t agree with the google has better results thing. sometimes it does. most of the time it’s just that google has the site i want higher in the ordering than DDG. personally i’m fine scrolling down a little bit more. it’s rare i need to go to google for something that DDG doesn’t have at all in their results, but it does happen. i do have to go to google for maps/directi…

> https://noai.duckduckgo.com is a thing fyi

or you can press the gear button -> "Ai features: Manage" -> Search assist

Re: As AI eats the web, the internet’s collective memory is disappearing

#90
post #40

Earlier quoted context omitted.

All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else. If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to re…

Gemini can just consume the device documents. There's an incentive for device makers to publish this content.

There had always been some incentive for manufactures to publish device documentation, and yet it has often been quite lacking either in quality or overall existence. I doubt LLM/agents being the readers will change that at all. What I expect AI scraping and using without credit will impact is people publishing their own unofficial help and guidance, and the affect there is likely to be negative. It won't stop all of them, but enough to be noticeable. Another possible negative is the manufactures documentation being AI generated without sufficient review, so possibly more erroneous than before, or intentionally not producing full documentation at all and expecting AI to fill the gap (MS seems to be heading this way: pushing "ask copilot" all over Azure instead of links direct to good reference material). All this would add up to a situation that is somewhere between "a little worse than pre-AI" and "an absolute shit show".
Post reply on HN