Live data from Hacker News

Google Cache is fully dead

seroundtable.com

231–240 of 249 posts

Re: Google Cache is fully dead

#232

Earlier quoted context omitted.

I hope Google is NOT going to be a significant source of funding for the Internet Archive. Because I want to trust Wayback Machine and the Internet Archive to be unbiased. Google likes to influence search results, hiding ones it doesn't like, and elevating those that the Company supports. Wayback Machine has been very reliable so far, I hope it stays that way.

Generally speaking, the Wayback Machine is not searchable in the fashion that Google is, there isn't a scale to put the thumb on. (There are some tiny subsets which are rudimentarily full-text searchable; and some efforts to make domains findable. But nothing remotely like even Google 1.0 mapping URIs to organic terms.)

Which is really unfortunate because a lot of useful information is now essentially unfindable unless you know where it is. Maybe someone should start crawling and indexing the internet archive.

And no, the IA is already not an impartial archive: https://archive.org/post/1126216/kiwi-farms-removed-from-way...

Re: Google Cache is fully dead

#233

Earlier quoted context omitted.

The “problem” is that society doesn’t see Nintendo/Disney/et al as copyright trolls - instead they’re successful businesses who made content and profit. Connecting those dots to archival work and historic preservation is a long slow process and won’t be successful in courts without legal changes.

We have to stop prioritizing it over everything else. You can't compete in the global playground if you have impossible to implement entitlement programs. Priority has to be new work not existing work and definitely not the work of dead people. We have countless similar schemes were people are to be rewarded for things done long ago. One can't pretend it isn't slowing everything down.

I wonder if reframing it as getting rid of hereditary wealth would make more people open to copyright reform.

Re: Google Cache is fully dead

#234

Earlier quoted context omitted.

Everyone just ignoring bad laws and contradicting them can remove laws too. But of course this is a niche topic that would never get such broad support. A lot of people smoking weed is certainly a component for the prohibition to fail at some point. Writing mails to legislative members isn't enough if you don't have any form of leverage.

Jury nullification is the real mechanism for We The People when we don't consent to be held to laws passed by They The Wealthy/Bribed Lawmakers It requires that people refuse plea deals and demand jury trials, and that the jury is educated on what jury nullification is but when prosecutors can't get a conviction regardless of much evidence they have of guilt the laws will get changed or at least they stop being enfor…

AFAIU the jury doesn't have final say in civil trials so this would only work partially (copyright infringement can be both a civil dispute as well as a criminal matter).

Re: Google Cache is fully dead

#235
post #167
post #135

Earlier quoted context omitted.

I disagree. I am happy the Internet Archive are fighting the draconian copy right laws that exist.

Back over here in the real world, there is no possible way the IA will win this fight through the courts. It has to be dealt with by legislation.

So far no one has won the fight by legislation. Or even a single battle. In fact, copyright terms have only gotten worse over the years - FAR worse.

Re: Google Cache is fully dead

#236

> Then a couple of weeks ago, added [direct] links to the Wayback Machine Hopefully they are also making substantial donations to the Internet Archive, since they will be directing a lot of traffic into it and basically using their infrastructure as a feature on their main product... EDIT: Apparently they are collaborating but there are not much details [0] [0] https://blog.archive.org/2024/09/11/new-feature-alert-ac…

IA needs an alternative - an independent backup archive - more than it needs funding. Unless IA funding exceeds the entire US copyright lobbying industry there is always a chance they will cease to exist without enough notice to save the data somewhere else.

There is also the matter what IA will be able to archive. The the machine learning gold rush more and more site operators see dollar bills in front of them and are restricting who can crawl their content. Google is in a special position here because almost no one can affort not to be crawled by Google which is what made their cache especially valuable in addition to the IA.

Re: Google Cache is fully dead

#237

Google Cache was useful because you could sometimes not find a term or keyword in the web site, but it would be in the cache. Or for sites that have gone offline, or no longer have the item. "It's still in the Google Cache!" you can't say that anymore. I use Google less and less these days. What's the point when you can just ask an LLM, and it gives you an answer within seconds, with no ads? You can ask for reference…

LLM… no ads…. *For now

*That you notice.

Assuming that the output isn't biased towards the operator's interests is naive.

Re: Google Cache is fully dead

#238

Earlier quoted context omitted.

This still happens all the time. * I search a keyword * I see a google result * I see the keyword IN THE PREVIEW on Google * I click on the link * No keyword And this isn't hidden SEO spam stuff, it was literally removed. The cache doesn't match the live result. No recourse.

Another annoying scenario is when the search result isn't the actual page/article/… itself, but only a snippet within a site's own (paginated) index. At the time Google had indexed that page, the article preview you were looking for was maybe on page 5 of that index, but by the time you're arriving, it might have moved to page 11 because of all the additional content that got added since then. With online shops it's…

Even worse are the related articles/posts some sites like to show next to the main content. It really shows Google's lack of progress that they still haven't figured out a way to handle these cases better.

Re: Google Cache is fully dead

#239

I would presume Google still has all this data. They just will not let anyone else use it. Could this be an advantage that Google can use to train their models on but others won't have access? Google wants it to be more difficult to notice rewrites? Journalists to often have found valuable information with it?

As I understand it, Google does a decent amount of rendering of a page before indexing; this a) allows it to index content loaded by JS and b) prevents some ways spammers show Google different content from users. Perhaps Google's main way of storing a page no longer matches something that can be easily served as a cache page. This might be a way to remove a legacy copy of each page and reduce storage costs.

> prevents some ways spammers show Google different content from users.

Google obviously hasn't cared about that for a long time.

Re: Google Cache is fully dead

#240
post #16

Any solid evidence on why, or why now? I have to assume the additional interest in crawling/scraping data for AI precipitated this. Why deal with all the messiness of crawling the web at large when you can use a Google search and cache: results as your RAG?

Probably yes. Or websites that google has made scraping deals with don't want a cache of their content to be publically available and the easiests thing to do was to just turn off the pulic cache completely.
Post reply on HN