Live data from Hacker News

google.com/goto: Google's anti-scraping update

autom.dev

401–410 of 545 posts

Re: google.com/goto: Google's anti-scraping update

#401

Earlier quoted context omitted.

How is that different for any web site on Earth?

As a general rule, if another country has an open war with your neighbor and openly says that you may be next, don’t put in private data into the sites controlled by that country.

But that's true for google too

Ok, maybe not the war *with a neighbor* (yet), but US has an open war with Iran (and has bombed a few other countries recently also) and it also threathens to annex the neigbor to the north plus another island country.

Sure, americans might consider themselves to be "the good guys", but the rest of the world doesn't.

Re: google.com/goto: Google's anti-scraping update

#402

Earlier quoted context omitted.

Kagi is mostly based on Bing and Google results, officially by API and not scraped, and ranked by there own mechanism I think.

Officially Kagi pays SerpAPI to scrape Google for them, since Google doesn't have an API they'll let Kagi use.

Interesting. So this means the new google.com/goto will exclude Kagi from the results?

Re: google.com/goto: Google's anti-scraping update

#403

Earlier quoted context omitted.

Why are you defending the ruination of one of the basic principles of hypertext? Why are Google apologists flagging my other comments?

Because Google search isn't a hypertext document, it's a web app that presents dynamic ephemeral content to the user. There's no need for stable hyperlinks, since either you will click them within a few seconds, or they will disappear forever, and it shouldn't be a problem to indirect via Google when clicking them, since you just made a request to Google to get the link in the first place.

[deleted]

Re: google.com/goto: Google's anti-scraping update

#404
post #22

Earlier quoted context omitted.

Because I wouldn't want Google to know what links I click on. When I used to use Google, and they added such redirects that included the destination URL as a parameter, I'd edit the link to make it just that URL before resolving it. This new scheme would make that impossible.

They own the JS on the page, they don't need the redirects to know where you clicked...

[deleted]

Re: google.com/goto: Google's anti-scraping update

#405

Earlier quoted context omitted.

Kagi works by scraping Google and other engines, so this should still worry you. I still support the use of Kagi though, as a market signal to Google that we'd even pay them for their product if it wasn't dogshit.

Kagi don't scrape, they pay other engines for API access: https://help.kagi.com/kagi/search-details/search-sources.htm... >Our search results also include anonymized API calls to all major search result providers worldwide

The only “other engine” that matters is really a thin paid SERP façade in front of a scraped Google result. In other words, Kagi pays for access to an API whose implementation is basically live scraping Google and throwing away the ads.

There are a few other indexes that Kagi’s aggregator mixes in (such as Bing Search) but those haven’t been contributing much in terms of meaningful results. The sheer size of Google’s index still dwarfs all the others, including Bing Search.

So yes, Kagi doesn’t scrape but they pay SERP providers who do.

Re: google.com/goto: Google's anti-scraping update

#406
post #97
post #22

Earlier quoted context omitted.

Because I wouldn't want Google to know what links I click on. When I used to use Google, and they added such redirects that included the destination URL as a parameter, I'd edit the link to make it just that URL before resolving it. This new scheme would make that impossible.

They already knew which links you clicked on though…

[deleted]

Re: google.com/goto: Google's anti-scraping update

#407
post #83

Earlier quoted context omitted.

Yup, same here. Been using Kagi for years. It's boring, it just works. Hoping they can stay that way.

My ongoing concern with Kagi is they always seem to be focused on sidequests like their Orion browser and their LLM-powered Translate tool. Maybe that's interesting for some people, but I can't help but feel like I just want a damn search engine than works.

I also have no interest in their side projects. It doesn't bother me that they have them, though. As you say, as long as the search engine works and the price is acceptable, I'm happy.

Re: google.com/goto: Google's anti-scraping update

#408

Earlier quoted context omitted.

Officially Kagi pays SerpAPI to scrape Google for them, since Google doesn't have an API they'll let Kagi use.

Interesting. So this means the new google.com/goto will exclude Kagi from the results?

Not necessarily, as evidenced by TFA, whose scraper still seems to work even after it visits each individual SERP entry to harvest the URLs.

But Kagi’s SERP API provider is going to have to adapt, I guess.

Re: google.com/goto: Google's anti-scraping update

#409
post #97
post #22

Earlier quoted context omitted.

Because I wouldn't want Google to know what links I click on. When I used to use Google, and they added such redirects that included the destination URL as a parameter, I'd edit the link to make it just that URL before resolving it. This new scheme would make that impossible.

They already knew which links you clicked on though…

How? I never allowed Javascript or anything.

Re: google.com/goto: Google's anti-scraping update

#410

While a lot of people are concerned with local model performance, I wonder how feasible is it now to run a local indexed web search? Surely running an old school Google is possible with the beefy AI rigs today. I know the problem will be crawling which would be bottlenecked by the ISP but I use Google to search SO, Wikipedia, programming language docs, Github issues, and AWS docs. I think a feasible workflow would be…

Wayback Machine full archive is like less than 50PB. Let's say you could strip multimedia and remove every patterned data to compress that into about a petabyte. The per-bit cheapest disk right now is consumer Seagate 24TB(SI; 21.8TiB usable) at ~$500, or around $1200k for just the disks.

Doable if you had couple million dollars to burn. Cheaper than private jets new.

Post reply on HN