Live data from Hacker News

google.com/goto: Google's anti-scraping update

autom.dev

521–530 of 545 posts

Re: google.com/goto: Google's anti-scraping update

#521

While a lot of people are concerned with local model performance, I wonder how feasible is it now to run a local indexed web search? Surely running an old school Google is possible with the beefy AI rigs today. I know the problem will be crawling which would be bottlenecked by the ISP but I use Google to search SO, Wikipedia, programming language docs, Github issues, and AWS docs. I think a feasible workflow would be…

For some of that, you don't need to crawl. Wikipedia offers database dumps which you can download in one go. Lots of programming docs are managed in repos, so you can clone the repo instead. Even stackoverflow seems to have a snapshot dump (https://archive.org/details/stackexchange).

Re: google.com/goto: Google's anti-scraping update

#522

Earlier quoted context omitted.

Hamas claims that the journalists are Hamas, not Israel. And I'm not declaring any journalist "good" or "bad". I'm stating that a Hamas member is a valid target. Him moonlighting as a journalist does not change that.

Every claim need to be proven and everyone deserves a process. In a democracy of course, it may be different in a fascist state committing a holocaust.

That's not how things work when you are participating in hostilities in an active war zone. What a ridiculous statement to make.

Re: google.com/goto: Google's anti-scraping update

#523

Earlier quoted context omitted.

Everything uses energy. AI is uniquely bad due to its scale. Comparing AI inference and training with posting a one-line comment on a website is like comparing wildfires with candles because they both produce heat. It's true but useless as an argument.

Complaining about AI is just a meme. There are tons of things that use way more energy for arguably way less utility to humans.

AI was not even a thing a few years ago so almost anything that use more energy than AI is probably far more useful

Re: google.com/goto: Google's anti-scraping update

#524

As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our cor…

+1 for Yandex. I was not expecting it at all but agreed you get most of organic search results which is what search used to be based on relevancy.

Re: google.com/goto: Google's anti-scraping update

#525
"Combined with earlier moves like removing &num=100 and tightening BotGuard/SearchGuard, Google is steadily raising the cost of naive SERP scraping."

Another "move" is suing companies like Autom, e.g., SerpApi

Google's Amended Complaint from their suit against SerpApi

https://ia801008.us.archive.org/25/items/gov.uscourts.cand.4...

"30. Copyright holders have authorized Google to implement access controls like SearchGuard for the content they license to Google, and in some cases insisted that Google do so. Googles authorization takes many forms. For example, Google has an agreement with a prominent licensing partner that holds copyrights to millions of works that it licenses Google to use in its Search results. Under the parties agreement, versions of which date back to 2017, Google is not only authorized, it is obligated to use commercially reasonable efforts to safeguard the licensed content against unauthorized third-party access. Other license agreements contain similar obligations. For example, another major content provider requires that Google ensure the content it licenses will not be available for download by third parties, thereby authorizing the implementation of technical access controls."

"31. In other cases, Googles authorization to implement access control measures like SearchGuard is part and parcel of the grant of licenses themselves, as Google and its licensors recognize that the value of the licensed rights would be undermined if others were free to access, take and resell the licensed content without restriction. For example, Google has a licensing agreement with Reddit, under which Reddit licenses Google to use the copyrighted content of both Reddit and its users in Search Services."

"32. Googles licensing partners have also expressly requested that Google prevent unauthorized access to licensed content. For example, when Reddit suspected that scrapers like SerpApi were accessing, taking, and reselling the content that Reddit had licensed to Google, it specifically asked Google to employ technical measures to prevent such unauthorized appropriation."

But this does not account for material that is not covered by the "license with a prominent licensing partner", its license with "another major content provider" or its agreement with Reddit

Google not only uses SearchGuard on SERPs containing links to the content covered by these licenses, it uses SearchGuard on _all_ SERPs

Google needs more than a "goto" update. It needs to update its terms to require _all_ copyright holders for the materials it has indexed and cached to give Google authorisation to use "technological protection measures" to deny access to certain members of the public, e.g., Google's perceived competitors including any Google user who "searches too fast"

Re: google.com/goto: Google's anti-scraping update

#527
post #174

Earlier quoted context omitted.

You can block Google Analytics with your browser. You can't block needing to ask Google for the URL of the result you want to visit.

If you're that concerned, use another search engine

My responding to comments doesn't mean I use them.

Re: google.com/goto: Google's anti-scraping update

#528
post #493

Earlier quoted context omitted.

FWIW Kagi has worked hard on their pricing over the last few years and has been trying different models. They're not "big search" the price is what it is so it can exist as a business. 100% fine to not be a customer obviously, but then you forfeit your license to complain about Google spying on you and ruining the web. Like they're trying to solve the problem. IMHO speaking just for myself I feel a moral duty to supp…

You can not find the value of 10 searches a day for 5 dollars and complain about Google. Yandex,DDG and Bing are major search engine choices. Kagi has limitations like not showing sites with ads which many people with ad blockers don't mind.

What do you mean Kagi doesn't show sites with ads ?

Re: google.com/goto: Google's anti-scraping update

#529

Earlier quoted context omitted.

Kagi don't scrape, they pay other engines for API access: https://help.kagi.com/kagi/search-details/search-sources.htm... >Our search results also include anonymized API calls to all major search result providers worldwide

They pay scraping sites for results, including SERP API. When I pointed this out, last time Kagi was discussed on Hacker News, an employee of Kagi said that they're trying to build their own internal index, but he didn't provide details.

They use a ton of different search index's from what I've seen. Kagi Small Web is also awesome

I believe they are building their own index but targeted towards useful results not included in the others

Re: google.com/goto: Google's anti-scraping update

#530

Earlier quoted context omitted.

Impossible. The majority of websites firewall automated crawler traffic (because of the rise of the bots), only making exceptions for the largest search engines. There is no possibility of starting a new crawler.

It would be interesting to see a decentralised, residential collective that builds and publishes an index. There are surely enough interested people on HN alone that would be willing to run software at home to scrape a small slice of the internet.

residential proxies are one of the leading causes of what's killing the internet right now
Post reply on HN