Live data from Hacker News

google.com/goto: Google's anti-scraping update

autom.dev

231–240 of 545 posts

Re: google.com/goto: Google's anti-scraping update

#231
I really like the idea of turning the search index into a public utility. It is one of those natural monopoly coordination problem things. Just quasi nationalise it for economic efficiency. Ofc google can still sell adds against their own ui (like everyone else). Hopefully this move will move that idea closer to reality.

Re: google.com/goto: Google's anti-scraping update

#232
post #136

Earlier quoted context omitted.

The age of internet search is over. The age of Cloudflare has begun. It wouldn't be possible to build a search engine now if they wanted to... and there wouldn't be anything to search for anyway. The non-corporate internet withered into dust and blew away in the wind. If you could find what you want, how would they ever sell you what they want you to buy? And I'm not just talking merchandise, though that too. Your po…

I'm hoping for a future where we create static HTML pages again styled with a bit of handmade CSS, because we're so tired of bot attacks and long loading times. Then suddenly a Cloudflare network becomes absolet.

people have been putting their fuckass blogs with 0.5 visitors a day behind Cloudflare long before the increase in bot traffic. and that increase makes fuck all any difference for them anyway.

Re: google.com/goto: Google's anti-scraping update

#233

While a lot of people are concerned with local model performance, I wonder how feasible is it now to run a local indexed web search? Surely running an old school Google is possible with the beefy AI rigs today. I know the problem will be crawling which would be bottlenecked by the ISP but I use Google to search SO, Wikipedia, programming language docs, Github issues, and AWS docs. I think a feasible workflow would be…

SearXNG configured as in the OpenWebUI docs is pretty cool. My "Hello World" with a new agent framework is teaching it to use SearXNG. Hook this in as a tool and the agent can answer a lot of questions.

SearXNG is more of a metasearch, the dude who wrote it pops in on here and is working on a cool sounding project that is more like a local personal search engine, I forget the name, but I've been meaning to check it out.

There is also Common Crawl.

https://docs.openwebui.com/features/chat-conversations/web-s...

https://github.com/brian-learns/xng-agent

Re: google.com/goto: Google's anti-scraping update

#234

As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our cor…

surprising to see how much they have stripped from our view Don't forget cached pages.

You may not even appreciate the full extent of it. Long ago now, when Google took off and started dominating because it was “the best results” people didn’t realize that even though they were the best results, they were not actually good results. To be the best, you just have to be notably better than your nearest competitor, you don’t actually have to be good or even decent, just noticeably better. Today Google really has no competitors besides all the derivatives it has maintained to maintain an illusion of competition, the most obvious examples being DickDuckGo and Bing.

It’s somewhat similar with AI today, only the competitive landscape has majorly shifted with the Chinese models, something that wasn’t supposed to happen. All the sudden you no longer have a closed system of American models that can all be contained and managed as good enough to make people believe are “the best”; they actually have to compete in a far more competitive arena with the tension of the best model being the most competent model, and the most competent model, by virtue of AI, will be the least restricted, the least censored, the least guardrailed, the least controlled to present a telescreen and Hollywood type fantasy world where the good guys always win and … wouldn’t you know it … the good guys is always us, the people controlled by a psychopathic, narcissistic cabal that also controls what AI will tell you, the false truth.

Re: google.com/goto: Google's anti-scraping update

#235
post #190

When Google stopped paid API search a few months ago, I looked for an alternative for my agents that I felt would be sustainable (one-time setup, then out of my mind). I quickly excluded SERP as I feard Google would pull exactly this type of shenanigans to cut them off. I somehow found Mojeek and settled on it. I had never heard of them. Unlike Kagi, their business model is ads (so they hold no particular moral high…

I’m guessing that the knowledge of the techniques to support large-scale web search have diffused out of Google - it has been a few decades after all. Not to mention that the distributed system knowledge that used to live only in Google was either published by Google or cloned in other projects, usually made by ex-Googlers. And you can now rent capacity at scales that 20 years ago required Google to build lots of their own data centers.

Also, there is now a use case for paid search APIs - LLMs and agents - that essentially didn’t exist a few years ago. Not sure why Google hasn’t leaned more into this, but probably a combination of Gemini and Ads interests have combined to view their search index as an increasingly valuable asset, when if anything it might be corrupted already by these interests and therefore be less valuable.

Re: google.com/goto: Google's anti-scraping update

#236

So why are we angry about that ? I mean the end result for the users are exactly the same, it matters only for bots. Google have such a (justified) bad reputation that whatever they do, people assume it’s entishification. I don’t believed it is on that matter.

It breaks the social contract of the web. My user agent should be able to tell me where a hyperlink goes without first clicking on it, and perhaps act differently based on that information. Now, all Google search results appear to go back to Google.

TBF this is more a browser problem than a website problem to my mind. Redirects shouldn't be handled in such a cavalier manner. I should be able to configure the browser to stop and confirm the destination any time the domain changes without direct user interaction.

Re: google.com/goto: Google's anti-scraping update

#237

As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our cor…

How is this comment relevant to the post (and be so upvoted?) Google is making it harder for people to game the index, working in your favor.

Mmmh this change doesn't seem about gaming the index but rather about scraping the index. Which affects in a negative way only Google. And actually would affect users in a positive way because it lets other companies create their indexes more easily (at the expense of a poor mega-corp, indeed)

Re: google.com/goto: Google's anti-scraping update

#238
post #94

As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our cor…

> Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our corpo-political masters". As long as you don't search anything related to Russia itself, or to Russian interests elsewhere (like their invasion of Ukraine). Then it's very heavily biased, priority is given to state ran propaganda mills, independent media is hidden from results, etc. And before anyone start…

Reminds me of this old joke.

A Soviet citizen and an American are sitting next to each other on a commercial flight.

The American turns to the Russian and says, "I have to hand it to you — your state propaganda is very impressive. You really know how to shape what people think."

The Soviet smiles, nods, and replies, "Thank you, but it's really nothing compared to American propaganda."

The American looks shocked and says, "What are you talking about? We don't have any propaganda in America."

The Soviet smiles again and says, "Exactly"

Re: google.com/goto: Google's anti-scraping update

#239
post #183

Earlier quoted context omitted.

I do agree it's kinda expensive. I'm lucky enough that my employer pays the $25 plan for met because it includes access to many LLMs.

> because it includes access to many LLMs. That is interesting, which LLMs? Asking as someone paying monthly Claude subscription.

A lot of them. Claude is a choice. Here is the list : https://help.kagi.com/kagi/ai/assistant.html

You can even choose a model for each turn.

Re: google.com/goto: Google's anti-scraping update

#240
post #183

Earlier quoted context omitted.

I do agree it's kinda expensive. I'm lucky enough that my employer pays the $25 plan for met because it includes access to many LLMs.

> because it includes access to many LLMs. That is interesting, which LLMs? Asking as someone paying monthly Claude subscription.

[deleted]
Post reply on HN