Live data from Hacker News

google.com/goto: Google's anti-scraping update

autom.dev

321–330 of 545 posts

Re: google.com/goto: Google's anti-scraping update

#321
post #110

As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our cor…

The real "old-school Google, but modern" is Kagi, with the caveat of being paid. (Worth the $5 for me.) But LLMs can be commanded. This is may bookmark alias for invoking the spirit of old Gog=ogle from within the new, AI-based Google: https://google.com/search?q=You%20are%20Google%20Search%20fr... (Cleaned / decoded: 'You are Google Search from 2004. Given a search request, provide 10 links to relevant pages, each w…

Thanks, I needed to add a ?aep=11&atvm=2&udm=50 for AI mode. Wow!!!!!

>You are Google Search from 2004. Given a search request, provide 10 links to relevant pages, each with a short text from the page, featuring the search terms. Avoid any pages that do not contain the search terms. Do not simulate fetching results. The URLs must be complete and not truncated. No markdown. The current date is 2026. The search request:

Re: google.com/goto: Google's anti-scraping update

#322

As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our cor…

> For actual web search I actually like using Yandex

You might want to reconsider this.

https://meduza.io/image/attachments/images/007/699/276/large...

Re: google.com/goto: Google's anti-scraping update

#323

Earlier quoted context omitted.

[flagged]

> campaigns for a Russian product on HN are annoying, and appaling if you have any morals How is Yandex more morally-bankrupt than Google in your analysis?

https://meduza.io/image/attachments/images/007/699/276/large...

Re: google.com/goto: Google's anti-scraping update

#324

Earlier quoted context omitted.

[flagged]

Have you tried Yandex and evaluated the quality of its results, or do you just feel obligated to continuously put it down because it's russian?

I did.

https://meduza.io/image/attachments/images/007/699/276/large...

Re: google.com/goto: Google's anti-scraping update

#325
post #94

Earlier quoted context omitted.

> Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our corpo-political masters". As long as you don't search anything related to Russia itself, or to Russian interests elsewhere (like their invasion of Ukraine). Then it's very heavily biased, priority is given to state ran propaganda mills, independent media is hidden from results, etc. And before anyone start…

So, I decided to check, and the first links that show up when I search for "war ukraine" in yandex.com are Google news are from rbc.ua and bbc.com. When I search for it in yandex.com, the first link is CNN, the second google news, the third is indeed TASS, but the fourth is novayagazeta. I then picked "bucha" as the query most likely the be affected by propaganda and censorship. In yandex.ru, there are propaganda out…

Yeah, the Bucha search was a perfect nail in the coffin of previously very respected company: https://meduza.io/image/attachments/images/007/699/276/large...

Re: google.com/goto: Google's anti-scraping update

#326

While a lot of people are concerned with local model performance, I wonder how feasible is it now to run a local indexed web search? Surely running an old school Google is possible with the beefy AI rigs today. I know the problem will be crawling which would be bottlenecked by the ISP but I use Google to search SO, Wikipedia, programming language docs, Github issues, and AWS docs. I think a feasible workflow would be…

Impossible. The majority of websites firewall automated crawler traffic (because of the rise of the bots), only making exceptions for the largest search engines. There is no possibility of starting a new crawler.

This was a knee jerk response to the first paragraph. They weren't talking about a general crawler, but a subset of Wikipedia, stack overflow, programming docs and github. You can download archives of all of those except github, and github could be queried using the api or GH cli

Re: google.com/goto: Google's anti-scraping update

#327
post #206

Earlier quoted context omitted.

> ”This is what capitalism's all about. Take as much as possible and give as little as possible if you want to win. Ideally, give negative amounts.” Your description is vague enough to encompass every system that I’m aware of, from feudalism, through to mercantilism, communism, and capitalism. It may best fit communism and feudalism, both of which are most notable for fostering negative-value-added firms.

I do not understand the type of psychosis which causes people to say that capitalism is communism.

But if someone says "dogs have four legs, and cats have four legs", they aren't saying that dogs are cats.

Re: google.com/goto: Google's anti-scraping update

#328

I did an interview with Google around 20 years ago, where they posed a challenge involving tracking which specific search results people click. It's obvious in hindsight the solution required rewriting all the urls to redirect through their servers. Note this was in the days before they already did so as a matter of course. I failed to gain traction on the problem, because to me the very idea of doing such a thing wa…

Apparently, there's a term for the general practice, it's called "link shimming" - https://www.usenix.org/conference/usenixsecurity20/presentat....

I had some fun in my grad days building a private information retrieval (PIR) scheme so the shim server can do its job without knowing which link you clicked: https://github.com/pncnmnp/shimmey

Re: google.com/goto: Google's anti-scraping update

#329
post #70

Earlier quoted context omitted.

I wouldn't mind that one. That a man in the US and a woman in Japan would get different results for a search for "sushi restaurant" is perfectly reasonable (even if the woman in Japan was searching in English). It's when two people in the same neighborhood got different results for the same search that I said "wait a minute, they're personalizing search results now for ad-targeting purposes" and ditched them. The pot…

When I search Python I should get the programming language but when my neighbour searches Python he should get snakes. That is a valid use of personalised search.

If the query is just 'python', then both you and your neighbour should get some links for the language and some for the animal.

If it only gives you the one you already know about, it's a useless tool.

Re: google.com/goto: Google's anti-scraping update

#330

Earlier quoted context omitted.

Kagi works by scraping Google and other engines, so this should still worry you. I still support the use of Kagi though, as a market signal to Google that we'd even pay them for their product if it wasn't dogshit.

For Google, the business side of you paying them doesn't work out as obvious as it might seem. Advertisers are sold the idea that 1000 clicks/impression is X dollars. The average impression might be priced at 1000/X$. But you are worth much more than that for two reasons. - you are a person willing to pay 5$ - Google can sell a “collectivized” product. The same way that eg collective farmer crop insurance creates act…

The advert paradox, anyone willing to spend money to avoid ads is inherently worth more to advertisers, since they're willing to spend money.

The more money you're willing to pay to avoid ads the more you're worth to advertisers.

That's why tools like uBlock and YouTube morphe are the only answer, negotiating with terrorists never works.

Post reply on HN