Live data from Hacker News

google.com/goto: Google's anti-scraping update

autom.dev

151–160 of 545 posts

Re: google.com/goto: Google's anti-scraping update

#151

Note: Article published by Autom.dev which, from a quick read of their homepage, seems like it scrapes Google search results in violation of Google's terms of service and sells those results to customers via an API. That's just my quick read of it, though, so this could be wrong.

Note: Irrelevant.

The reported behavior exists, does it not?

Anecdata: I've observed this behavior for several weeks already as a regular user without a Google account and there are countless comments from regular users reporting the same behavior.

Re: google.com/goto: Google's anti-scraping update

#152

While a lot of people are concerned with local model performance, I wonder how feasible is it now to run a local indexed web search? Surely running an old school Google is possible with the beefy AI rigs today. I know the problem will be crawling which would be bottlenecked by the ISP but I use Google to search SO, Wikipedia, programming language docs, Github issues, and AWS docs. I think a feasible workflow would be…

Impossible. The majority of websites firewall automated crawler traffic (because of the rise of the bots), only making exceptions for the largest search engines. There is no possibility of starting a new crawler.

Re: google.com/goto: Google's anti-scraping update

#154
post #94

As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our cor…

> Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our corpo-political masters". As long as you don't search anything related to Russia itself, or to Russian interests elsewhere (like their invasion of Ukraine). Then it's very heavily biased, priority is given to state ran propaganda mills, independent media is hidden from results, etc. And before anyone start…

So, I decided to check, and the first links that show up when I search for "war ukraine" in yandex.com are Google news are from rbc.ua and bbc.com. When I search for it in yandex.com, the first link is CNN, the second google news, the third is indeed TASS, but the fourth is novayagazeta.

I then picked "bucha" as the query most likely the be affected by propaganda and censorship. In yandex.ru, there are propaganda outfits in the result (something called ruwiki.ru comes third), most links are of fairly direct accounts of the massacre by independent media. The AI summary says the town was "occupied" by the Russian army and "liberated" by the Ukrainian one and "После отступления российских войск в городе были обнаружены многочисленные свидетельства массовых убийств мирных жителей." On yandex.com on the other hand, the first page of results does look fairly propagandistic, including "globalresearch.ca" and "donbass-insider.com", which seem like propaganda outfits. But it also does include accounts from novayagazeta, al Jazeera a video of killings from Radio Free Europe.

While they do put either own propaganda, neither Western, nor Ukraininan nor independent media is not hidden from the results and it's not visibly de-prioritized.

Of course, I wouldn't discount the possibility that the results would look quite different from Russian territory. And, if anything, half the reason why Russian propaganda is so effective is that its self-aware and capable of subtlety when needed. "Fuck you if you can't handle the truth, this version of Biden is the best version ever." would never happen there. Unlike Scarborough, Soloviev knows exactly what he is.

Open and constant suppression of information is not the regime's usual strategy in information space (though, they'd ratchet up the level of control when they deem it necessary), which is why I knew the claims here wouldn't stand up to scrutiny.

Re: google.com/goto: Google's anti-scraping update

#157

So why are we angry about that ? I mean the end result for the users are exactly the same, it matters only for bots. Google have such a (justified) bad reputation that whatever they do, people assume it’s entishification. I don’t believed it is on that matter.

Privacy-wise, it enables them to track the result you click on.

Also, though more niche, it would make archived search result pages (e.g. on the Wayback Machine) less useful.

Re: google.com/goto: Google's anti-scraping update

#158
I did an interview with Google around 20 years ago, where they posed a challenge involving tracking which specific search results people click. It's obvious in hindsight the solution required rewriting all the urls to redirect through their servers. Note this was in the days before they already did so as a matter of course.

I failed to gain traction on the problem, because to me the very idea of doing such a thing was too reprehensible to seriously consider. It broke an unwritten contract between the company and the user's expectation of how websites worked. You expect to be able to do things like right-click a link and copy the authentic URL, or hover to see where it wants to take you. The notion of obfuscating the link beyond easy recognition and polluting it with tracking markers felt misleading and, well, evil. A move that would mainly only benefit Google, and not it's users. I (quite mistakenly) presumed this opinion would be obvious and self-evident to anyone who spent enough time around the early web to understand its norms.

I explored other ways of achieving the goal, but it clearly wasn't the answer the interviewer sought.

I'm more seasoned now, and experienced enough to say with confidence the approach was wrong. This may seem like a small thing, but a series of misteps and chronic failure to adequately advocate for users is what has led us to the toxic waste dump that so much of the Internet has become today.

I'm really glad to have fresh alternatives (like Kagi), and can't wait for the cultural zeitgeist among developers to swing back around to valuing users as human beings and living up to the trust they place in us.

Re: google.com/goto: Google's anti-scraping update

#160

As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our cor…

Sometimes I use my own index of domains and channels

https://github.com/rumca-js/Internet-Places-Database

I hate the very idea that destination location is opaque, and user can't verify if you are funneled toward malvertizing.

I hate even base64 encoded links in the results.

Post reply on HN