Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

261–270 of 365 posts

Re: Only Google is really allowed to crawl the web

#261

Earlier quoted context omitted.

Interesting that the most comments it got before was 11, and today it succeeds and makes it to the front page! This is a good illustration of whether or not submissions get any traction can be fairly stochastic. On topic, stack overflow does exactly what the article is talking about; They lock down their sitemap and make special exceptions for the Google bot: https://meta.stackexchange.com/a/98087 https://meta.stacke…

I think it's partly because they create a website which reported on the status of the Ever Given which rose to 1. on the front page. I feel like I often see submissions which are, even tangentially, related to front page material rise very quickly. Regardless, congrats to Knuckleheads Club for fighting the good fight.

You are right, that was how I found it.

Re: Only Google is really allowed to crawl the web

#262

Maybe a naïve question but what prevents Knuckleheads’ from ignoring the robots.txt and crawl the side anyway? And if it's so easy to do, how does Google have a monopoly on crawling then?

On smaller sites, nothing usually. But on bigger sites you will be blocked. You will probably be blocked even if you do follow robots.txt

Re: Only Google is really allowed to crawl the web

#263
post #202

Earlier quoted context omitted.

So, one more reason to hate Cloudflare and every single website that uses it.

Or maybe don’t “hate” folks who are just trying to put some content online and don’t want to deal with botnets taking down their work? You know, like what the internet was intended for.

Internet was certainly not intended for centralization. I hit Cloudflare captchas and error pages so often it's almost sickening. So many things are behind Cloudflare, things you least expect to be behind Cloudflare.

Re: Only Google is really allowed to crawl the web

#264
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

Broadly speaking, robots.txt files are often ignored. I used to run a fairly large job ad scraping organization, and we would be hired by companies (700 of the fortune 1000 used us) to scrape the job ads from their career pages, and then post those jobs on job boards. 99 of 100 times, the robots file would disallow us to scrape. Since we were being paid by that company's HR team to scrape, we just ignored it because…

> Broadly speaking, robots.txt files are often ignored.

If you wanna go nuclear on people who do that, include an invisible link in your html and forbid access to that URL in your robots.txt, then block every IP who accesses that URL for X amount of time.

Don't do this if you actually rely on search engine traffic though. Google may get pissed and send you lots of angry mail like "There's a problem with your site".

Re: Only Google is really allowed to crawl the web

#265
post #63
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

I swear something like 50% of those digests are totally incorrect as well. It's amazing they have kept the feature because it has never had a very high signal-to-noise ratio. I never trust what's presented in these digests without double-checking the source page.

Have you heard the story of Thomas Running? It’s a story Google will tell you.

(Search who invented running)

Re: Only Google is really allowed to crawl the web

#266
post #211

Earlier quoted context omitted.

Not sure why you're being downvoted; I completely agree with what you're saying (modulo questionable usage of "sponsor"). If Wikipedia were to try to charge for this use of their data, Google would likely make it a priority to drop the Wikipedia blurbs, either without replacement, or with data sourced elsewhere.

Given the scale that google already operates at, I don't doubt that they would just take a copy of thr content and rebrand it as a google service, complete with user contribution. Then, after two or five years, let it fester then abandon it. Nobody gets promoted for keeping well oiled machines running.

Remember Knol? https://en.wikipedia.org/wiki/Knol?wprov=sfti1

It was actually good for writing stuff when I tried it. Never brought in enough traffic. Killed.

Re: Only Google is really allowed to crawl the web

#267

Earlier quoted context omitted.

BA had some tracking request inline on the “payment processing” page which when blocked by my pihole prevents me from ever getting to the confirmation page, just have to refresh your email and wait for the best. I have no idea how these companies, which make quite a decent amount of money at least up until 2020, can have such utterly poor sites. I once counted some 20+ redirects on a single request during this proces…

I don’t know what they’re doing but most every single sign on tool I’ve seen redirects 10-20 times during the sign on process (and then dumps you to the homepage to navigate your way back).

Probably to get first party cookies on a handful of domains

Re: Only Google is really allowed to crawl the web

#268
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

Large swaths of web are garbage. Wasting people's time and attention on visiting pointless sites for something presentable in a small box is obviously not economical.

And if some of the sources somehow die? New sources will spring up. It doesn't matter.

Re: Only Google is really allowed to crawl the web

#269

Earlier quoted context omitted.

Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…

>> Google donated $2M [1]. In 2010, Google also donated $2m [2]. $2 Million a year? Now I know why Googlers complained about having one less olive in their lunch salad. How much does Google PROFIT from Wikipedia and how much does Wikipedia loses in fundraising when Google fails to send users to the info provider?

Wikipedia is drowning in money so this whole line of discussion is weird.

And most of the value of wikipedia is created by its unpaid users, not Wikimedia foundation.

Re: Only Google is really allowed to crawl the web

#270
post #155

Earlier quoted context omitted.

I don't think google muscling out intermediaries like Expedia is a good thing. Just for example, Expedia is probably 5% of Google's total revenue and Google doesn't like slim margin services by and large that can't be automated. Travel is fairly high-touch - people centric. It doesn't fit Google's "MO". But... its shitty that google can play all sides of the markets while holding people ransom to mass sums of money t…

>In essence, you're advocating that eBay goes away because google could do it... they could.. and eBay is technically just an intermediary, but do we want everything to be googlefied? I don't think I'm really advocating for it as much as I see as a more or less neutral change. That said, I'm pretty ambivalent about Google. Their size is a concern, but they also tend to be pretty low on the dark pattern nonsense. eBay…

Companies opt in to sites like Expedia and list their properties/flights/vacations on their marketplace and they pay a commission for those being booked. Expedia doesn't just crawl them and demand a royalty for sending them traffic...

Google has a huge pay 2 play problem with PPC... i've worked for Expedia so that's the only reason i know this :)

It's the reason companies work with Expedia many times because they don't have the leverage expedia group does...

i see it as unnatural change btw... "borg" if you will.

Post reply on HN