Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

171–180 of 365 posts

Re: Only Google is really allowed to crawl the web

#171
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…

Couldn't you make a similar argument about for-profit uses of free/libre software? The software serves a useful purpose, who cares where it came from?

Re: Only Google is really allowed to crawl the web

#172

This isn’t illegal and it isn’t Google’s fault Right there in the article..

Again, with critical context.

This isn’t illegal and it isn’t Google’s fault, but this monopoly on web crawling that has naturally emerged prevents any other company from being able to effectively compete with Google in the search engine market.

Re: Only Google is really allowed to crawl the web

#173
post #23

If the shared cache ever became significant enough to matter it would be devastated by marketers, scammers and other abusers. Google employs the groomers that make their index at least tolerable, if still clearly imperfect. Without that cadre of well compensated expertise to win the arms race against such abusers the scheme is not feasible. I suppose this could be crowdsourced if I didn't know about politics and how…

I don't really understand your comment. Marketers, scammers and other abusers already publish to the web with the intention to be included in a crawl. Postprocessing crawl data is already a thing. Assuming this hypothetical shared crawl cache were to exist, it does not preclude google (and all consumers of that cache) doing their own processing downstream of that cache. Does it? What's the new attack vector?

> I don't really understand your comment.

If you don't then you fail to appreciate the amount of labor it takes to thwart bad actors from ruining indexes. Abusers do publish to the web, and we enjoy not wallowing in their crap because small army of experienced and expensive people at a select few Big Tech companies are actively shielding us from it.

It's easy to anticipate the malcontent view; 'Google spends all its resources on ads and ranking and we don't need all that.' That is naïve; if Google completely neglected grooming out the bad actors people wouldn't use Google and Google's business model wouldn't be viable.

So the obvious question is; where is this mechanism without Google et. al? Will the published caches be 99% crap (and without an active defense against crap you can bet your life it will) and anything derived from it hopelessly polluted? If so then it isn't viable.

Now the instinct will be to find a groomer. Guess what; that's probably doomed too. No selection will be impartial to all, so you get to fight that battle. Good luck.

Re: Only Google is really allowed to crawl the web

#174
post #93

Earlier quoted context omitted.

So from the website's point of view there is no difference between 'crawling' and 'scraping'. Census.gov I assume has a ton of very useful information which is in the public domain which a host of potential companies could monetize by regularly scraping census.gov. Census.gov's purpose to make this information available to people is served by google, yahoo and bing. On the other hand if I have a business which is bas…

I'm generally anti business. But I have to disagree. "The Public" that the government serves includes businesses. Businesses (ignoring corporate personhood bullshit) are owned and operated by people. I do not want the government deciding "what purposes" e.g. non-commercial, serve the public good. The public gets to decide that. (charging a license for commercial use is maybe ok (assuming supporting that use costs gov…

> I do not want the government deciding "what purposes" e.g. non-commercial, serve the public good. The public gets to decide that.

the public's "decision" on things like this is made manifest by government policy, no?

Re: Only Google is really allowed to crawl the web

#175
post #93
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

So from the website's point of view there is no difference between 'crawling' and 'scraping'. Census.gov I assume has a ton of very useful information which is in the public domain which a host of potential companies could monetize by regularly scraping census.gov. Census.gov's purpose to make this information available to people is served by google, yahoo and bing. On the other hand if I have a business which is bas…

Having data in the right format as a download or via an API would be the best way to go for public data.

If people have to 'scrape' that data from a public resource, I'd say they're presenting the data in the wrong way.

Re: Only Google is really allowed to crawl the web

#176
post #5

The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…

An alternative but similar idea, apply your own algorithms to a crawler/index. That's half the problem with these large platforms commanding the majority of eyeballs, you search the entire web for something and you get results back via a black box. Alternatives in general are most definitely a good thing.

Knuckleheads' Club at the very least are doing a great job of raising awareness and the potential barriers to entry for alternatives.

Re: Only Google is really allowed to crawl the web

#177
post #93
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

So from the website's point of view there is no difference between 'crawling' and 'scraping'. Census.gov I assume has a ton of very useful information which is in the public domain which a host of potential companies could monetize by regularly scraping census.gov. Census.gov's purpose to make this information available to people is served by google, yahoo and bing. On the other hand if I have a business which is bas…

But Google, Yahoo and Bing are also monetizing the data. Why are they allowed to provide “benefits” but “scrapers” are not? Why is it wrong to monetize public data?

Re: Only Google is really allowed to crawl the web

#178
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…

Well then they can't nag users to donate to Jimmy Wales' trust fund.

Re: Only Google is really allowed to crawl the web

#179
post #75
post #63

Earlier quoted context omitted.

I swear something like 50% of those digests are totally incorrect as well. It's amazing they have kept the feature because it has never had a very high signal-to-noise ratio. I never trust what's presented in these digests without double-checking the source page.

I remember when rich snippets (one type of those widgets) came out there were a lot of funny examples. One for a common query about cancer treatments that pulled data from a dodgy holistic site saying that "carrots cured most types of cancer" (or something like that). There was a similar one where Google emphatically claimed a US quarter was worth five cents in a pretty and large snippet graphic.

The most memorable rich snippet humor I've seen is a horse breeder sharing a story of how her searches gave snippets with my little ponies as the preview image.

Re: Only Google is really allowed to crawl the web

#180
I tried to set up YaCy [1] at home to index a few of may favorite smaller websites, so I could quickly search just them. That turned out to be a bad idea. Some ended up blocking my home IP address and others reported me to my ISP. None of these sites were that large, and I wasn't continuously crawling them...

[1] https://yacy.net/

Post reply on HN