Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

221–230 of 365 posts

Re: Only Google is really allowed to crawl the web

#221
post #5

The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…

Common Crawl is attempting to offer this as a non-profit: https://commoncrawl.org

o/t but what the hell are they doing to scroll on that page? I move my fingers a centimeter on my trackpad and the page is already scrolled all the way to the bottom.

Hijacking scroll like this is one of the biggest turnoffs a website can have for me, up there with being plastered with ads and crap. It's ok imo in the context of doing some flashy branding stuff (think Google Pixel, Tesla splashes) but contentful pages shouldn't ever do this.

Re: Only Google is really allowed to crawl the web

#222
post #209

Earlier quoted context omitted.

Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…

> Wikipedia isn't monetized. No, but they often ask for donations when you visit the site, which people won't see if they just see the in-line blurb from Wikipedia on the Google results page. > In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. $2M is a pittance compared to what I expect Google believes is the value of their Wikipedia blurbs. If Wikipedia could charge for use of this data (which…

Unlikely that Wikipedia will be able to charge for content, seeing as all of their content is CC-BY-SA licensed. https://en.wikipedia.org/wiki/Wikipedia:Licensing_update

They may be able to charge for bandwidth (if you want to use a Wikipedia image, you can use Wikipedia's enterprise CDN instead of their own), but their licensing allows me to rehost content as long as I follow the attribution & sublicensing terms.

Google has no problem operating their own CDNs, so I find it unlikely that Wikipedia will be able to monetize Google search results in such a manner as you described.

Disclaimer: I work for Google; opinions are my own.

Re: Only Google is really allowed to crawl the web

#223

On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…

I wonder what happens to RSS feeds in this situation. Programs I run that process RSS feeds will just fetch them over HTTP completely headlessly, so if there are any CAPTCHAs, I'm not going to see them.

Re: Only Google is really allowed to crawl the web

#224
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

> In short, I don't think the crawler is the problem. Except that, allow other companies to crawl/compete, and you can take eyeballs away from Google (which may well then return eyeballs to Wikipedia so long as the Google competitors don't also present scraped data).

[deleted]

Re: Only Google is really allowed to crawl the web

#225
post #93
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

So from the website's point of view there is no difference between 'crawling' and 'scraping'. Census.gov I assume has a ton of very useful information which is in the public domain which a host of potential companies could monetize by regularly scraping census.gov. Census.gov's purpose to make this information available to people is served by google, yahoo and bing. On the other hand if I have a business which is bas…

I used to run a fairly large job ad scraping operation. Our scraped data was used by many US state and federal job sites. "Scraping" is just using software to load a page and extracting content. "Crawling" is just load a page, find hyperlinks (hmm... a kind of content), and then crawling those links. Crawling is just a kind of scraping.

Re: Only Google is really allowed to crawl the web

#227
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

I've noticed that sometimes Google had updated flight information before the displays at the airport.

For the most part individual airports own that infrastructure. So it's hard to generalize. For most types of notable flight status/time changes, however, airlines usually know first.

There are exceptions, like an airport-called ground stop.

Re: Only Google is really allowed to crawl the web

#228
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

Broadly speaking, robots.txt files are often ignored. I used to run a fairly large job ad scraping organization, and we would be hired by companies (700 of the fortune 1000 used us) to scrape the job ads from their career pages, and then post those jobs on job boards. 99 of 100 times, the robots file would disallow us to scrape. Since we were being paid by that company's HR team to scrape, we just ignored it because getting it fixed would take six months and 22 meetings.

Re: Only Google is really allowed to crawl the web

#229
post #28
post #5

The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…

Imagine if Donald Trump decided that indexing Joe Biden's campaign site was unacceptable. A mandated singular public cache has potential slippery slopes.

>A mandated singular public cache has potential slippery slopes.

That may be, but it seems like everything has a slippery slope - if the wrong person gets into power, or if the public look the other way/complacence/ignorance/indifference, etc, etc. It shouldn't stop us evaluating choices on their merits, and there is a lot of merit to entrusting 'core infrastructure' type entities to the government - or at-least having an option.

Re: Only Google is really allowed to crawl the web

#230

On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…

How hard is it to ask Cloudflare to let you crawl?
Post reply on HN