The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
Perhaps I am misunderstanding or over simplifying things but it always surprises me that there are legal cases brought against companies who scrape data when so many of Google's products are doing exactly this. It definitely feels like one set of rules for them and a different set for everyone else.
Only Google is really allowed to crawl the web
271–280 of 365 posts
Re: Only Google is really allowed to crawl the web
#272The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
What makes you think they care? Killing off the sources of content might even be there goal. If they kill off sources of content, they'd be more than happy to create an easier-to-datamine replacement.
Hypothetically, if they killed off wikipedia, they are best placed to use the actual wikipedia content[1] in a replacement, which they can use for more intrusive data-mining.
Google sells eyeballs to advertisers; being the source of all content makes them more money from advertisers while making it cheaper to acquire each eyeball.
[1] AFAIK, wikipedia content is free to reuse.
Re: Only Google is really allowed to crawl the web
#273The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…
Keeping users from clicking through to organic results helps them generate more revenue.
Re: Only Google is really allowed to crawl the web
#274Earlier quoted context omitted.
You can screwed any time you book a connecting flight on two different airlines even if the times aren't tight. For instance if one is cancelled. If you use the same airline they will make sure you get to the destination.
That's true, but it can save you a ton of money. You just have to be aware of the risks and plan accordingly. I have typically used this strategy when flying back to the US from the EU. Take an EZJet or similar low cost airline from random small EU city to a larger EU city like Paris, London, Frankfurt, etc... and book the return trip to the US from the larger city. I've also been forced to do this from some EU citie…
* UA or Lufthansa round trip (single carrier) $3K
* UA round trip SFO - Paris + Aeroflot round trip Paris - Moscow: $1K
No amount of search could reduce the gap. I went with the second option. The gap is even bigger if you have a route with multiple segments.
Re: Only Google is really allowed to crawl the web
#275Earlier quoted context omitted.
Broadly speaking, robots.txt files are often ignored. I used to run a fairly large job ad scraping organization, and we would be hired by companies (700 of the fortune 1000 used us) to scrape the job ads from their career pages, and then post those jobs on job boards. 99 of 100 times, the robots file would disallow us to scrape. Since we were being paid by that company's HR team to scrape, we just ignored it because…
> Broadly speaking, robots.txt files are often ignored. If you wanna go nuclear on people who do that, include an invisible link in your html and forbid access to that URL in your robots.txt, then block every IP who accesses that URL for X amount of time. Don't do this if you actually rely on search engine traffic though. Google may get pissed and send you lots of angry mail like "There's a problem with your site".
Ah, but of course you would exclude Google's published crawler IPs from this restriction, because that is exactly what they want you to do.
Re: Only Google is really allowed to crawl the web
#276Earlier quoted context omitted.
Googlebot is pretty careful and generally doesn’t cause these problems.
Right, then they shouldn't be effected by the rate-limiting, as long as its reasonable. If it was applied evenly to all clients/crawlers, it'd at least allow the possibility for a respectful, well designed crawler to compete.
Re: Only Google is really allowed to crawl the web
#277Earlier quoted context omitted.
Google was/is also the largest sponsor of Mozilla. This doesn't stop Google from sabotaging Mozilla. 2 mln is probably Google's hourly profit. For that they get one of the biggest knowledge bases in the world. It's basically free as far as Google is concerned. The instant Google becomes confident they can supplant Wikipedia, they will.
Not sure why you're being downvoted; I completely agree with what you're saying (modulo questionable usage of "sponsor"). If Wikipedia were to try to charge for this use of their data, Google would likely make it a priority to drop the Wikipedia blurbs, either without replacement, or with data sourced elsewhere.
That's an odd way of phrasing things. If Wikipedia were to take away free access to their data, Google wouldn't be dropping Wikipedia, Wikipedia would be dropping Google. This line of thinking "you took this when I was giving it away for free, but now I want to charge for it, so you are expected to keep paying for it" is incorrect.
Re: Only Google is really allowed to crawl the web
#278I can't really trust a website that spells its own name wrong on their homepage. "Knucklesheads’ Club" Edit: https://imgur.com/a/inqYrjV
Re: Only Google is really allowed to crawl the web
#279I tried to set up YaCy [1] at home to index a few of may favorite smaller websites, so I could quickly search just them. That turned out to be a bad idea. Some ended up blocking my home IP address and others reported me to my ISP. None of these sites were that large, and I wasn't continuously crawling them... [1] https://yacy.net/
Re: Only Google is really allowed to crawl the web
#280dupe/posted earlier etc I also got confused about this page as there's another project of theirs around right now about RIP Google Reader that's on a seperate domain... Funny a site that's all about google this and that doesn't have clear URL/pages for their articles that can be linked to easily geez Original post/discussion from the source, 3 months ago: https://news.ycombinator.com/item?id=25417067
That seems like an easy link?