Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

301–310 of 365 posts

Re: Only Google is really allowed to crawl the web

#301
post #209

Earlier quoted context omitted.

> Wikipedia isn't monetized. No, but they often ask for donations when you visit the site, which people won't see if they just see the in-line blurb from Wikipedia on the Google results page. > In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. $2M is a pittance compared to what I expect Google believes is the value of their Wikipedia blurbs. If Wikipedia could charge for use of this data (which…

Unlikely that Wikipedia will be able to charge for content, seeing as all of their content is CC-BY-SA licensed. https://en.wikipedia.org/wiki/Wikipedia:Licensing_update They may be able to charge for bandwidth (if you want to use a Wikipedia image, you can use Wikipedia's enterprise CDN instead of their own), but their licensing allows me to rehost content as long as I follow the attribution & sublicensing terms. Go…

Fine, the content is free. But if your crawlers want access to the content, then pay! Simple as that.

Re: Only Google is really allowed to crawl the web

#302
post #253
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

Does the concergie of a hotel take anything away when he informs you that your flight has been delayed?

It's hard for me to make that an apt analogy. She's not well known as a portal to find websites, which is what Google had been for most of its existence.

It's pretty difficult to come up with a non-computer analogy for how Google works now. Pick a different space, and the power imbalance is quite clear. If they wanted, they could destroy StackExchange very quickly with these widgets.

Re: Only Google is really allowed to crawl the web

#303
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

> You see this recently with Wikipedia. Google's widgets have been reducing traffic to Wikipedia pretty dramatically. Enough so that Wikipedia is now pushing back with a product that the Googles of the world will have to pay for.

Do you have a link of that product/service from Wikipedia?

Re: Only Google is really allowed to crawl the web

#304

Earlier quoted context omitted.

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

Even before it gets to that point, they routinely display snippets off regular websites and show ads next to it. Keeping users from clicking through to organic results helps them generate more revenue.

Here’a a thought: most companies don’t actually want to serve a ton of extra pages. For example, airlines just want to fly passengers. They don’t care who puts those butts in seats and they would fully acknowledge that they aren’t able to deliver a better flight search than Google can. I mean, sure, some small team of web developers at every airline is pissed, but the CEO needs butts in seats to keep the pilot and service unions off her back. Her own web dev team is the least of her problems.

Re: Only Google is really allowed to crawl the web

#305
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

> The airline web site could then present things Google can't do. Like "hey, we see you haven't checked in yet" or "TSA wait times are longer than usual" or "We have a more-legroom seat upgrade if you want it".

If I'm a passenger, there's plenty of ways for airlines to notify me. If I'm searching for a flight status online, it's because I'm picking someone up. If I want more information, I'll click through. I don't see how either me or the airline are hurt.

> You see this recently with Wikipedia. Google's widgets have been reducing traffic to Wikipedia pretty dramatically. Enough so that Wikipedia is now pushing back with a product that the Googles of the world will have to pay for.

Why is that even a problem for Wikipedia?

Re: Only Google is really allowed to crawl the web

#306

Earlier quoted context omitted.

Google was/is also the largest sponsor of Mozilla. This doesn't stop Google from sabotaging Mozilla. 2 mln is probably Google's hourly profit. For that they get one of the biggest knowledge bases in the world. It's basically free as far as Google is concerned. The instant Google becomes confident they can supplant Wikipedia, they will.

> Google was/is also the largest sponsor of Mozilla. This doesn't stop Google from sabotaging Mozilla. Google isn't a sponsor of Mozilla, they're a customer. Do people think Google is "sponsoring" Apple with $1.5 billion a year too?

$1.5 billion a year? You're off by an order of magnitude; the number is thought to be over $10 billion a year.

Re: Only Google is really allowed to crawl the web

#307
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

> You see this recently with Wikipedia. Google's widgets have been reducing traffic to Wikipedia pretty dramatically. Enough so that Wikipedia is now pushing back with a product that the Googles of the world will have to pay for. Do you have a link of that product/service from Wikipedia?

https://diff.wikimedia.org/2021/03/16/introducing-the-wikime...

Discussed not long ago: https://news.ycombinator.com/item?id=26484080

Re: Only Google is really allowed to crawl the web

#308
post #68
post #61

Earlier quoted context omitted.

Only if Google stays around long term. I wouldn't be surprised if each free product on its graveyard took down a dozen of competing products before it was killed of.

Then someone can start a competitor up again, right? Assuming there's actually a market for it.

[deleted]

Re: Only Google is really allowed to crawl the web

#309
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

> And I don't think Google will realize what the problem is until they start accidentally killing off large swaths of the actual sources of this content by taking the audience away. What makes you think they care? Killing off the sources of content might even be there goal. If they kill off sources of content, they'd be more than happy to create an easier-to-datamine replacement. Hypothetically, if they killed off wi…

I'm not suggesting it's illegal. There are a great many practices that are legal that I dislike.

Re: Only Google is really allowed to crawl the web

#310
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

I think the flight arrivals/departures is a bad example. A good example might be putting flights.google.com on the first page or even allow it to exist.

Not sure what you mean. Both for flight status as well as flight shopping, Google drops a huge widget at the top and pushes everything else down, below the visible fold.
Post reply on HN