Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

71–80 of 365 posts

Re: Only Google is really allowed to crawl the web

#71
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia??

And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2].

[1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli...

[2] https://en.wikipedia.org/wiki/Wikimedia_Foundation

Re: Only Google is really allowed to crawl the web

#72
Google makes $150+ billion from Google Search per year. Running Google Search could be operated for likely (much less than) $10 billion per year.

So, Google is in effect taxing us all $140 billion per year.

It's not dissimilar from how Wall Street effectively taxes us all for an even larger amount.

In both cases, we could use some kind of non-profit open system to facilitate web search and stock trading.

The Great Lie that Google is doing a good thing by charging money to insert "relevant ads" above the search results is totally wrong. If those ads are the most relevant, they should just be the top organic results, obviously.

Google mostly solved search 20 years ago. There's really nothing that impressive about Google Search in 2021. It should be relatively easy to replace it with something open, leveraging the massive improvements in hardware and software. It could operate like Wikipedia or Archive.org. The hard part is probably getting the right team and funding assembled.

Re: Only Google is really allowed to crawl the web

#73
I just don't see this working out legally. How would it even work?

From the "learn more"

> Sometime soon we will be publishing what we think should happen and what we think will happen. These two futures diverge and we believe that, while the gap between them exists, it will entrench Google’s control over the internet further. We believe that nothing short of socialization of these resources will work to remove Google’s control over the internet. Our hope is that in publishing this work right now we will let the genie out of the bottle and start a process towards socialization that cannot be undone.

Sorry, but I deeply skeptical of this. This sounds like the first step towards a non-free internet. At the end of the day, it is your box on the web, and if you want or don't want someone/something to crawl it, that is your call to make.

Re: Only Google is really allowed to crawl the web

#74
I have an idea: remove the art of web crawling from the domain of a single company and instead create a international group of interested parties to run it instead. I'm thinking broadly along the lines of the Bluetooth SIG. Maybe it will be a bit more complicated, and require international political efforts, but it will make the search engine market way more democratic.

Re: Only Google is really allowed to crawl the web

#75
post #63
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

I swear something like 50% of those digests are totally incorrect as well. It's amazing they have kept the feature because it has never had a very high signal-to-noise ratio. I never trust what's presented in these digests without double-checking the source page.

I remember when rich snippets (one type of those widgets) came out there were a lot of funny examples. One for a common query about cancer treatments that pulled data from a dodgy holistic site saying that "carrots cured most types of cancer" (or something like that).

There was a similar one where Google emphatically claimed a US quarter was worth five cents in a pretty and large snippet graphic.

Re: Only Google is really allowed to crawl the web

#76

Earlier quoted context omitted.

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

Aren't there anti trust laws to prevent this kind of thing?

That depends, would Google let us know?

Re: Only Google is really allowed to crawl the web

#77

Earlier quoted context omitted.

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

Aren't there anti trust laws to prevent this kind of thing?

Yes, but they lack enforcement.

Re: Only Google is really allowed to crawl the web

#78
post #5

The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…

I'd have to look more but maybe running a cache isn't dead simple. I can imagine that the benefits of manipulating what's in the cache either adding or removing would be very high. Google and the others are private companies so they're not required to do everything in the public view.

A public cache wouldn't be able - indeed shouldn't - to play cat and mouse games with potential opponents. I suspect most of the games played require not explaining exactly what you're doing.

Re: Only Google is really allowed to crawl the web

#79
> Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission.

This was eyebrow-raising. Actually seems like this is not (any longer?) true:

https://census.gov/robots.txt:

User-agent: *

User-agent: W3C-checklink

Disallow: /cgi-bin/

Disallow: /libs/

...

That first line wildcards for any user agent but does nothing with it. It should say "Disallow /" on the next line if it blocked all unnamed robots. It looks like someone found out about it and told the operators, rightfully so, that government webpages with public information (especially the census) shouldn't have such restrictions. They then removed only the second line and left the first. Leaving the first line has no impact on the meaning of the file.

Re: Only Google is really allowed to crawl the web

#80
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

Be careful when using Google Flight, last time I checked they use significantly less margins between flights so it’s shorter trips but much riskier.
Post reply on HN