Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

191–200 of 365 posts

Re: Only Google is really allowed to crawl the web

#191

Can we take a moment to talk about this club's business model? There's not even any information to see what the "private forum access" that you have to pay for is about, what kind of people are in it...or even to know about what happens with the money. For me, this sounds like a scam. I mean, no information about any company. No imprint. No privacy policy. No non-profit organization. And just a copy/paste wordpress i…

Not being set up as a 527 nonprofit[0] is the biggest red flag - no donation or membership money has to be spent for political purposes. They also use memberful for their membership/payment system, which doesn't require owning a business, so you might be paying out to the owner directly instead of to a business with its own bank account. Maybe the owner is looking at HN and can clarify.

To add, there are a lot of businesses that use the terms 'Knucklehead' so finding their business on secretary of state business searches might be impossible.

0: https://www.irs.gov/charities-non-profits/political-organiza...

Re: Only Google is really allowed to crawl the web

#192
post #57

Earlier quoted context omitted.

Google doesn’t give anyone access to said cache. I mean one crawler with a shared api among competitors. So exactly the same as the public cache, but run my a private company and accessed for a small fee

You can API google search results to make a meta-search engine if you want to but it's like $5 / 1k requests.

Google's TOS prevents blending (alterations, etc.) though.

Re: Only Google is really allowed to crawl the web

#193

Earlier quoted context omitted.

Perhaps there could be some kind of 'Crawler consortium'? Under this consortium, website owners would be allowed to either allow all crawlers (approved by the consortium) or none at all (that is, none that is in the consortium, i.e. you could allow a specific researcher or something to crawl your website on a case-by-case basis). This consortium would be composed of the search engines (Google, MS, other industry memb…

> Perhaps there could be some kind of 'Crawler consortium'? An industry-wide agreement not to compete for commercially valuable access to suppliers of data? Comprised of companies that are current (and in some cases perennial) focusses of antitrust attention? I think there might be a problem with that plan.

I don't see the problem. If a bunch of non-google companies pooled resources to make a crawl, that would reduce market concentration, not increase it.

Re: Only Google is really allowed to crawl the web

#194

Earlier quoted context omitted.

Ignoring robots.txt is trivial, that's why some(many?) sites enforce it by verifying source IP and recognize Googlebot from its IP addresses - how will you get access to one of those?

What does "recognize Googlebot from its IP addresses" mean? If I'm a human and I access a site, I have some other IP than Googlebot, how should this side know if I'm a human or knuckleheadsbot?

https://developers.google.com/search/docs/advanced/crawling/...

Re: Only Google is really allowed to crawl the web

#195
post #180

I tried to set up YaCy [1] at home to index a few of may favorite smaller websites, so I could quickly search just them. That turned out to be a bad idea. Some ended up blocking my home IP address and others reported me to my ISP. None of these sites were that large, and I wasn't continuously crawling them... [1] https://yacy.net/

How often were you searching?

I was regularly searching, but I was rarely indexing any of the sites. I struggled to even get an initial index of many of the sites, due to being blocked or being reported.

Re: Only Google is really allowed to crawl the web

#197
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

They're making it easier to search for flights and arrange a trip. It's UX and makes me not hate the airlines/travel process as much. And I end up buying the flight from the airline anyways, and in many cases doing the arranging on the airline site in the end once it's determined, so Google is giving that back. They're not taking stuff from the airlines, I mean what ads and stuff are on the airline sites anyways specifically during the search process. Where they are taking away is from the Expedia's and other aggregation sites that offer a garbage/hodgepodge experience that drives people crazy.

Re: Only Google is really allowed to crawl the web

#198
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

The way to look at this from Google’s point of view is to realise that most websites are slow and bad[1], so if Google sent you there you would have a bad experience with a bad slow website trying to find the information you want. Google want to make it better for you. [1] it feels like Google have contributed a lot to websites being slow and bad with eg ads, amp, angular, and probably more things for the other 25 le…

> Google want to make it better for you.

Hehe, sure, nothing nefarious or greedy here... move along, move along, nothing to see...

Re: Only Google is really allowed to crawl the web

#199

Earlier quoted context omitted.

Wikimedia recently announced Wikimedia Enterprise for "organizations that want to repurpose Wikimedia content in other contexts, providing data services at a large scale". So they're pretty clearly looking to monetize organizations which consume their data in a for-profit context.

monetizing != for-profit You could e.g. just cover operational cost and/or improve the service quality from it.

I think they may have meant "(organizations) (which consume their data in a for-profit context)."

Re: Only Google is really allowed to crawl the web

#200

Earlier quoted context omitted.

I'm generally anti business. But I have to disagree. "The Public" that the government serves includes businesses. Businesses (ignoring corporate personhood bullshit) are owned and operated by people. I do not want the government deciding "what purposes" e.g. non-commercial, serve the public good. The public gets to decide that. (charging a license for commercial use is maybe ok (assuming supporting that use costs gov…

> I do not want the government deciding "what purposes" e.g. non-commercial, serve the public good. The public gets to decide that. the public's "decision" on things like this is made manifest by government policy, no?

In theory. In practice, is every single policy that our government upholds currently popular with the majority of people?

It's possible to have government policies that the majority of people disagree with, that remain for complicated reasons related to apathy, lobbying, party ideology, or just because those issues get drowned out by more important debates.

Government is an extension of the will of the people, but the farther out that extension gets, the more divorced from the will of the people it's possible to be. That's not to say that businesses are immune from that effect either -- there are markets where the majority of people participating in them aren't happy with what the market is offering. All of these systems are abstractions, they're ways of trying to get closer to public will, and they're all imperfect. But government is particularly abstracted, especially because the US is not a direct democracy.

I'm personally of the opinion that this discussion is moot, because I think that people have a fundamental Right to Delegate[0], and I include web scraping public content under that right. But ignoring that, because not everyone agrees with me that delegation is right, allowing the government to unilaterally rule on who isn't allowed to access public information is still particularly susceptible to abuse above and beyond what the market is capable of.

[0]: https://anewdigitalmanifesto.com/#right-to-delegate

Post reply on HN