Earlier quoted context omitted.
Interesting that the most comments it got before was 11, and today it succeeds and makes it to the front page! This is a good illustration of whether or not submissions get any traction can be fairly stochastic. On topic, stack overflow does exactly what the article is talking about; They lock down their sitemap and make special exceptions for the Google bot: https://meta.stackexchange.com/a/98087 https://meta.stacke…
I think it's partly because they create a website which reported on the status of the Ever Given which rose to 1. on the front page. I feel like I often see submissions which are, even tangentially, related to front page material rise very quickly. Regardless, congrats to Knuckleheads Club for fighting the good fight.
Only Google is really allowed to crawl the web
261–270 of 365 posts
Re: Only Google is really allowed to crawl the web
#262Maybe a naïve question but what prevents Knuckleheads’ from ignoring the robots.txt and crawl the side anyway? And if it's so easy to do, how does Google have a monopoly on crawling then?
Re: Only Google is really allowed to crawl the web
#263Earlier quoted context omitted.
So, one more reason to hate Cloudflare and every single website that uses it.
Or maybe don’t “hate” folks who are just trying to put some content online and don’t want to deal with botnets taking down their work? You know, like what the internet was intended for.
Re: Only Google is really allowed to crawl the web
#264https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…
Broadly speaking, robots.txt files are often ignored. I used to run a fairly large job ad scraping organization, and we would be hired by companies (700 of the fortune 1000 used us) to scrape the job ads from their career pages, and then post those jobs on job boards. 99 of 100 times, the robots file would disallow us to scrape. Since we were being paid by that company's HR team to scrape, we just ignored it because…
If you wanna go nuclear on people who do that, include an invisible link in your html and forbid access to that URL in your robots.txt, then block every IP who accesses that URL for X amount of time.
Don't do this if you actually rely on search engine traffic though. Google may get pissed and send you lots of angry mail like "There's a problem with your site".
Re: Only Google is really allowed to crawl the web
#265The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
I swear something like 50% of those digests are totally incorrect as well. It's amazing they have kept the feature because it has never had a very high signal-to-noise ratio. I never trust what's presented in these digests without double-checking the source page.
(Search who invented running)
Re: Only Google is really allowed to crawl the web
#266Earlier quoted context omitted.
Not sure why you're being downvoted; I completely agree with what you're saying (modulo questionable usage of "sponsor"). If Wikipedia were to try to charge for this use of their data, Google would likely make it a priority to drop the Wikipedia blurbs, either without replacement, or with data sourced elsewhere.
Given the scale that google already operates at, I don't doubt that they would just take a copy of thr content and rebrand it as a google service, complete with user contribution. Then, after two or five years, let it fester then abandon it. Nobody gets promoted for keeping well oiled machines running.
It was actually good for writing stuff when I tried it. Never brought in enough traffic. Killed.
Re: Only Google is really allowed to crawl the web
#267Earlier quoted context omitted.
BA had some tracking request inline on the “payment processing” page which when blocked by my pihole prevents me from ever getting to the confirmation page, just have to refresh your email and wait for the best. I have no idea how these companies, which make quite a decent amount of money at least up until 2020, can have such utterly poor sites. I once counted some 20+ redirects on a single request during this proces…
I don’t know what they’re doing but most every single sign on tool I’ve seen redirects 10-20 times during the sign on process (and then dumps you to the homepage to navigate your way back).
Re: Only Google is really allowed to crawl the web
#268The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
And if some of the sources somehow die? New sources will spring up. It doesn't matter.
Re: Only Google is really allowed to crawl the web
#269Earlier quoted context omitted.
Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…
>> Google donated $2M [1]. In 2010, Google also donated $2m [2]. $2 Million a year? Now I know why Googlers complained about having one less olive in their lunch salad. How much does Google PROFIT from Wikipedia and how much does Wikipedia loses in fundraising when Google fails to send users to the info provider?
And most of the value of wikipedia is created by its unpaid users, not Wikimedia foundation.
Re: Only Google is really allowed to crawl the web
#270Earlier quoted context omitted.
I don't think google muscling out intermediaries like Expedia is a good thing. Just for example, Expedia is probably 5% of Google's total revenue and Google doesn't like slim margin services by and large that can't be automated. Travel is fairly high-touch - people centric. It doesn't fit Google's "MO". But... its shitty that google can play all sides of the markets while holding people ransom to mass sums of money t…
>In essence, you're advocating that eBay goes away because google could do it... they could.. and eBay is technically just an intermediary, but do we want everything to be googlefied? I don't think I'm really advocating for it as much as I see as a more or less neutral change. That said, I'm pretty ambivalent about Google. Their size is a concern, but they also tend to be pretty low on the dark pattern nonsense. eBay…
Google has a huge pay 2 play problem with PPC... i've worked for Expedia so that's the only reason i know this :)
It's the reason companies work with Expedia many times because they don't have the leverage expedia group does...
i see it as unnatural change btw... "borg" if you will.