And yet I constantly get CAPTCHA's just from blocking tracking cookies. The biggest scraper of them all goes to incredible lengths to prevent scraping. How ironic. I hope the US govt demolishes these monopolies. It's not just in web either, the closest historical precedent I can think of is the Robber Barons of the 30's
There’s a thing called a robot.txt.
Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]
211–220 of 271 posts
Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]
#212"Don't be evil" ;-) Is it still Google's motto?
They deprecated that years ago.
It's listed in the code of conduct which is easily found online.
Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]
#213Surely they can find a stronger example than a spammy SEO business that didn't exist before google, couldn't exist without and doesn't exist after? Google sucks now but not finding this compelling
Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]
#214Doesn't this behavior eventually hurt Google? If there's no incentive for third parties to collate, verify, and organize facts...eventually there isn't anyone to scrape. And the facts go stale.
Which (for worse) has been the reason to roll back all the antitrust laws that would have stopped Google from ever becoming powerful enough to do this in the first place.
Google of course argues "faster results" are good for the consumer, but if you strip away the incentives for publishing information in the first place, you quickly end up with "bad results" then "no results."
Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]
#215Doesn't this behavior eventually hurt Google? If there's no incentive for third parties to collate, verify, and organize facts...eventually there isn't anyone to scrape. And the facts go stale.
It would depend on how the websites google are collecting data from operate. For example, somewhere like wikipedia isn't harmed through google's re-use of the facts they collate.
Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]
#216Earlier quoted context omitted.
Your distinction between "measure" and "estimate" seems meaningless to me, and probably to most scientists and engineers. All physical measurements have uncertainty, and many people are engaged as we speak in creative (in the colloquial sense of the word) work seeking to reduce that uncertainty. We tend to say "measurement" when the uncertainty is negligibly small for whatever purpose is at hand, and "estimate" when…
>So does that make a new estimate of the mass of the Earth copyrightable? Likely not on its own (due to fair use). But in a "collection" of estimates for the mass of all known celestial bodies then yes, it absolutely would be copyrightable. The same way a collection of stock price targets produced by a research company are copyrightable. I agree with the rest of your comment in general but it's relevance to the discu…
In any case, large companies with near-infinite legal budgets still don't get to (openly) violate well-settled law, since judges do see what's happening and seek to minimize the burden on their opponents--and even award legal fees where possible, which it probably would be here under the Copyright Act. Their advantage comes when the law is at least slightly unclear and a victory would be worth a lot more dollars to the large company (for the precedent) than to their opponent. I believe that's the situation here, and the reason why CNW chose not to pursue Google for copyright infringement; though even with matched legal teams, I'd still bet that CNW would lose.
Finally, if that mass of the Earth isn't copyrightable, then no concept of fair use exists, because that concept exists only for copyrighted material; so I don't think your statement there makes sense. If a copyright for the CNW dataset did exist (which I doubt, but maybe), then Google could still argue fair use. I'd bet Google would lose that one, since they're using the whole work, without transformation of the original, in a way that destroys the commercial value of the original.
Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]
#217Earlier quoted context omitted.
Maybe if they served up some of Google's AMP™ pages they would get higher search rankings.
I know you are being sarcastic but it would result in a 10x better experience for the user.
Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]
#218Earlier quoted context omitted.
> (You should read up a bit on this; it’s obvious you have not since you call it “copy written” when the term is “copyright” and a work can therefore be “copyrighted”. The term is about who has the right to copy a given work, and has nothing to do with “writing”.) You used someone’s grammar as a reason to piss all over a reasonable comment, merely because it didn’t 100% validate your own? You should read up a bit on…
Maybe you should realize that merely commenting upon someone’s word choices, inferring inexperience, and providing some elucidation, is not the same as “piss all over”. I mean, nowhere did I say anything else negative about anything in mfer’s comment; I did not even disagree with anything they said! Ergo, please take your lecturing about “basic human decency” somewhere else.
Food for thought.
Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]
#219Google is incredibly shady. There’s been a manual penalty in place against my website The Online Slang Dictionary ( http://onlineslangdictionary.com ) for the better part of a decade. I know this because the data suggested it - and then a Google employee confirmed it. It would be easy to explain it away because Aaron Peckham, owner of Urban Dictionary, worked at Google when Matt Cutts was head of the Web Spam team, a…
Here, I've marked up a typical content page of yours, showing what's wrong very blatantly. It took about 30 seconds to diagnose the whole site based on the repeating content farm structures and thin content problems.
https://i.imgur.com/k5brlYr.png
That page maybe has 20 or 30 words of quality content on it. That will get you an extreme demotion from Google if you're running a whole site like that in a content farm structure.
Your site was very likely penalized due to low quality, shallow content. Google tagged it as being a content farm. No conspiracy is necessary, it's obvious why it's penalized. Any SEO expert would point out the countless problems in the first few minutes of a discussion with them.
I looked at numerous pages, all were of very low quality, and thin / shallow on content.
On one page I checked there were 471 words, most of which were repeating low value words or from low value segments of the page (eg the "Link to this slang definition" section). Google knows the page is thin on quality content, aka shallow. I found this to hold true across all the pages I checked.
On most of the pages I checked the most frequent non-common words were "definitions" and "include" - ie hyper low value content. Your typical page looks like it might have 50 or fewer words of quality non-repeating content (content unique to just that page, written by a human). That's hundreds of words shy of what it needs to be at an absolute minimum.
The pages all have repeating content farm structures (which were increasingly common 8-12 years ago) with shallow content in those areas of the page, which is something Google sought to clean up around the time you're claiming the penalty occurred, which makes perfect sense.
You built a content farm, with a lot of shallow content pages, Google dramatically changed how it treats those types of sites, and your site most likely got nuked because of it. If you build a content farm structured site, where all the pages have repeating segments and your offset of positive value only consists of dozens of words of quality text on each page, you stand zero chance of getting traffic. That started being true 8-10 years ago as Google updated its algorithms to target content farms, to wipe them out of existence. They killed a billion dollar site, eHow, circa 2010-2011 in the process:
Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]
#220In June 2019, search engine analyst Rand Fishkin put together a report about Google using data from web analytics firm Jumpshot. The data show that today an estimated 48.96% of all Google searches end with the searcher NOT clicking through to a website. The same report estimates that 7% of all search clicks go to a paid ad result and 12% go to properties owned by Google’s parent company Alphabet. Moreover, those stats do not even show the full extent of the problem because the data largely relied upon desktop devices and could not track searches that took users to a Google-owned app like the YouTube or Google Maps.