Live data from Hacker News

Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

docs.house.gov

201–210 of 271 posts

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#201
post #119
post #41

The primary case here made against Google might be questionable behavior from Google, but might not be illegal, since simple numbers like a dollar amount can’t be copyrighted, and AFAIK there is no other legal impediment to Google doing what they did (IANAL). What looks far more inculpatory to me is this, on the last page: […] Because it controls essentially the entire internet, Google has endless levers at its dispo…

> This is far more damning of Google abusing its monopoly, not only on search, but also on the Android platform. No, it’s really not. I just visited celebritynetworth.com. One article took 462 http requests, loaded 9mb of content and though DOM content loaded in 434ms, it took 1 minute to finish loading everything. The sheer amount of adware on this site is staggering. Let’s put down the pitchforks and put on our cri…

Maybe if they served up some of Google's AMP™ pages they would get higher search rankings.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#202

Earlier quoted context omitted.

actually this is interesting right here, because if Google takes CNW's fake celebrity, then that is definite copyright infringement - the fake celebrity was created by CNW and as such is not a fact so Google can't argue that it's not copyrighted because facts are not copyrighted.

That's not how the courts view these issues. Fake facts hidden within real facts have been litigated many times and has always been shown to be uncopyrightable. See https://en.m.wikipedia.org/wiki/Trap_street#Legal_issues .

You are correct, about Trap Streets. I'm not sure if the language is fake facts hidden in real facts, or fake streets in maps. Because it seems pretty specific to trap steets.

If the language is indeed not general beyond trap streets - if I create a fake celebrity and data regarding them, I would argue that is more like making a story. A trap street is a small line on the map (generally hidden somewhere unimportant) and a couple words. A fake celebrity requires much more content.

From the PDF "I added fivecompletely conjured celebrities to the site. I used stock photos with entirely made up names, biographies and net worth numbers. I published the pages backdated several yearsso they were nearly invisible to the world."

Now, let's look at a real celebrity on their site https://www.celebritynetworth.com/richest-celebrities/actors...

It seems reasonable to assume from the pdf that they put the kind of effort into an article on the fake celebrities to make them indistinguishable from real celebrities, it may be decided that this is the same as a trap street. But I don't think it is just a foregone conclusion.

Aside from all that - https://www.theguardian.com/uk/2001/mar/06/andrewclark they might be able to sue in some other place than the U.S and find a more sympathetic court.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#203

> And what about those conjured celebrities I added as a precaution? All five were all scraped right into Google’s search result pages. It provided undeniable proof that after being turned down, Google simply went ahead and stole the entire database of content CNW took eight years and over a million dollars to build. I remember a few years ago Google had a big blog post about how they’d injected some fake search resu…

I worked at Microsoft from 2010-2012, and the crux of the issue there was that, for Microsoft users who had the Bing toolbar installed, the toolbar would track what you were visiting in your browser and use that session/journey information to try and improve the relevancy of Bing results. If a lot of people who searched for "best pancake house" (regardless of what search engine you used) end up visiting ihop.com within the same browser session, well then Bing would want to rank ihop.com higher for that search phrase.

Google noticed this, and for unique/low-traffic search terms, was able to synthetically generate enough "fake" traffic that the "fake" traffic became the dominating signal for those terms, and therefore Bing started directing users to the fake results.

The bad thing here IMO is the level of tracking of users via this toolbar, but fundamentally this seems as bad as any other digital fingerprinting or advertiser tracking as anything else that's become common on the web. This is not to excuse Microsoft for doing a bad thing, but it really had very little to do with "scraping Google", which somehow became the popular media takeaway for this.

EDIT: a decent contemporaneous article in Wired: https://www.wired.com/2011/02/bing-copies-google/, mentioning how Microsoft was using the clickstream data from the browser/toolbar, not scraping Google results per se

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#204
post #171

Earlier quoted context omitted.

This isn't theft of data or copyright infringement because you can't copyright facts in the US.

These aren't facts though. They are researched estimates. Different researchers will come up with different numbers.

Is the fact “CNW has estimated that celebrity A has a net worth of $X” copyrightable? I sincerely doubt it.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#206

Google is incredibly shady. There’s been a manual penalty in place against my website The Online Slang Dictionary ( http://onlineslangdictionary.com ) for the better part of a decade. I know this because the data suggested it - and then a Google employee confirmed it. It would be easy to explain it away because Aaron Peckham, owner of Urban Dictionary, worked at Google when Matt Cutts was head of the Web Spam team, a…

> I do know that when I confronted Cutts about the penalty here on HN he lied about it

That's an extreme thing to say and crosses into personal attack. The thread you seem to be referring to (https://news.ycombinator.com/item?id=5418864) doesn't seem to support what you've said here. At a minimum, you're breaking the site guideline which says: "Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith." https://news.ycombinator.com/newsguidelines.html

I'm not saying your claims about Google and your site are false—I have no idea; it sounds like you have every right to be frustrated. Also, I don't know Matt. But these sorts of accusations ought not to be slung carelessly on HN. Since we're trying to have a site that staves off internet-default outcomes, we all need to be careful about this.

I have experience with these dynamics in the HN context and can tell you that people jump to conclusions all the time about why they were penalized or banned. Their conclusions are almost always wrong and overly dramatized. Frequently they take to the forum to declaim about how badly they were treated, and they always put it in strikingly specific and factual-sounding terms, as if they know for sure. But they don't know for sure; in fact they don't know at all. They've just completely made it up.

The HN context is simpler and has lower stakes than Google. (At least here, people can get factual answers to specific questions if they ask.) But that only strengthens the point: these domains are complicated and the dramatic explanations that people come up with out of frustration are almost always wrong. Of course it's frustrating is Google has penalized your site, but please don't let that boil over into accusations against specific people unless you can strictly demonstrate what you're saying.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#208
post #119

Earlier quoted context omitted.

> This is far more damning of Google abusing its monopoly, not only on search, but also on the Android platform. No, it’s really not. I just visited celebritynetworth.com. One article took 462 http requests, loaded 9mb of content and though DOM content loaded in 434ms, it took 1 minute to finish loading everything. The sheer amount of adware on this site is staggering. Let’s put down the pitchforks and put on our cri…

Maybe if they served up some of Google's AMP™ pages they would get higher search rankings.

I know you are being sarcastic but it would result in a 10x better experience for the user.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#209
post #119
post #41

The primary case here made against Google might be questionable behavior from Google, but might not be illegal, since simple numbers like a dollar amount can’t be copyrighted, and AFAIK there is no other legal impediment to Google doing what they did (IANAL). What looks far more inculpatory to me is this, on the last page: […] Because it controls essentially the entire internet, Google has endless levers at its dispo…

> This is far more damning of Google abusing its monopoly, not only on search, but also on the Android platform. No, it’s really not. I just visited celebritynetworth.com. One article took 462 http requests, loaded 9mb of content and though DOM content loaded in 434ms, it took 1 minute to finish loading everything. The sheer amount of adware on this site is staggering. Let’s put down the pitchforks and put on our cri…

That's not the issue OP is addressing. If Google is ranking them lower because it's a slow, crappy, ad-ridden site, that's understandable. It's stealing their data and presenting it as their own that's the problem here.

A site being badly designed does not mean they're suddenly not protected by copyright.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#210
post #41

The primary case here made against Google might be questionable behavior from Google, but might not be illegal, since simple numbers like a dollar amount can’t be copyrighted, and AFAIK there is no other legal impediment to Google doing what they did (IANAL). What looks far more inculpatory to me is this, on the last page: […] Because it controls essentially the entire internet, Google has endless levers at its dispo…

Just because it wasn't illegal (or questionably legal) yesterday doesn't mean we can't make it illegal tomorrow.

I would surmise that the most popular laws ever written by humanity, collectively, started with a bunch of people thinking some version of "That behavior is questionable."

So sure, the last part is bad too. But that statement was being provide to people who write new laws. Not quite understanding why the current legality of Google's copying is that significant in that context.

Post reply on HN