Live data from Hacker News

Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

docs.house.gov

151–160 of 271 posts

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#151

Earlier quoted context omitted.

Right on, Jeff8. And when Teddyh says "simple numbers like a dollar amount can’t be copyrighted" he's ignoring the fact that google isn't just copying dollar amounts. Google's copying a data set which establishes a dollar amount as a net worth of a particular person. That's a database. The only difficulty for CNW will be coming up with the money, time, and emotional energy to achieve a result in court over as many ye…

Nobody forces you to (checks notes) create celebrity net worth databases. I just feel like we live in a time with more information than ever, the problem clearly being inaccurate information or propaganda. The problem isn’t insufficient marginal knowledge. Obviously we live in this high knowledge world because of, not despite, Google, so it’s hard to see if you’d really be on the right side of justice if you’re again…

How is Google "Big Knowledge" when the actively play a role in deciding what knowledge gets disseminated or buried? You said yourself the problem is inaccurate information or propaganda which Google is a part of if they're suppressing information they don't like.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#152
post #138
post #112

Earlier quoted context omitted.

Recently one of the lyrics sites had a similar complaint that might be the same thing happening in this case (and maybe it was the same issue for Bing's case too). If you specify that you don't want Google to take snippets, but a third party ignores the robots.txt or meta tags and has a permissive one of its own Google can just scrape it from there. A missing detail here is what did the textbox for the fake celebriti…

Playing around with real celebrities: * https://www.google.com/search?q=larry+david+net+worth -> featured snippet from Wikipedia, showing an estimate * https://www.google.com/search?q=bill+gates+net+worth -> database style response, showing answer with no link

[deleted]

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#153
post #85

Earlier quoted context omitted.

> Taking someone else's copy written content […] (You should read up a bit on this; it’s obvious you have not since you call it “copy written” when the term is “copyright” and a work can therefore be “copyrighted”. The term is about who has the right to copy a given work, and has nothing to do with “writing”.) Copying text someone else has written is indeed copyright infringement (commonly, but technically inaccurate…

>(You should read up a bit on this; it’s obvious you have not since you call it “copy written” when the term is “copyright” and a work can therefore be “copyrighted”. The term is about who has the right to copy a given work, and has nothing to do with “writing”.) Perhaps the parent comment was from some country where things are different, like the language.

Perhaps, but this case was explicitly about US copyright, and may even hinge on the specific intricacies on whether database copyright in US law covers this specific case.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#154
post #39

Earlier quoted context omitted.

IANAL: 1. Google doesn't show in-line'd data box with Net worth 2. Google asks site (CNW) for access, is denied. 3. Google then shows site's data in data box (with CNW site proof of theft with canary tokens being the fake celebrity's net worth) that are now presented as Google's data (no source or fake source), without license from original site. If that isn't theft of data, what is it? From the PDFs, there is no lic…

actually this is interesting right here, because if Google takes CNW's fake celebrity, then that is definite copyright infringement - the fake celebrity was created by CNW and as such is not a fact so Google can't argue that it's not copyrighted because facts are not copyrighted.

That's not how the courts view these issues. Fake facts hidden within real facts have been litigated many times and has always been shown to be uncopyrightable. See https://en.m.wikipedia.org/wiki/Trap_street#Legal_issues.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#155
post #39

Earlier quoted context omitted.

>> Now it's clearly IPR theft, present it how you like, but theft is theft. I don't quite see the legal basis for it being theft. Can you explain?

IANAL: 1. Google doesn't show in-line'd data box with Net worth 2. Google asks site (CNW) for access, is denied. 3. Google then shows site's data in data box (with CNW site proof of theft with canary tokens being the fake celebrity's net worth) that are now presented as Google's data (no source or fake source), without license from original site. If that isn't theft of data, what is it? From the PDFs, there is no lic…

This isn't theft of data or copyright infringement because you can't copyright facts in the US.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#156

Earlier quoted context omitted.

Very true. Due to the penalty, the site gets less and less traffic every year. Enabling SSL/TLS might help a bit, but I have to face facts: the site is dead.

Not to belittle your experience, but it seems this may be a bit of a negative feedback loop. Haven't updated the homepage because "what's the point I'm already punished" etc...

I spent 80 hours a week from when the penalties started in 2011 until I went broke and got outside employment in 2014. I’ve worked on it for years worth of time since then, including updating the home page. (It’s a content site. Updating the home page is a thing, but people come to the site via searches for slang terms, not for the home page...)

There is nothing I can do with the manual penalty in place, and it doesn’t make sense to throw good time after bad. It’s been 9 years since the penalties started and there’s been no improvement since then. Not only is it not healthy to continue banging my head against the wall of Google, since we’re on a site for a startup incubator: it doesn’t make monetary sense.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#157
post #148

Earlier quoted context omitted.

The mere fact that he compiled the database isn't sufficient to qualify for copyright protection. He'll have to show that he used a sufficient level of creativity to arrive at the numbers he displayed. Based on his testimony, I suspect he'd have a reasonable case to argue, but he'd have to show that his researchers carefully weighed different sources and made other judgements for each individual celebrity rather than…

The article says they employed a number of researchers to come up with the numbers. I don't think property records (or any other publicly available information) would be enough to come up with an accurate figure in an automated fashion. So, there's some educated guesses (aka "informed estimations") going on and therefore a form of "creativity" (IMHO).

I agree if the testimony is taken at face value, it makes a pretty reasonable case. However, there's nothing to say that the values he arrived at were accurate (not that it'd necessarily be material to the case) or that his 'researchers' were not just some content farm to write up articles around those numbers in order to enhance SEO. These are the kind of things that'd be deliberated upon in a court case.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#158
post #93
post #41

The primary case here made against Google might be questionable behavior from Google, but might not be illegal, since simple numbers like a dollar amount can’t be copyrighted, and AFAIK there is no other legal impediment to Google doing what they did (IANAL). What looks far more inculpatory to me is this, on the last page: […] Because it controls essentially the entire internet, Google has endless levers at its dispo…

Given enough time tinkering with the prompt, and assuming the Wired article and website were in the training data, I think GPT-3 could probably suss out the relationship between the critical comments in Wired and the net worth website. It clearly can identify sentiment and maybe from that even infer the strained relationship of the site to Google Seems plausible that Search could autonomously and possibly without hum…

Just tried this through AI Dungeon to see what happens:

Prompt

You are Bob Warner, founder of CelebrityNetWorth.com, and are speaking at a hearing before the House Judiciary Subcommittee on Antitrust, Commercial and Administrative Law. The subject is Google. You start your testimony as follows:

"Thank you, Chairman Cicilline, Ranking Member Sensenbrenner, Committee Chair Nadler and Ranking Member Collins,

GPT-3 Response

for inviting me to testify today about Google's anticompetitive conduct in the search market. I am honored that my company, CelebrityNetWorth.com, has been chosen by the subcommittee to speak on this important issue.

I want to begin by thanking the members of this committee for their interest in our industry. As you know, we are an information technology business that provides Internet marketing services to small businesses and individuals. We have been providing these services since 1996. Our service allows users to post advertisements on the Internet in exchange for money. In addition, we provide other services such as web hosting and e-mail services.

CelebrityNetWorth.com is an information website that provides public information about celebrities. It was founded in 2008 and has been very successful. We do not collect user information; we simply provide the public with information about celebrities. For example, on our website, you can find out how much money Kim Kardashian is worth.

We have thousands of users and track over 3 million individuals. We get a combined 6 million monthly page views. We receive an average of 75,000 unique visitors per month. We are very successful and I thank this subcommittee for the help it has given us over the years.

Our website has had many successes, but our crowning achievement was exposing Google for the shakedown racket that it is. Our story began in 2006 when we tried to enter the search engine market. At that time, Google had a near monopoly in the search engine market.

In 2009, we applied to be listed in the Google search engine as a secondary link. We went through all the proper channels and even paid money to Google. After we paid them, our website's popularity increased substantially.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#159
post #112

> And what about those conjured celebrities I added as a precaution? All five were all scraped right into Google’s search result pages. It provided undeniable proof that after being turned down, Google simply went ahead and stole the entire database of content CNW took eight years and over a million dollars to build. I remember a few years ago Google had a big blog post about how they’d injected some fake search resu…

Recently one of the lyrics sites had a similar complaint that might be the same thing happening in this case (and maybe it was the same issue for Bing's case too). If you specify that you don't want Google to take snippets, but a third party ignores the robots.txt or meta tags and has a permissive one of its own Google can just scrape it from there. A missing detail here is what did the textbox for the fake celebriti…

I believe it was Genius that was using rare unicode characters in the lyrics that don't really show up when viewing the page. They were able to see that Google had those same characters in their lyric results and therefore Google was just scraping Genius.

Re: Statement on Google’s conduct by founder of CelebrityNetWorth.com (2019) [pdf]

#160

Google is incredibly shady. There’s been a manual penalty in place against my website The Online Slang Dictionary ( http://onlineslangdictionary.com ) for the better part of a decade. I know this because the data suggested it - and then a Google employee confirmed it. It would be easy to explain it away because Aaron Peckham, owner of Urban Dictionary, worked at Google when Matt Cutts was head of the Web Spam team, a…

How did you site rank on other search engines? Do you have evidence it outranked UD on yahoo, etc?
Post reply on HN