Live data from Hacker News

Google doesn't recognise or penalise stolen content

pi-datametrics.com

71–80 of 80 posts

Re: Google doesn't recognise or penalise stolen content

#71
post #66

Earlier quoted context omitted.

So back when Blekko was a consumer search engine we could 100% figure out who owned content on sites we crawled often. And even when we didn't we could often guess correctly more often than not based on the domain registration dates. (not to mention registry owners). That is because few people who rip off content rip off just one web site, they will rip off dozens of web sites and they will all share the same AdSense…

Do you know if any search engine is actively filtering for this?

Sadly no, there are only a few actual derived general search indexes (english language search) in the world, Microsoft's, Google's, Baidu's, and Yandex's. They are expensive to build and maintain and the only way to monetize them requires driving search traffic your way. Google is paying $4B/year to third parties to send search traffic their way.

My guess, having been at both Google and Blekko, is that "whitelisted" search will be the next wave in the industry. For those old enough to remember Yahoo!'s original "directory" model, once Yahoo!'s contract with Microsoft is up one could hope they rebuild their search team and technology into something with a strong editorial bias for "quality" content.

Re: Google doesn't recognise or penalise stolen content

#72
post #23

Earlier quoted context omitted.

For physical property, yes. For intellectual property, it can still be stolen even if the original owner still has a copy. e.g. The Soviet spies stole the plans for the hydrogen bomb.

No, for intellectual property the word "stolen" is not appropriate. (Well, unless you're talking about getting the courts to tell the original owner it's yours instead. Which has been at least attempted a few times.)

It is semantically appropriate for the word steal. I think you're confusing it with the criminal implications.

Re: Google doesn't recognise or penalise stolen content

#73
post #66

Earlier quoted context omitted.

Do you know if any search engine is actively filtering for this?

Sadly no, there are only a few actual derived general search indexes (english language search) in the world, Microsoft's, Google's, Baidu's, and Yandex's. They are expensive to build and maintain and the only way to monetize them requires driving search traffic your way. Google is paying $4B/year to third parties to send search traffic their way. My guess, having been at both Google and Blekko, is that "whitelisted"…

> "whitelisted" search will be the next wave in the industry

> ... strong editorial bias ...

Interesting. Care to elaborate?

When you use words like 'whitelisted' and 'editorial', I imagine humans adding something to a database one by one. But the volume of useful pages (and the number of site) is really large now, so I guess that's not what you mean.

One thing I like about search today is that it's almost comprehensive. If I know something exists on the (open) web, I can usually find it with a few searches, even if it's very recent or obscure. I don't want to go back to the days when I browsed gopher directories, or even to the days when finding good quality content meant a hierarchical journey from a directory to a site, to a site map, to an individual page.

Re: Google doesn't recognise or penalise stolen content

#74
post #66

Earlier quoted context omitted.

Do you know if any search engine is actively filtering for this?

Sadly no, there are only a few actual derived general search indexes (english language search) in the world, Microsoft's, Google's, Baidu's, and Yandex's. They are expensive to build and maintain and the only way to monetize them requires driving search traffic your way. Google is paying $4B/year to third parties to send search traffic their way. My guess, having been at both Google and Blekko, is that "whitelisted"…

Did you get that $4 billion number from that quarterly results? Does that include things like their payments to Opera and Apple? Does it include search rev share deals with entities like AOL and Ask (& soon to be Yahoo)? Does it include paying from Chrome distribution bundled with Flash security updates & such? I have never seen the overall numbers broken down in terms of what percent goes where on the different sorts of syndication deals.

Three things which would be a major issue for Yahoo! on that sort of search would perhaps be first that they themselves rely so heavily on content syndication to power their various verticals, second they keep losing search market share (especially as more search happens on mobile devices and Google has mobile locked down with their Android contracts), and they also screwed up their old directory before they moved it to Yahoo! small business as part of the Alibaba share spinco.

I also don't see how Yahoo would effectively differentiate their search engine enough to be able to (profitably) buy share at prices set by Google, particularly if they over-promote their internal results & rely on a smaller search index.

Re: Google doesn't recognise or penalise stolen content

#75
post #54
post #49

Earlier quoted context omitted.

> the current algorithm's obsession with "freshness." Which is how Google makes blogspam such a good business to be in, even if your content is inferior to the post you used for "research".

Most of the spam I see in the wild these days is indeed (established) dropped domains which were picked up and then loaded with thousands of pages of "fresh" spun content, with an incestuous backlink profile if any. So indeed 'blogspam'. Everything old is new again; it feels just like twelve years ago. Soon people will be keyword stuffing in a font the same color as the background... But Google certainly isn't intend…

"The SERPs are clean these days."

Here's an alternate take on that http://www.johnon.com/1075/bullish-on-seo-rankbrain-vs-seobr...

Re: Google doesn't recognise or penalise stolen content

#76
post #15

Earlier quoted context omitted.

The book has not yet been published , so this is stealing. But once you put up something in the internet, it is officially available to everyone. Doing stuff with public information is fine IMO. Same as analyzing tweet data (tweets are public).

What if the book has been published in paper form? The book is public as long as you pay for it. Since internet isn't free, you're still paying for content. At what point are you paying "enough" that the information isn't public anymore? Are you saying no one should monetize their content using ads unless they're willing to allow anyone else to do that as well?

It is worth pointing out just how pissed off Google engineers were publicly when they felt Bing was copying their search results.

https://googleblog.blogspot.com/2011/02/microsofts-bing-uses...

http://searchengineland.com/google-bing-is-cheating-copying-...

Re: Google doesn't recognise or penalise stolen content

#77
post #54

Earlier quoted context omitted.

Most of the spam I see in the wild these days is indeed (established) dropped domains which were picked up and then loaded with thousands of pages of "fresh" spun content, with an incestuous backlink profile if any. So indeed 'blogspam'. Everything old is new again; it feels just like twelve years ago. Soon people will be keyword stuffing in a font the same color as the background... But Google certainly isn't intend…

"The SERPs are clean these days." Here's an alternate take on that http://www.johnon.com/1075/bullish-on-seo-rankbrain-vs-seobr...

That's fascinating because I really don't understand it. Maybe I'm just out of touch with SEO, but things like this escape me completely:

"This is because SEOs follow and influence the intent of searchers in the marketplace, while Google’s algorithm (and AI) merely monetizes it."

Where does the extra monetization on page 1 results come from? Unless he's implying that Google provides bad search results so that people will click the ads instead.....

Re: Google doesn't recognise or penalise stolen content

#78

Earlier quoted context omitted.

Sadly no, there are only a few actual derived general search indexes (english language search) in the world, Microsoft's, Google's, Baidu's, and Yandex's. They are expensive to build and maintain and the only way to monetize them requires driving search traffic your way. Google is paying $4B/year to third parties to send search traffic their way. My guess, having been at both Google and Blekko, is that "whitelisted"…

Did you get that $4 billion number from that quarterly results? Does that include things like their payments to Opera and Apple? Does it include search rev share deals with entities like AOL and Ask (& soon to be Yahoo)? Does it include paying from Chrome distribution bundled with Flash security updates & such? I have never seen the overall numbers broken down in terms of what percent goes where on the different sort…

Prior to restructuring their reporting, Google reported as a cost paid distribution. I left in 2010, I started tracking the number in Q1 2011. It was $337M for the quarter. by Q4 of 2014 that number had ballooned to $968M for the quarter. In 2015 they changed the way the reported this number making future comparisons problematic.

I expect it does include fees paid to Apple so that Apple would send search traffic to Google, and fees paid to browser vendors.

Our experience as a search results provider was that there was demand for a more 'functional' search capability (not casual searching) many of the techniques we used have been adopted by Microsoft in their Bing engine which has improved both their recall and quality with respect to Google results on highly contested searches.

I certainly agree that Yahoo! has made a number of missteps with their search technology. I talked with them once (post Marissa's arrival) and in many ways they were confused as ever about how search engines generate value for the parent company, but such things are rarely permanent.

Re: Google doesn't recognise or penalise stolen content

#79
post #77

Earlier quoted context omitted.

"The SERPs are clean these days." Here's an alternate take on that http://www.johnon.com/1075/bullish-on-seo-rankbrain-vs-seobr...

That's fascinating because I really don't understand it. Maybe I'm just out of touch with SEO, but things like this escape me completely: "This is because SEOs follow and influence the intent of searchers in the marketplace, while Google’s algorithm (and AI) merely monetizes it." Where does the extra monetization on page 1 results come from? Unless he's implying that Google provides bad search results so that people…

There are numerous ways to interpret that. At a base level, one could look at how the mobile search results are sometimes a screen full of ads, or how in some verticals they are a screen full of ads followed by yet another screen full of ads.

And then there is the knowledge graph & other flavors of scrape-n-displace, which is largely content recycled from elsewhere, given prominent positioning not based on merit or editorial quality, but based on who the publisher (or recycler) is.

Another parallel trend would be the confirmation bias / brand bias factors promoting older and staler sites. Or simplified "take" articles in the mainstream media rather than the original source articles on niche hobbyist blogs and forums or such.

And in taking broad sets of new niche intents and trying to guide those streams of users back down well worn paths. For example, sometimes when you want to find a particular news story about a broad & well-known web platform like Apple, Amazon, Facebook, or Google it can be hard to find sites other than the official site. And on some other longtail queries Google rewrites what is being searched for in a way that brings up some results that don't match the true searcher intent. Probably the best example I can come up with on this front is say you wanted a pair of shoes of a specific brand, size, width, and model number. If they are not the most recent and most heavily marketed versions it can be tough. Auto-generated internal search pages on trusted brand sites rank well, while a small retailer carrying that specific shoe might be penalized by Panda.

Re: Google doesn't recognise or penalise stolen content

#80

Earlier quoted context omitted.

Did you get that $4 billion number from that quarterly results? Does that include things like their payments to Opera and Apple? Does it include search rev share deals with entities like AOL and Ask (& soon to be Yahoo)? Does it include paying from Chrome distribution bundled with Flash security updates & such? I have never seen the overall numbers broken down in terms of what percent goes where on the different sort…

Prior to restructuring their reporting, Google reported as a cost paid distribution. I left in 2010, I started tracking the number in Q1 2011. It was $337M for the quarter. by Q4 of 2014 that number had ballooned to $968M for the quarter. In 2015 they changed the way the reported this number making future comparisons problematic. I expect it does include fees paid to Apple so that Apple would send search traffic to G…

Thanks for sharing that :)

One interesting bit from the most recent IAC investor conference call is on it they mentioned that their search deal with Google was renewed for another 4 years & that the rev share on mobile was lower than it was in the past. An analyst asking a question mentioned both Google and Yahoo! were lowering revenue share on mobile.

Post reply on HN