Live data from Hacker News

To show how easy it is for plagiarized news sites to get ad revenue, I made one

cnbc.com

11–20 of 40 posts

Re: To show how easy it is for plagiarized news sites to get ad revenue, I made one

#12
OK first try. But needs more work.

Not much proven so far.

Many site seem to translate to language X and back to English to clean the data.

Research this.

Anyone using GANs yet?

How do you stop sites blocking your scraper?

There's money for the ad companies to allow you to plod along then steal your hard earned money because you are breaking the rules. Are they?

Re: To show how easy it is for plagiarized news sites to get ad revenue, I made one

#13
post #11

> These firms mostly sold “popunder” ads, which pop up a new link in a browser tab when you click something Who is buying ads on these networks? There cannot possibly be any returns can there?

It might be a "victim filter", some scams are created to avoid wasting time in people smart enough not to fall for the scam in the following steps.

Re: To show how easy it is for plagiarized news sites to get ad revenue, I made one

#14
post #10
post #5

We're working on another way for disseminating news. It might make plagiarism a little more difficult, while also working a little better for our audience: https://blog.nillium.com/what-can-napster-teach-local-news/

How does that help prevent plagiarism?

There's not much in the post so I'm gonna guess it's a form of content fingerprinting like we see with YouTube's Content ID, plus whatever is used in plagiarism-detection software used in schools and universities.

Re: To show how easy it is for plagiarized news sites to get ad revenue, I made one

#15
post #10
post #5

We're working on another way for disseminating news. It might make plagiarism a little more difficult, while also working a little better for our audience: https://blog.nillium.com/what-can-napster-teach-local-news/

How does that help prevent plagiarism?

Because it isn't full articles -- just updates as they happen coming straight from the newsroom, more like tweets. It's not to say that people can't plagiarize, but it wouldn't be as easy or make as much sense as just copy and pasting an article.

Re: To show how easy it is for plagiarized news sites to get ad revenue, I made one

#16
post #11

> These firms mostly sold “popunder” ads, which pop up a new link in a browser tab when you click something Who is buying ads on these networks? There cannot possibly be any returns can there?

Often it's just affiliate fraud, you load casinos, aliexpress and what not affiliate links as be hope for the payout. That is why they redirect like crazy in order to hide the tracks since the sites offering affiliate services don't want it.

Re: To show how easy it is for plagiarized news sites to get ad revenue, I made one

#17

I suspect one of the hard parts of this for Google is that many news sites legitimately publish the same articles because of wire services and correspondence arrangements, like AP and Reuters. Hard to tell whether the new site is plagiarizing or syndicating.

Don't they always include the source as AP or Reuters in the body of text somewhere?

Re: To show how easy it is for plagiarized news sites to get ad revenue, I made one

#18
post #4

Earlier quoted context omitted.

It may be a low bar, but being able to automate this (scraping website ansible playbook?), makes the effort required as near-nil. They only have to clear $50 to pay for ALOT of domains and hosting.

Sure but you need thousands of legitimate-seeming pageviews to get that $50 back, and the networks - even or especially the bottom tier ones - are likely to be hotter on click fraud than scraped content.

+1

its really not easy to get the ad revenue, scraped websites can't rank good and get enough visitors, of course one can make fraud clicks system, but if one can do this level, he may easily find more interesting things, but not peanut $$

Re: To show how easy it is for plagiarized news sites to get ad revenue, I made one

#19

Doesn't say how much she made from it. Guess: very close to zero, if not zero.

Including hosting? I'm sure it's actually in red.

I'm not. I knew a developer who, in his spare time, developed a clever scraper. He scraped the top stories and results from Google, then scraped similar content based on Google's own ranking, then submitted that content to his own aggregator sites (all resolving to the same server). He ran ads on it. He got plenty of traffic and was net It's not that expensive to run a site and the right advertising partners (cough Taboola cough) pay nicely.
Post reply on HN