Lately google search results are full of such low quality results, computer generated text, etc.
Same applies also to YT - videos with 2x speed, inverted colors etc.
61–70 of 169 posts
Lately google search results are full of such low quality results, computer generated text, etc.
Same applies also to YT - videos with 2x speed, inverted colors etc.
The ad revenue estimate could be off by as much as an order of magnitude. This is a non-story. There are millions of spam websites on the internet.
Lot's of garbage in it, but it also has some interesting docs. I'm more interested on the technical details of it. How it works internally for scrapping, text extraction, indexing, random users creation and uploading of stuff and so on.
I understand the privacy and copyright infringement concerns, but it's not like they've hacked into those other sites and uploaded their files. They're using existing and "open for everyone" pages. Also the estimates for revenue are highly exaggerated IMHO. Sure they're making some €€€, but not in the millions range :-)
On a final note, how is this vastly different from, let's say, https://www.scribd.com ?
Earlier quoted context omitted.
> If a file is available through non-auth http it’s unclear what the copyright is. Protocol has nothing to do with copyright. “Non-auth http” doesn't suddenly make copyright murky. > But by releasing a document publicly with unlimited access via URL the author explicitly allows unlimited distribution (and due to the nature of tcp/ip redistribution). There is an implicit license to exactly the redistribution necessary…
There is no implicit licence. This is an oft-circulated but fallacious rationalization made by computer people that is not in line with the law. The law does not recognize any such thing, and until the turn of the 21st century this was a recognized hole. Technically, the World Wide Web (and indeed other systems from FidoNet to SMTP) was a violation of many countries' copyright laws. Legislators fixed the hole, but no…
There is implicit copyright, I assume that’s what the parent comment was referring to. In the US, it is automatically illegal to copy something and redistribute it under copyright law [1]. The same is true in the EU [2]. While copyrights are not licenses, the parent comment is correct in the sense that one does not have a license to distribute copied content until one is granted that license explicitly by the copyright holder.
https://www.copyright.gov/help/faq/faq-general.html#register
https://euipo.europa.eu/tunnel-web/secure/webdav/guest/docum...
meh.
dude really shouldn't have doxxed him though. we dont even know if the guy in the whois is really the guy behind the site, since you can put anything in there.
Misleading title. As the article points out, the Alexa ranking puts it outside the top 200k most visited sites.
You do realize that if it's in top 10 million, out of estimated ~640 million websites, it's right at the top? It's not misleading at all, it's quite down to the point, just like the article itself.
Earlier quoted context omitted.
Just because they're published on the internet doesn't mean you can rehost them - that's copyright infringement. (The Internet Archive is very careful to skirt the edge of what's permissible and what they can get away with in this regard)
So copyright infringement of this form may or may not be immoral depending on who you ask. And may even be legal depending on where you live
Sorry if i don't understand this totally, but why is scraping for pdf and ppt/pptx files and mirroring them illegal? If you can reach that files just scraping it means they are somehow open to public access. No joke, i am genuinely asking.
Just because they're published on the internet doesn't mean you can rehost them - that's copyright infringement. (The Internet Archive is very careful to skirt the edge of what's permissible and what they can get away with in this regard)
TL;DR it's not a crime if nobody complains. Now that someone complained though, things might end up badly for the owner.