Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

401–410 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#401

If you're an author with a book likely to have be hoovered up, I wonder what you'd get from the fb models if you asked "complete this in the style of [author] in [book]: [quite a long excerpt]" If you get a direct quote then you're good with your claim, surely.

The way it works counts if you bring prompting into it. It could easily have learned enough style chops of [author] from other sources to mimic/predict those stanzas from raw data points.

Whatever the ruling one thing is for sure, plagiarism is no longer the sincerest form of flattery. The human authors are out for AI blood on this.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#402

Earlier quoted context omitted.

> But I prefer looser intellectual property rights anyway so Im ok with it I think more people, potentially anyways, would feel similar to to this if it applied even somewhat equally. Instead, companies can seemingly do whatever they please whereas lawyers will send letters to your home for downloading a single episode of game of thrones.

From the article, they took steps to avoid using their IP addresses. Individuals doing the same using a VPN are pretty much immune from any legal issues.

This is one small blip in an incredibly long history of companies being able to not care about copyright while individuals must.

Workarounds with a VPN are great and all, but they are a band-aid on a systemic problem.

(You are not immune, by the way, if your VPN company is subject to a subpoena and isn't one of very few actually no-log services)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#403
post #156

Earlier quoted context omitted.

What happens in US right now shows that change is achieved through voting. There are other examples as well in Europe where things did change because of how people voted. If the change is good or bad depends on your perspective. For me the annoying part is that people vote for a guy because of a couple heavily advertised issues, ignoring all the other plans or the fact that he might not keep his word. Then they are u…

It's unclear yet whether anything will really change. It is a perfect example though of how the rich are above the law.

USAID has already been shut down

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#404

Earlier quoted context omitted.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.

I remember calling a taxi 3 hours before my flight to get to SFO. After an hour and four different phone calls to the taxi company, I took BART and barely made it before the counter closed.

The feedback system incentivizes drivers and riders to behave.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#405

Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.

Anna's Archive encourages (and monetizes!!) the use of their shadow library for LLM training. They have a page dedicated to it on their site. You pay them, and they give you high download speeds to entire datasets.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#406
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life. Just to point, but the material in question was public domain, so nobody had even a copyrights claim over it.

Do you have a citation for that claim? I've not seen a claim that none of the material had copyright before.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#407
"Say they hood robin, ain't that a b*, take from the poor and give to the rich."

- Ice Cube.

Meta will face no consequences. Say your a small publisher and you'd like a bit of compensation. If you dare sue Meta can just blacklist your books on its platforms. Even if they don't, you probably don't have the money to sue one of the biggest companies on earth.

I think copyrights should be limited to 25 years after first publication. This would fix plenty of issues and give the AIs of the world plenty to learn from.

Who am I kidding, Meta will take what they will. For that author making 20k a year, be honored to be of use to Meta.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#408
post #124

Earlier quoted context omitted.

> Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. I don't understand why you wouldn't just buy copies of the books. Seems like such a relatively inexpensive way to strengthen your legal case.

Too much paperwork, too much effort. These are important people, doing much more important stuff than whatever book authors do. Or so they think, I think.

I doubt they think that way, but even if they did, they'd be right - for 99% of the works in question, the biggest value they gave to the world is, by far, being part of the LLM training corpus.

There's lots of content out there. Most of it is noise. People forget because they're only ever exposed to an aggressively curated fraction of it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#410
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life. In case anybody here doesn't know, that's a reference to Aaron Swartz, an activist (and Reddit co-founder) that was risking 35 years in prison and a $1 million fine just for downloading a lot of academic papers from JSTOR. He eventually took his life because of the pressure. May his soul rest in peace.

Thanks for sharing this. Reading his story is kind of insane honestly. He created the CC licenses which I did not know. What an icon, truly.
Post reply on HN