Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

841–850 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#841
post #185
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

You're conflating different problems. Big corporations are too big, they should just not exist. When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature. The IP laws may need rethinking. Saying that they should disappear because big corporations are above the law doesn't help, though. First kill the big corporations, then think about fair laws. Changing the law now…

> When you have corporations more powerful than the government of the biggest states

I don’t know how you define powerful, but I highly doubt it is at that point.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#842

Earlier quoted context omitted.

Because the wealth was transferred from India to the "antipode" through Colonization. GDP reduced from 25% Pre-Colonization (and 30% if you take Pre-Islamic Colonization) to merely 4% Post-India's Independence. At least Indians are not reverse-colonizing the West.

They are throwing away their own culture chasing the "wealth" though. Seems to be the same lust for money that drove the colonists.

What are you even talking about? Indians carry their culture/traditions everywhere they go. No one is "throwing it away". You are talking as if Indians have started to emigrate in just the past few decades. Indians have been navigating the World for the past 6000+ years at the very least (recorded history). The word "Navigation" is itself a Sanskrit word "Nava gatih". We are an Ancient Civilization and the oldest surviving Civilization. Everyone else either converted and destroyed their own civilization or were destroyed by invaders.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#843

Unless Meta 'fessed up to this (which seems unlikely), the headline here is missing the word "allegedly".

Meta admitted to the torrenting more than a month ago. The reason this is in the news is because some of the emails discussing it have been unsealed.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#844

Earlier quoted context omitted.

I don't think I've heard the term "English empire". Is it an attempt by the Scottish to pretend they weren't involved?

Is this an attempt to imply the Scots had imperial ambitions and have not been fighting to keep their homes free of invasion for several thousand years? Fuck this sounds familiar right now

Didn't Scotland try to make an empire in Eastern Canada, eastern USA, Africa, and Panama, then bankrupt themselves and agree to the Act of Union with England making Great Britain?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#845
post #193
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

i know of a company that poisoned an entire town! thats terrorism if done by an individual. the company still exists, just paid a settlement and carried on...

Russia has a town called Asbest which has an open-pit asbestos mine half the size of Manhattan where they mine with explosives.

https://en.wikipedia.org/wiki/Asbest

https://www.youtube.com/watch?v=cy3piCUPIkc - VICE documentary and visit video. I think it contains an interview with an American woman who suffered from WR Grace and Company's asbestos mining and manufacturing in the USA, she says "they knew, they knew". WRG faced 129,000 personal injury claims and set asude $3 Bn for settling asbestos related lawsuits.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#846

I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…

Trained on doesn't mean significant inclusion in the final state.

Is it truly a violation of copyright when a user hacks out bits and pieces of easily restyled raw data points from a model to look samey? what about if it takes two models? Might be time to accept humans are just cooked in their ability to discern attempts at direct plagiarism - just as it is hard to discern Sky voice from Her voice.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#847

Earlier quoted context omitted.

i will simply disagree that the dominant social dynamic leading to this favoring of foriegn capital over us law is not systemic corruption.

So there are two different categories of things here. One is, they ban cannabis and put individuals in prison for it, but then if you pay thousands a year for overpriced health insurance and the insurance pays thousands of dollars for a doctor to ask you some cursory questions and a pharma company to manufacture the drug, you can get a prescription for opioids, which are way more dangerous. But that isn't the big guy…

I mean there are examples in the same category as well: how many years was it illegally exported from dispensaties between states? New state legalizes? day one the corporate dispensary is stocked which is curious since it takes several months to grow. Also, lots of foriegn capital involved in the industry.

Not just meta, Open AI, spotify, youtube...its become a routine exception and can now be relied upon.

I agree that the legal fees could be a big factor, but it seems cases aren't even filed.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#848

Earlier quoted context omitted.

You know what would happen right? Copywrite expiring in 20 years doesn't mean access is democratized. Publishers would likely keep the price the same, but instead is the author getting a cut, they just take everything. Besides. The public isn't owed the fruits of my labor for free.

I honestly suspect fairly little would change. The US operates with a 20 year copyright for nearly 200 years, these long copyrights are actually far newer. Also, you are not owed a monopoly on arrangement of words enforced by the public. There are plenty of other places to spend tax dollars.

What tax dollars are we spending enforcing copywrite laws?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#849

Earlier quoted context omitted.

Not only was it inevitable, if we were so inclined and willing to use the regulatory pen, we could've simply written into law that for Taxi's to operate, they must be well maintained and must accept all major forms of payment. And yeah, the Taxi industry would've fought it because every company ever has fought every regulation ever no matter how much it stands to benefit both their customers and they them-fucking-sel…

> Not only was it inevitable, if we were so inclined and willing to use the regulatory pen, we could've simply written into law that for Taxi's to operate, they must be well maintained and must accept all major forms of payment. That was frequently already the case. They were required to accept credit cards but then the card reader would be "broken" and it wasn't worth anybody's time to dispute it instead of just pay…

> That was frequently already the case. ...the card reader would be "broken"

I traveled a lot to a smallish town for work before Uber got there and ran into this several times. After the second or third time, I started just saying "well that sucks for you" and starting to leave. Suddenly it would work.

Yes it sucked, but it didn't really impact much.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#850
post #839

Earlier quoted context omitted.

> Getting more comprehensive enforcement of laws in general against well-heeled players is a good thing. Whether something is good independent of what it takes to achieve it is a separate question from whether that's where you should focus your efforts. > We would have a lot less bad law if laws were enforced more evenly, because people would more quickly see their true effects, rather than having to wait until compa…

So I get your argument, but by that logic the only bad laws that get repealed will be those that affect big business, and the laws against individuals without resources will still be in place. I think we all agree that it’s unfair there’s a very large difference in enforcement of law between those with resources and those without, but I think to prevent that we need to figure out how to prevent the capture of the gov…

> It’s almost like we need to ensure no one has much more resources than anyone else (ya know, workers owning the means of production) so there’s a more level field!

Heh, yeah, I was going to post to add this as well. That is the underlying problem. I don't necessarily think it has to mean "workers owning the means of production" per se, more like "the richest person's wealth cannot be more than X times the poorest" and "the largest participant in a market cannot have more than Y% market share", but the idea is similar. :-)

Post reply on HN