Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

891–900 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#891

Earlier quoted context omitted.

Nothing stops you from downloading Ann’s archive and training a model on it, right? The likelihood that you, as an individual, get sued over is is virtually zero. This is what Meta tried to do, quietly download and use the data, to do research and advance their LLMs, without trying to establish any legal precedents or pick up fights.

In Germany people are sued for illegal movie downloading all the time. It's hard to imagine the companies behind that operation aren't aware that you can also download books.

People are also sued over books

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#892

Earlier quoted context omitted.

I really don't think that Meta did this because the alternative would have been too onerous; they are a huge org, they could work through whatever loopholes required. They did it because it would have cost money and there will be no penalty for not paying.

So, if they're sued in Japan, or France, do you think that the courts will take any special measures because it's a valuable American corporation? I suspect that if the case is reasonable they will just convict, and quickly-- appeal denied and all simply because the laws are so straightforward.

I must have failed to clearly express myself - I don't think Meta should be doing what they are doing, I hope they do end up being punished. But the only way that Meta is going to change its behaviour is by being held accountable in a way that's much more difficult and costly than if they'd simply followed the law in the first place.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#893

Earlier quoted context omitted.

So, if they're sued in Japan, or France, do you think that the courts will take any special measures because it's a valuable American corporation? I suspect that if the case is reasonable they will just convict, and quickly-- appeal denied and all simply because the laws are so straightforward.

I must have failed to clearly express myself - I don't think Meta should be doing what they are doing, I hope they do end up being punished. But the only way that Meta is going to change its behaviour is by being held accountable in a way that's much more difficult and costly than if they'd simply followed the law in the first place.

Ah, I'm not sure exactly what I believe here, but this kind of torrenting is obviously illegal-- I'm personally split on how I feel about it morally, because some of these people really are trying to preserve knowledge, and I think that's commendable, at the same time, commercial piracy is something which really does screw over authors with it being some kind of theft-of-service type thing where people exploit other people's work-- and if they felt that the work had no value they could have written another text themselves.

I only really wanted to convey that I believed that it probably isn't obviously easy for Meta to get away with anything in this, even if the US government decides to be lenient for the sake of a high market-cap US company simply because other countries are a viable place to sue as well.

I think I misinterpreted your comment as that you thought that Meta thought that costs would be low because they imagined a US court system that simply ignored the illegality because it's they who committed it, when nothing like that is actually implied in your comment.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#894
post #609

Earlier quoted context omitted.

lol I absolutely do not want non digital goods nor pirating. Ever. It's 2025. I don't have a cdplayer, a tape player, a blue ray player, I don't even know what the most modern "blue ray" disc would be. I have $2k worth of vinyls that are just unique copies I display as art I'll never put in my record player, that's also never been used. I don't want to constantly worry about 60gb of mp3 files. Oh no, that TV show I'l…

I don't know how you went from "don't pay for overpriced digital goods, just pirate them instead" to "hurr durr start using blurays and vinyls". Reading comprehension is a lost art nowadays.

He just wantrd to express how superior to the young kids he is .. while all the points about ethics, freedom, and privacy went over his head.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#895

Earlier quoted context omitted.

Swartz wasn't the kind of person to accept a plea bargain from an overzealous prosecutor who was indicting him on 13 felony charges with a possible sentence up to 35 years along with a $1 million fine. I assume he wanted a trial because he wanted to continue his fight for open access. And maybe he thought he might lose, but wouldn't lose on all counts, and would make the prosecution look unreasonable in the public ey…

His criminality is one matter, but the full weight of the Federal Government on him was an entirely separate matter. A federal prosecutor's job is to jail you regardless of whether it is for downloading a file from a server or for trafficking in humans, and they will come at you with the same vigor regardless of the crime. And nothing has changed about that.

That is complete bullshit as evidenced by the federal criminal sitting president.... The ONLY time I ever see this argument is to try and paint over blatant police and state injustice and tyranny.

"A federal prosecutor's job is to jail you regardless of whether it is for downloading a file from a server or for trafficking in humans, and they will come at you with the same vigor regardless of the crime."

You have to be malicious to put forward this statement in the current environment. Or you are so propagandized you think it's true? Either one is very frightening

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#896
We're starting to find out that Meta ruined LibGen for the rest of use who used it like a library. Just like how Google screwed over libraries by sending interns to the Stanford library to checkout books they scanned into Google Books. Not to increase shared knowledge or preserve human artificats, but to put them all in a museum and, to paraphrase Joni Mitchell, charge the people a dollar and a half just to see 'em.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#897

Earlier quoted context omitted.

But the point of the response is that "getting money from selling music" is, in digital era, artificial scarcity. I.e. the copyright laws that big corporations are lobbying for continued enforcement and tightening, are the very thing that create this artificial scarcity that they are best positioned to profit off. Cut out copyright, and no one will be getting any money from selling music per copy (or equivalent) - as…

digital music is not artificial scarcity, because it's not the copied bits that are the resource, it's attention. we only have so much time and attention for consuming media, and only so much attention and memory space in our brains for keeping track of where to find it. large budgets can easily dominate these channels and limit the average person's apparent choice. this is what I mean when large players would outcom…

Are you talking about mere distribution? In that case, a few large players leveraging scale to drive costs down to near zero sounds great.

I’m still not seeing how lack of copyright hurts small artists or consumers. Small distributors, maybe, but that’s not doing harm to the arts.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#898

[flagged]

You can't post like this here, regardless of whom you want to kill. Between this and other abusive comments, we've banned the account.

https://news.ycombinator.com/item?id=42946919

https://news.ycombinator.com/item?id=42690711

If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#899
post #868

Earlier quoted context omitted.

> so it isn't a restriction of your freedom of speech when you choose to seek out and repeat somebody's very particular text. I hadn't made that claim, but I will in now that you've brought it up. Art operates as part of a discussion, the reference to and re-use of prior art is a key part of the how that happens. There are sooo many cases of copyright being used to limit the freedom of expression, that this really is…

But we're talking about extremely direct copying. Actual computerized copying, typically verbatim. Doing things relating to discussion of a work are typically permitted, but you have no reason to use anybody's particular work other than to make use of the work he did in creating it.

> But we're talking about extremely direct copying. Actual computerized copying, typically verbatim.

Copyright doesn't just extend to "literal direct copying". When you claim copyright doesn't harm anyone, you can't ignore all the other types of activity it prohibits.

> Doing things relating to discussion of a work are typically permitted,

Only if you limit the meaning of "discussion" so much that it no longer includes the process of making art.

> but you have no reason to use anybody's particular work other than to make use of the work he did in creating it.

Did you not ready my comment? I already explained the reason. Creative works become part of our culture, you can't choose which works will do that, you can only choose to participate in that culture or not.

Copyright is a social system for artificially limiting access to our shared culture and thus also limits participation in that culture.

I understand the value of a limited copyright system, but anyone that claims that our copyright system doesn't cause harm or cost us anything isn't being realistic. Copyright duration should be far more limited and we need significant reforms to the DMCA. Personally, I think even all non-commercial distribution should be legal as copyright should only grant commercial rights.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#900
Copyright law needs major reform. We need to figure out a way to let authors monetize their work while not making complying with the law so burdensome. We've created a system where people who (understandably) ignore the law benefit at the expense of people trying to do the right thing.
Post reply on HN