Live data from Hacker News

Meta pirated books to train its AI

theatlantic.com

21–30 of 77 posts

Re: Meta pirated books to train its AI

#21
post #5

There seems to be a precedent with tech companies that if you do a lot of illegal stuff in the beginning you can count on either failing before it matters or you can buy your way out of the consequences. Bonus points if you induce regulatory capture to make sure no one else can follow in your footsteps. Please HN do your thing and prove me wrong or tell me the benefits to society outweigh the cons, because otherwise…

Well, the fruit of it, Llama 3, is public. Maybe it comes with some sort of license, but considering how it was made, I wouldn't feel guilty about violating it.

Then again, I would happily train a model on Anna's Archive myself without feeling guilty about it, if I had the resources to do so.

I think it ranks very, very low on the list of bad things Facebook has done.

Re: Meta pirated books to train its AI

#22

> Bulk downloading is often done with BitTorrent, the file-sharing protocol popular with pirates for its anonymity, and downloading with BitTorrent typically involves uploading to other users simultaneously. Isn't this a twofold misunderstanding of BitTorrent? I haven't used it much, but I've never believed BitTorrent to be popular for anonymity (is it even truly anonymous?), I thought it was popular because it makes…

> [D]ownloading with BitTorrent typically involves uploading to other users simultaneously. The key word here is "typically". Mutual sharing is designed into the protocol itself, particularly in the early seeding period, where higher priority is given to peers that re-share their portions of the torrent. Yes, you can turn off uploads, but that's not the default in any client I've ever used. So I don't see anything wr…

But Meta specifically says they took precautions to not do that, and in every client I've seen the only thing they would have to do is make sure a single checkbox which is prominently visible on the main download modal isn't checked!

I'm not objecting to the use of the word "typically", I'm objecting to the explicit suggestion that Meta might have seeded the pirated works. It's unlikely and unnecessary to suggest—there's plenty else wrong with this picture so there was no need to include this idea, it's just a distraction from the real problem.

Re: Meta pirated books to train its AI

#26
post #5

There seems to be a precedent with tech companies that if you do a lot of illegal stuff in the beginning you can count on either failing before it matters or you can buy your way out of the consequences. Bonus points if you induce regulatory capture to make sure no one else can follow in your footsteps. Please HN do your thing and prove me wrong or tell me the benefits to society outweigh the cons, because otherwise…

Well, the fruit of it, Llama 3, is public. Maybe it comes with some sort of license, but considering how it was made, I wouldn't feel guilty about violating it. Then again, I would happily train a model on Anna's Archive myself without feeling guilty about it, if I had the resources to do so. I think it ranks very, very low on the list of bad things Facebook has done.

When one end of the scale is "widespread and systematic contribution to genocide", everything else starts to look a little rosier, doesn't it?

Re: Meta pirated books to train its AI

#27
post #5

There seems to be a precedent with tech companies that if you do a lot of illegal stuff in the beginning you can count on either failing before it matters or you can buy your way out of the consequences. Bonus points if you induce regulatory capture to make sure no one else can follow in your footsteps. Please HN do your thing and prove me wrong or tell me the benefits to society outweigh the cons, because otherwise…

It is the golden rule. Whoever has the gold rules. Laws are for mortals, while big-tech can do whatever they please while channeling their masculine energy. Basically, behave in a self-serving manner violating basic norms such as dont do unto others what you dont want done to you. As content gets polluted with ai-garbage, human-generated content will regain value. Here is a startup idea: A startup that protects creat…

> while channeling their masculine energy

What has this to do with masculinity? Is this cheap misandry?

Re: Meta pirated books to train its AI

#28
post #5

There seems to be a precedent with tech companies that if you do a lot of illegal stuff in the beginning you can count on either failing before it matters or you can buy your way out of the consequences. Bonus points if you induce regulatory capture to make sure no one else can follow in your footsteps. Please HN do your thing and prove me wrong or tell me the benefits to society outweigh the cons, because otherwise…

The benefits to society outweigh the cons. If AI and copyright are really destined to fight each other -- an exercise in question-begging if there ever was one -- then copyright must lose.

If only because any other outcome will radically empower corporations (and entire countries!) that don't GAF about copyright.

There. We good?

Post reply on HN