Live data from Hacker News

Meta pirated books to train its AI

theatlantic.com

71–77 of 77 posts

Re: Meta pirated books to train its AI

#71
post #5

There seems to be a precedent with tech companies that if you do a lot of illegal stuff in the beginning you can count on either failing before it matters or you can buy your way out of the consequences. Bonus points if you induce regulatory capture to make sure no one else can follow in your footsteps. Please HN do your thing and prove me wrong or tell me the benefits to society outweigh the cons, because otherwise…

Often it happens because there is a lot of market demand for the thing they are doing even if there are regulations against it, so you get Uber instead of normal taxis because the taxi service was no good, in this case you get informed AI instead of the information being tied up in court battles for years. So often the benefits to society do outweigh the cons, though not necessarily in every case.

I've been quite glad of Ubers and LLMs. I think it would be good if the LLMs could read the stuff Google Books scanned but was unable to make public. (this stuff https://www.reddit.com/r/books/comments/67fkkj/somewhere_at_...)

Re: Meta pirated books to train its AI

#72

Earlier quoted context omitted.

Well, the fruit of it, Llama 3, is public. Maybe it comes with some sort of license, but considering how it was made, I wouldn't feel guilty about violating it. Then again, I would happily train a model on Anna's Archive myself without feeling guilty about it, if I had the resources to do so. I think it ranks very, very low on the list of bad things Facebook has done.

> Well, the fruit of it, Llama 3, is public. Maybe it comes with some sort of license, but considering how it was made, I wouldn't feel guilty about violating it. Facebook would quite literally sue you until you killed yourself if you did that. It’s not the same rules for them and for us.

They have no real way of knowing though, unless someone goes advertising it.

Re: Meta pirated books to train its AI

#73

Earlier quoted context omitted.

I don't think there is any benefit to society when a company with a black box doesn't give back to the community They are at least giving the model they trained back to the community. https://www.llama.com/llama-downloads/

I have found throughout my life that the best gifts were ones that required me to give PII to the giver, and sign a binding agreement that legally restricts my use of said gift. Bonus points if you harass me about your newsletter.

This is obviously annoying, but I suspect most people use them by downloading them off HF, Ollama, or similar sites which require no agreement. I also wonder why there is so much attention put on Facebook rather than OpenAI for example. Facebook is at least giving the weights to the public.

Re: Meta pirated books to train its AI

#74

This is one of the reasons there should simply be a big tax on AI profits, assuming they come to pass: they literally require uncompensated training on the sum of human intellectual effort. There is no AI without standing on the shoulders of both giants and thousands of everyday authors. If you're just giving the result back to humanity, OK, there's a case for it being a fair trade, though there's still the question…

I'm pretty sure Facebook/Meta isn't earning any profits from this. They are mostly doing it to "commoditize their complement" and probably prevent the rise of another company to the Big Tech status.

Re: Meta pirated books to train its AI

#75
post #5

There seems to be a precedent with tech companies that if you do a lot of illegal stuff in the beginning you can count on either failing before it matters or you can buy your way out of the consequences. Bonus points if you induce regulatory capture to make sure no one else can follow in your footsteps. Please HN do your thing and prove me wrong or tell me the benefits to society outweigh the cons, because otherwise…

Russians have a phrase for this: "Never ask a billionaire how they made their first million".

Re: Meta pirated books to train its AI

#76
post #5

There seems to be a precedent with tech companies that if you do a lot of illegal stuff in the beginning you can count on either failing before it matters or you can buy your way out of the consequences. Bonus points if you induce regulatory capture to make sure no one else can follow in your footsteps. Please HN do your thing and prove me wrong or tell me the benefits to society outweigh the cons, because otherwise…

I think it is a good thing to undermine copyright by creating various carve-outs that will eventually make it untenable to maintain, which is probably one of the faster ways to reform the obsolete system.

Piracy doesn't care about borders, and if US firms don't do it, Chinese firms will. It is a national competitiveness issue in that sense, which is the easiest argument to get the government to do things to weaken copyright.

Laying the groundworks for the eventual defeat of MAFIAA is a great benefit to society if it happens, and in my opinion outweighs the nonexistent "damage" they did. There wouldn't be llama if they didn't pirate the books, and the authors won't get paid anything either. Bonus points if we can get rid of DRM anti-circumvention as well.

For what it's worth Facebook doesn't seem to be doing the regulatory capture part, unlike "Open"AI and Anthropic.

Re: Meta pirated books to train its AI

#77
post #46

My take is pretty depressing for a different reason - I understand the outrage against mark zuckerberg for nurturing such culture and making the executive decision, but also understand that at least one engineer was involved in writing and executing the code that does the piracy (with Product Managers and other cross-function employees) And given the importance and visibility of the work, it's pretty obvious - that p…

That's a very interesting and refreshing take. But that's the culture that has been nurtured in the capitalist present - climb up to the top and push the ladder behind you.
Post reply on HN