Live data from Hacker News

Facebook LLAMA is being openly distributed via torrents

github.com

351–360 of 719 posts

Re: Facebook LLAMA is being openly distributed via torrents

#351
post #183

Earlier quoted context omitted.

The Repilka subreddit became one of the weirdest places on the internet when their model got capped for adult content. https://www.reddit.com/r/replika/ Hundreds of men (and yes women) full on acting like they lost a spouse and posting constantly about it for weeks. AI is going to create some unusual social situations the general public isn't ready to grasp. And we're only in the early alpha stages.

I posit to a friend that: a) As these AI constructs become more advanced (especially around memory and personalization), we will eventually be able to treat them as people b) Some business will eventually sell an off-the-shelf product (hardware and/or software) that is an AI you can bring into your home, that you can treat as a friend, confidant and partner c) Someone will eventually lose their AI friend of many mont…

Have you seen the movie "Her"?

Re: Facebook LLAMA is being openly distributed via torrents

#352

It seems that the leak originated from 4chan [1]. Two people in the same thread had access to the weights and verified that their hashes match [2][3] to make sure that the model isn't watermarked. However, the leaker made a mistake of adding the original download script which had his unique download URL to the torrent [4], so Meta can easily find them if they want to. [1]: https://boards.4channel.org/g/thread/9184826…

What kind of recourse would meta have here? Sue him for breach of contract?

Re: Facebook LLAMA is being openly distributed via torrents

#353
post #41

It’s interesting that these models are both massively expensive to produce and self-contained to a degree that you can distribute the end product in a torrent. This has not been the case for most commercial software for the past 20 years, during the cloud era. If you could steal a dump of random Facebook source code, it would be 99% useless because it’s so closely tied to the infrastructure. There’s almost nothing yo…

(My posts are dead by default, but for the good showdead people if someone knows and can answer in a sibling reply) Are models like this copyrightable? It seems like this falls under the realm of "fact", which can't be copyrighted.

Under Feist Publications, Inc., v. Rural Telephone Service Co. ... it gets tricky.

From Wikipedia:

> The ruling of the court was written by Justice Sandra Day O'Connor. It examined the purpose of copyright and explained the standard of copyrightability as based on originality.

> The case centered on two well-established principles in United States copyright law: that facts are not copyrightable, and that compilations of facts can be.

> "There is an undeniable tension between these two propositions", O'Connor wrote in her opinion. "Many compilations consist of nothing but raw data—i.e. wholly factual information not accompanied by any original expression. On what basis may one claim a copyright upon such work? Common sense tells us that 100 uncopyrightable facts do not magically change their status when gathered together in one place. … The key to resolving the tension lies in understanding why facts are not copyrightable: The ″Sine qua non of copyright is originality."

> ...

> The standard for creativity is extremely low. It need not be novel; it need only possess a "spark" or "minimal degree" of creativity to be protected by copyright.

> In regard to collections of facts, O'Connor wrote that copyright can apply only to the creative aspects of collection: the creative choice of what data to include or exclude, the order and style in which the information is presented, etc.—not to the information itself. If Feist were to take the directory and rearrange it, it would destroy the copyright owned in the data. "Notwithstanding a valid copyright, a subsequent compiler remains free to use the facts contained in another's publication to aid in preparing a competing work, so long as the competing work does not feature the same selection and arrangement", she wrote.

> The court held that Rural's directory was nothing more than an alphabetic list of all subscribers to its service, which it was required to compile under law, and that no creative expression was involved. That Rural spent considerable time and money collecting the data was irrelevant to copyright law, and Rural's copyright claim was dismissed.

---

And so, my (I am not a lawyer) take on this is that the numbers of the model are not copyrightable. The selection of the source material is... kind of. This gets into a "a recipe is not copyrightable, yet a recipe book is"

The model may, however, be a trade secret. ( https://en.wikipedia.org/wiki/Trade_secret )

Re: Facebook LLAMA is being openly distributed via torrents

#354

Earlier quoted context omitted.

I posit to a friend that: a) As these AI constructs become more advanced (especially around memory and personalization), we will eventually be able to treat them as people b) Some business will eventually sell an off-the-shelf product (hardware and/or software) that is an AI you can bring into your home, that you can treat as a friend, confidant and partner c) Someone will eventually lose their AI friend of many mont…

At the end of the day, the Turing Test for establishment of AI personhood is weak for two reasons. 1. We're seeing more and more systems that get very close to passing the Turing Test but fundamentally don't register to people as "People." When I was younger and learned of Searle's Chinese Room argument, I naively assumed it wasn't a thought experiment we would literally build in my lifetime. 2. Humanity has a histor…

I dont know if I understand this general take I see a lot. Why care about this "AI personhood" at all? What is the tacit endgame everyone is always referencing with this? Isn't there just so many more both interesting and problematic aspects already here? What is the use of diverting the focus to some other point. "I see you are talking about cows, but I have thoughts about the ocean."

Re: Facebook LLAMA is being openly distributed via torrents

#355

Earlier quoted context omitted.

And if we feel as if we were losing a real person, AIs will have to be treated to some degree as if they were real people (or at least pets) rather than objects. This could be interesting, because so far the question of personhood and sentience of AIs has revolved around what they are and what they feel rather than what we feel when we interact with one of them.

Kids can feel like they're losing a real friend if they lose a stuffed animal. What's the progress on making teddy bears people?

Eventual, but needed. Kids feel pretty isolated during various pandemic lockdowns and maybe their parents have a lot of childfree friends, so they'll need companions, more than just a toy, even if technology marches on so quickly they'll be outdated soon enough. One day, you'll hear that supertoys last all summer long.

Re: Facebook LLAMA is being openly distributed via torrents

#356

Maybe this is an intentional leak to damage OpenAI. A supposedly better model by some accounts that strikes right at the heart of their business plan of selling access for $250k/year. One month of access to their service could buy a machine capable of running this leaked model. Facebook nerfs a potential upstart competitor to keep current big-tech cartel stable. Maybe this is a bit conspiratorial, but we live in the…

Not a conspiracy at all. See also IE, Android, Kubernetes...

Re: Facebook LLAMA is being openly distributed via torrents

#357

Earlier quoted context omitted.

Yeah, that's probably the most dystopian thing. This is almost a guaranteed outcome - someone pays a high subscription cost and cultivates a model with their personal details for years, and then loses all of it when they can't keep up the subscription cost. Cue a month or two later - they buy back in and their model has been wiped and their AI friend now knows nothing about them. It's easy to poke fun at people who u…

Or maybe they sell that data to another company that operates kind of like a collections agency, which takes on the 'risk' of storing the data, then repeatedly calls and offers to give them their AI friend back at an extortionate rate. The data privacy side of this is an interesting conversation as well. Think of the information an employee or hacker could leak about a person after they spent some time with such an i…

Imagine if they could transform the AI companion model into an extortionist model.

Re: Facebook LLAMA is being openly distributed via torrents

#358
post #164

Earlier quoted context omitted.

Electricity costs are basically irrelevant because the cards are so expensive. A100 cards consume 250w each, with datacenter overheads we will call it 1000 kilowatts for all 2048 cards. 23 days is 552 hours, or 552,000 kilowatt hours total. Most dataceneters are between 7 and 10 cents per kilowatt hour for electricity. Some are below 4. At 10 cents, that's $53,000 in electricity costs, which is nothing next to $30 mi…

> Electricity costs are basically irrelevant because the cards are so expensive. You mean in terms of money. I think this is exactly the problem that we have in CS, nobody really cares about CO2.

I'm on some strong hopium that those DC's run on renewables or nuclear, green energy.

Re: Facebook LLAMA is being openly distributed via torrents

#359

Earlier quoted context omitted.

At the end of the day, the Turing Test for establishment of AI personhood is weak for two reasons. 1. We're seeing more and more systems that get very close to passing the Turing Test but fundamentally don't register to people as "People." When I was younger and learned of Searle's Chinese Room argument, I naively assumed it wasn't a thought experiment we would literally build in my lifetime. 2. Humanity has a histor…

I dont know if I understand this general take I see a lot. Why care about this "AI personhood" at all? What is the tacit endgame everyone is always referencing with this? Isn't there just so many more both interesting and problematic aspects already here? What is the use of diverting the focus to some other point. "I see you are talking about cows, but I have thoughts about the ocean."

I was responding primarily to parent's (a): "As these AI constructs become more advanced (especially around memory and personalization), we will eventually be able to treat them as people."

Re: Facebook LLAMA is being openly distributed via torrents

#360

Earlier quoted context omitted.

Your SHA256 hash won’t be able to summarize text, write poems, or make up plots for books. The crazy thing about these models is that the compute power going into them is at least somewhat reversible.

Are you saying they are like compact memoizers? What Stable Diffusion can fit into that model is amazing.

When SD1.4 dropped, someone here described how those models are a form of lossy compression.
Post reply on HN