Live data from Hacker News

LLaMA: A foundational, 65B-parameter large language model

ai.facebook.com

21–30 of 209 posts

Re: LLaMA: A foundational, 65B-parameter large language model

#21
post #17

Earlier quoted context omitted.

Which kind of suggests Microsoft made a really bad move antagonizing the open-source community with Gtihub Copilot. They got a few years of lead time in the "AI codes for you" market, but in exchange permanently soured a significant fraction of their potential userbase who will turn to open-source alternatives soon anyway. I wonder if they'd have been better served focusing on selling Azure usage and released Copilot…

How did Microsoft sour developers with Copilot? I know dozens of people that pay for it (including myself) and I feel like it is widely regarded as a "no brainer" for the price that it's offered at. Please help me understand!

Controversy about unlicensed use of source code as training data and lack of attributions in the generated output.

Re: LLaMA: A foundational, 65B-parameter large language model

#23

This blog post is terrible at listing the improvements offered by the model, the abstract is better: > We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In par…

> We release all our models to the research community. This is yet more evidence for the "AI isn't a competitive advantage" thesis. State-of-the-Art is a public resource, so competing with AI offers no "moat".

They are only releasing it on a case-by-case basis and with a non-commercial license.

Re: LLaMA: A foundational, 65B-parameter large language model

#24
post #17

Earlier quoted context omitted.

Which kind of suggests Microsoft made a really bad move antagonizing the open-source community with Gtihub Copilot. They got a few years of lead time in the "AI codes for you" market, but in exchange permanently soured a significant fraction of their potential userbase who will turn to open-source alternatives soon anyway. I wonder if they'd have been better served focusing on selling Azure usage and released Copilot…

How did Microsoft sour developers with Copilot? I know dozens of people that pay for it (including myself) and I feel like it is widely regarded as a "no brainer" for the price that it's offered at. Please help me understand!

The company that tried to kill Linux in the 90s, owned by the world's most famously rich man, is now stealing my code and selling it back to me? Yeah, fuck that.

Re: LLaMA: A foundational, 65B-parameter large language model

#26
post #11
post #10

Earlier quoted context omitted.

In terms of medieval warfare, what Facebook is doing here looks like filling the moat with rocks and dirt. OpenAI is worth billions, Microsoft is spending billions to retrofit most of their big offerings with AI, Google is doubtless also spending billions to integrate AI in their products. Their moat in all cases is “a blob of X billion weights trained on Y trillion tokens”. Facebook here is spending mere _millions_…

I don't see such grand stratagems as being a likely explanation. It seems more likely that a bunch of dorks are running around unsupervised trying their best to make lemonade with whatever budgets they are given as nobody can realistically manage AI researches. Or at least, it seems that much like in political governance, events are suddenly outpacing the ability of corporate to react. In the case of OpenAI it is a "…

They almost certainly spent at least a few million dollars on this research project. Hard to say when the decision was made to open source this (from the outset or after it started showing results), but the decision was conscious and calculated. Nothing this high-profile is going to escape the strategic decision making processes of upper management.

Re: LLaMA: A foundational, 65B-parameter large language model

#28
post #18

> To maintain integrity and prevent misuse, we are releasing our model under a noncommercial license focused on research use cases. if say i wanted to replicate this paper for commercial use, what would it take and how do i get started? would FB have a basis for objection?

82342 GPU(A-100 80GB) hours for training their 7B(smallest) model according to their paper.

Re: LLaMA: A foundational, 65B-parameter large language model

#29
post #22

>'We release all our models to the research community' Where release means you fill out a form and wait indefinitely. Also no use for commercial purposes - which means 95% of users are out - certainly doesn't democratize LLMs.

> doesn't democratize LLMs

open source always leads to more democratization in the long run

Re: LLaMA: A foundational, 65B-parameter large language model

#30
post #11
post #10

Earlier quoted context omitted.

In terms of medieval warfare, what Facebook is doing here looks like filling the moat with rocks and dirt. OpenAI is worth billions, Microsoft is spending billions to retrofit most of their big offerings with AI, Google is doubtless also spending billions to integrate AI in their products. Their moat in all cases is “a blob of X billion weights trained on Y trillion tokens”. Facebook here is spending mere _millions_…

I don't see such grand stratagems as being a likely explanation. It seems more likely that a bunch of dorks are running around unsupervised trying their best to make lemonade with whatever budgets they are given as nobody can realistically manage AI researches. Or at least, it seems that much like in political governance, events are suddenly outpacing the ability of corporate to react. In the case of OpenAI it is a "…

this is a pretty biased and uninformed opinion. Pretty condescending to call it a "bunch of dorks...running around unsupervised."
Post reply on HN