Live data from Hacker News

LLaMA: A foundational, 65B-parameter large language model

ai.facebook.com

1–10 of 209 posts

Re: LLaMA: A foundational, 65B-parameter large language model

#2
This blog post is terrible at listing the improvements offered by the model, the abstract is better:

> We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In particular, LLaMA-13B outperforms GPT-3 (175B) on most benchmarks, and LLaMA-65B is competitive with the best models, Chinchilla70B and PaLM-540B. We release all our models to the research community.

Re: LLaMA: A foundational, 65B-parameter large language model

#3

This blog post is terrible at listing the improvements offered by the model, the abstract is better: > We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In par…

> We release all our models to the research community.

This is yet more evidence for the "AI isn't a competitive advantage" thesis. State-of-the-Art is a public resource, so competing with AI offers no "moat".

Re: LLaMA: A foundational, 65B-parameter large language model

#4

This blog post is terrible at listing the improvements offered by the model, the abstract is better: > We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In par…

> We release all our models to the research community. This is yet more evidence for the "AI isn't a competitive advantage" thesis. State-of-the-Art is a public resource, so competing with AI offers no "moat".

The research is open, you still need

1) to have a good idea for a product that people want

2) lots of talent and resources to actually build and run everything

Re: LLaMA: A foundational, 65B-parameter large language model

#5
> To maintain integrity and prevent misuse, we are releasing our model under a noncommercial license focused on research use cases. Access to the model will be granted on a case-by-case basis to academic researchers; those affiliated with organizations in government, civil society, and academia; and industry research laboratories around the world. People interested in applying for access can find the link to the application in our research paper.

The closest you are going to get to the source is here: https://github.com/facebookresearch/llama

It is still unclear if you are even going to get access to the entire model as open source. Even if you did, you can't use it for your commercial product anyway.

Re: LLaMA: A foundational, 65B-parameter large language model

#6

This blog post is terrible at listing the improvements offered by the model, the abstract is better: > We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In par…

> We release all our models to the research community. This is yet more evidence for the "AI isn't a competitive advantage" thesis. State-of-the-Art is a public resource, so competing with AI offers no "moat".

Business orgs have been finetuning open-source models like these on their own internal data to create a moat since BERT in 2018.

Re: LLaMA: A foundational, 65B-parameter large language model

#8
post #4

Earlier quoted context omitted.

> We release all our models to the research community. This is yet more evidence for the "AI isn't a competitive advantage" thesis. State-of-the-Art is a public resource, so competing with AI offers no "moat".

The research is open, you still need 1) to have a good idea for a product that people want 2) lots of talent and resources to actually build and run everything

And then facebook will release a version which beats it ? :)

Re: LLaMA: A foundational, 65B-parameter large language model

#9
post #5

> To maintain integrity and prevent misuse, we are releasing our model under a noncommercial license focused on research use cases. Access to the model will be granted on a case-by-case basis to academic researchers; those affiliated with organizations in government, civil society, and academia; and industry research laboratories around the world. People interested in applying for access can find the link to the appl…

Facebook continues to appropriate the word “open” [1], which is sad really. There are plenty of good words they and others could use instead.

[1]: https://news.ycombinator.com/item?id=32079558

Re: LLaMA: A foundational, 65B-parameter large language model

#10

This blog post is terrible at listing the improvements offered by the model, the abstract is better: > We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In par…

> We release all our models to the research community. This is yet more evidence for the "AI isn't a competitive advantage" thesis. State-of-the-Art is a public resource, so competing with AI offers no "moat".

In terms of medieval warfare, what Facebook is doing here looks like filling the moat with rocks and dirt. OpenAI is worth billions, Microsoft is spending billions to retrofit most of their big offerings with AI, Google is doubtless also spending billions to integrate AI in their products. Their moat in all cases is “a blob of X billion weights trained on Y trillion tokens”. Facebook here is spending mere _millions_ to make and release competitive models, effectively filling in the moat by giving it to everyone.
Post reply on HN