Live data from Hacker News

Llama 2

ai.meta.com

201–210 of 860 posts

Re: Llama 2

#201
Intersting that they did not use any facebook data for training. Either they are "keeping the gud stuff for ourselves" or the entirety of facebook content is useless garbage.

Re: Llama 2

#202

Hey HN, we've released tools that make it easy to test LLaMa 2 and add it to your own app! Model playground here: https://llama2.ai Hosted chat API here: https://replicate.com/a16z-infra/llama13b-v2-chat If you want to just play with the model, llama2.ai is a very easy way to do it. So far, we’ve found the performance is similar to GPT-3.5 with far fewer parameters, especially for creative tasks and interactions. Dev…

I'm wondering how do people compare different models? I've been trying chatGPT 3.5, bing chat (chatgpt 4 I believe?), and bard, and now this one, and I'm not sure if there's a noticeable difference in terms of "this is better"

Re: Llama 2

#203

Hey HN, we've released tools that make it easy to test LLaMa 2 and add it to your own app! Model playground here: https://llama2.ai Hosted chat API here: https://replicate.com/a16z-infra/llama13b-v2-chat If you want to just play with the model, llama2.ai is a very easy way to do it. So far, we’ve found the performance is similar to GPT-3.5 with far fewer parameters, especially for creative tasks and interactions. Dev…

How are the model weights licensed?

Re: Llama 2

#204
post #74

Making advanced LLMs and releasing them for free like this is wonderful for the world. It saves a huge number of folks (companies, universities & individuals) vast amount of money and engineering time. It will enable many teams to do research and make products that they otherwise wouldn't be able to. It is interesting to ponder to what extent this is just a strategic move by Meta to make more money in the end, but wh…

In a free market economy everything is a strategic move to make the company more money. It's the nature of our incentive structure.

Yes, that's true. But also vast majority of transactions are win-win for both sides, creating more wealth for everyone involved.

Re: Llama 2

#205
post #74

Making advanced LLMs and releasing them for free like this is wonderful for the world. It saves a huge number of folks (companies, universities & individuals) vast amount of money and engineering time. It will enable many teams to do research and make products that they otherwise wouldn't be able to. It is interesting to ponder to what extent this is just a strategic move by Meta to make more money in the end, but wh…

In a free market economy everything is a strategic move to make the company more money. It's the nature of our incentive structure.

Most, but not all things are strategic moves.

Some moves are purely altruistic. Some moves are semi-altruistic - they don't harm the company, but help it increase its reputation or even just allows them to offer people ways to help in order to retain talent. (Which is also kind of strategic, but in a different way.)

Also, some things are just mistakes and miscalculations.

Re: Llama 2

#206
From a modeling perspective, I am impressed with the effects of training on 2T tokens rather than 1T. Seems like this was able to get LLAMA v2 7b param models equivalent to LLAMA v1's 13b performance, and the 13b similar to 30b. I wonder how far this can be scaled up - if it can, we can get powerful models on consumer GPUs that are easy to fine tune with QLORA. A RTX 4090 can serve an 8-bit quantized 13b parameter model or a 4-bit quantized 30b parameter model.

Disclaimer - I work on Databricks' ML Platform and open LLMs are good for our business since we help customers fine-tune and serve.

Re: Llama 2

#207
post #198

Earlier quoted context omitted.

And yet here we are a few weeks after that with a free to use model that cost millions to develop and is open to everyone. I think you’re taking an unwarranted entitled view.

You act like this is a gift of charity instead of attempts to stay relevant.

What? Tell me you don't follow the space. FB AI is one of the top labs..

Re: Llama 2

#209

Earlier quoted context omitted.

Apple does not have the capability to train a LLM currently.

I very much doubt that.

If they want to own the whole stack, I don't think they have much to work with. Their highest-end server chip is a duplex laptop SOC, with maxed-out memory that doesn't even match the lowest-end Grace CPU you can buy (nevermind a fully-networked GH200). Their consumer offerings are competitive, but I don't think Apple Silicon or CoreML is ready to seriously compete with Grace and CUDA.

Re: Llama 2

#210

In the things you can't do (at https://ai.meta.com/llama/use-policy/ ): "Military, warfare, *nuclear industries or applications*" Odd given the climate situation to say the least...

Same thing deep inside the Java TOS. I remember it from like 20 years ago.
Post reply on HN