Earlier quoted context omitted.
You don't benchmark foundation model against RLHF model, results aren't very useful.
This does seem to be a RLHF model, not a base model. Unless 'supervised fine-tuning' and 'human preference' mean something else.
Llama 2
91–100 of 860 posts
Re: Llama 2
#92Re: Llama 2
#93Why doesn't FB create an API around their model and launch OpenAPI competitor? It is not like they don't have resources, and the learnings (I am referring to actual learning from users' prompts) will improve their models over time.
To reduce the valuation of OpenAI.
Re: Llama 2
#94This is really exciting. I work at Replicate, where we've already setup a hosted version for anyone to try it: https://replicate.com/a16z-infra/llama13b-v2-chat
Re: Llama 2
#95Earlier quoted context omitted.
>The 4k context length is nice, but RoPE makes it irrelevant anyway. Can you elaborate on this?
Here's some more info on it: https://arxiv.org/pdf/2306.15595.pdf https://www.reddit.com/r/LocalLLaMA/comments/14lz7j5/ntkawar... https://www.reddit.com/r/LocalLLaMA/comments/14mrgpr/dynamic... In short, the context is just an array of indexes passed along with the data, which can be changed to floats and encode more sparsely to scale to an arbitrarily small or large context. It does need some tuning of the model to…
Re: Llama 2
#96Hey HN, we've released tools that make it easy to test LLaMa 2 and add it to your own app! Model playground here: https://llama2.ai Hosted chat API here: https://replicate.com/a16z-infra/llama13b-v2-chat If you want to just play with the model, llama2.ai is a very easy way to do it. So far, we’ve found the performance is similar to GPT-3.5 with far fewer parameters, especially for creative tasks and interactions. Dev…
Re: Llama 2
#97Earlier quoted context omitted.
I think it's aimed at other social networks. TikTok has 1 billion monthly active users for instance
I think TikTok would just use it anyway even if they were denied a license (if they even bothered asking for one). They've never really cared about that kind of stuff.
Re: Llama 2
#98"Military, warfare, *nuclear industries or applications*"
Odd given the climate situation to say the least...
Re: Llama 2
#99Earlier quoted context omitted.
Google's model is not as capable as llama-derived models, so I think they would actually benefit from this. > I wouldn't be surprised if Amazon does as well. I would - they are not a very major player in this space. TikTok also meets this definition and probably doesn't have LLM.
Google has far better models than llama based models. They just simply don't put them facing the public. It is pretty ridiculous that they essentially just set a marketing team with no programming experience to write Bard, but that shouldn't fool anyone into believing they don't have capable models in Google. If Deepmind were to actually provide what they have in some usable form, it would likely be quite good. Despi…
All the AlphaGo/AlphaFold stuff is very cool, but since no one has seen their LLMs this is about as convincing as my claiming I've donated billions to charity.
Re: Llama 2
#100Earlier quoted context omitted.
> greater than 700 million monthly active users Hmm. Sounds like specifically a FAANG ban. I personally don't mind. But would this be considered anti-competitive and illegal? Not that Google/MS/etc. don't already have their own LLMs.
Most likely they want cloud cloud providers (Google, AWS, and MS) to pay for selling this as a service.