Live data from Hacker News

LLaMA2 Chat 70B outperformed ChatGPT

tatsu-lab.github.io

1–10 of 135 posts

Re: LLaMA2 Chat 70B outperformed ChatGPT

#4
Yeah, my experience has been that every one of these freely downloadable models can be measured as "percent of chatgpt quality". and getting up to 85% is shockingly good.

*edit: oops, my brain inserted "by" in the middle of "outperformed chatgpt". I'll leave my wrong comment up as a testament to shame.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#5
It looks like ChatGPT length is 827 while LLaMA2 length is more than double at 1790.

Disclaimer from the site:

> Caution: GPT-4 may favor models with longer outputs and/or those that were fine-tuned on GPT-4 outputs.

> While AlpacaEval provides a useful comparison of model capabilities in following instructions, it is not a comprehensive or gold-standard evaluation of model abilities. For one, as detailed in the AlpacaFarm paper, the auto annotator winrates are correlated with length.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#7
post #3

I haven't had a chance to use the GPT-4 API yet - is it that much better than the GPT-4 available via ChatGPT? Or am I misunderstanding?

ChatGPT uses the GPT-4 API, so it's the same. With the API directly though you can change the system prompt, which can enable better results if you know what you're doing.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#10
post #3

I haven't had a chance to use the GPT-4 API yet - is it that much better than the GPT-4 available via ChatGPT? Or am I misunderstanding?

ChatGPT uses the GPT-4, but there are conspiracy theories circling that ChatGPT is neutered and thus not as good as GPT-4 through the API. The theory being that OpenAI are thottling the free version of GPT-4 (ChatGPT)
Post reply on HN