Live data from Hacker News

LLaMA2 Chat 70B outperformed ChatGPT

tatsu-lab.github.io

61–70 of 135 posts

Re: LLaMA2 Chat 70B outperformed ChatGPT

#61

Earlier quoted context omitted.

There is one: "The agreement between GPT-4 and humans reaches 85%, which is even higher than the agreement among humans (81%). This means GPT-4’s judgments closely align with the majority of humans. We also show that GPT-4’s judgments may help humans make better judgments. During our data collection, when a human’s choice deviated from GPT-4, we presented GPT-4’s judgments to humans and ask if they are reasonable. De…

80% agreement is high, but the margins between models at the top are so low that even that remaining 20% could be enough to alter the final rankings, depending on which direction it errs.

Is perfect agreement possible? And what is the definition of agreement? Humans don't agree about much.. are we saying agreement means it matches the truth after intensive investigation by humans?

Re: LLaMA2 Chat 70B outperformed ChatGPT

#62
post #13

Better evaluation paints a bit different picture: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb... *FreeWilly2 is a Llama2 70B model finetuned on an Orca style Dataset EDIT: actually, impressive: FreeWilly2 GPT-3.5 GPT-4 ARC 71.1 85.2 96.3 HellaSwag 86.4 85.5 95.3 MMLU 68.8 70.0 86.4 TruthfulQA 59.4 47.0 59.0 So reasoning (ARC) is lagging behind, but the other evaluations are at GPT-3.5 level and closi…

LLaMA2 is far and away from GPT 3.5. Just look at HumanEval and other code generation metrics. All these GPT-4 based "chat evals" are extremely misleading and people should take it with a bag of salt.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#63

Earlier quoted context omitted.

There is one: "The agreement between GPT-4 and humans reaches 85%, which is even higher than the agreement among humans (81%). This means GPT-4’s judgments closely align with the majority of humans. We also show that GPT-4’s judgments may help humans make better judgments. During our data collection, when a human’s choice deviated from GPT-4, we presented GPT-4’s judgments to humans and ask if they are reasonable. De…

It's funny how ChatGPT really does give you the most balanced, middle of the road answers. It feels like a distillation of all human knowledge and sentiments. I use it constantly to get advice on plans, architectures, thoughts, etc.. to get an idea of pretty much what the average person would think. It often points out things I've overlooked which I'll improve my design with and go back and forth with ChatGPT until w…

> to get an idea of pretty much what the average person would think

There is no such thing as an average person.

https://www.thestar.com/news/insight/when-u-s-air-force-disc...

Re: LLaMA2 Chat 70B outperformed ChatGPT

#64
post #40

Earlier quoted context omitted.

That seems more in-line with my experience. I have been using GPT-3.5 and GPT-4 for data cleaning pipelines, and have tried to swap out LLaMA2 70B in a few of the "easier" tasks, and it hasn't performed well enough yet for any of my tasks done by GPT-3.5.

Hi! Could you please share a few words on what type of data you are cleaning using GPT? It is an intriguing idea and I would love to learn more to see if I could use a similar approach.

Youtube transcripts. It only works with single-person channels at the moment, as I haven't worked on disambiguating multiple speakers. They are very messy if they are just an auto transcription from Google. Practically unusable in most cases.

So first step in the pipeline is cleaning up the transcripts for incorrectly transcribed words or sentences. Using the context of the rest of the transcript, it is able to fix the vast majority of them. Then we add punctuation and format it with paragraphs. Then I have another check over the whole transcript for any remaining issues.

After all of that, I have a relatively clean transcript that represents the original audio very closely. From there, I am doing things like: 1. creating question/answer pairs from the transcript 2. creating a document of additional context that fills in details about what the speaker is talking about but may not have explicitly said 3. creating a summary of the transcript identifying the main purpose 4. creating a knowledge graph from the transcript with nodes and edges 5. creating an annotated version of the transcript using that knowledge graph

I plan on putting some of this data into a vector database, and some of it will be used for fine tuning LLaMA2 models on specific tasks (like knowledge graph creation, annotation using a knowledge graph, and writing using a knowledge graph to keep track of events)

Re: LLaMA2 Chat 70B outperformed ChatGPT

#66

Earlier quoted context omitted.

It's funny how ChatGPT really does give you the most balanced, middle of the road answers. It feels like a distillation of all human knowledge and sentiments. I use it constantly to get advice on plans, architectures, thoughts, etc.. to get an idea of pretty much what the average person would think. It often points out things I've overlooked which I'll improve my design with and go back and forth with ChatGPT until w…

> to get an idea of pretty much what the average person would think There is no such thing as an average person. https://www.thestar.com/news/insight/when-u-s-air-force-disc...

Funnily enough I think you both might be right here, there isn't such a thing as an average person, but ChatGPT may be the synthesis of the average opinion.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#67

Earlier quoted context omitted.

It's funny how ChatGPT really does give you the most balanced, middle of the road answers. It feels like a distillation of all human knowledge and sentiments. I use it constantly to get advice on plans, architectures, thoughts, etc.. to get an idea of pretty much what the average person would think. It often points out things I've overlooked which I'll improve my design with and go back and forth with ChatGPT until w…

> to get an idea of pretty much what the average person would think There is no such thing as an average person. https://www.thestar.com/news/insight/when-u-s-air-force-disc...

The other way to phrase this is high-dimensional spheres are "spikey".

Re: LLaMA2 Chat 70B outperformed ChatGPT

#68

There is a cool website where you can blind judge the outputs from LLaMa 2 vs ChatGPT-3.5: https://llmboxing.com/ Surprisingly, LLaMa 2 won 5-0 for me.

It was much closer to me. But llama 2 did surprisingly good. It’s looks like it’s a great alternative of chatGPT 3.5.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#69
post #38
post #8

Does this mean it may be possible to self-host a ChatGPT clone assuming you have a 70B model? I've used a 13B model with LLaMA1 and it's surprisingly good, but still nowhere near ChatGPT for coding questions.

You will want to look at HumanEval ( https://github.com/abacaj/code-eval ) and Eval+ ( https://github.com/my-other-github-account/llm-humaneval-ben... ) results for coding. While Llama2 is an improvement over LLaMA v1, it's still nowhere near even the best open models (currently, sans test contamination, WizardCoder-15B, a StarCoder fine tune is at top). It's really not a competition atm though, ChatGPT-4 wipes the f…

this all numbers can be missleading, and simply indicate that gpt have these tasks in training data, and another model doesn't.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#70

*When asked by GPT4 to compare the outputs. I'm a staunch believer that it would be foolish to rely on GPT4 for quality comparisons, and it has been mind boggling to see so many people do it and treat it as perfect proof of anything. It would be slightly more understandable if there was a study to see how human and gpt4 preferences compare, but I'm unaware of any such thing.

[deleted]
Post reply on HN