Live data from Hacker News

LLaMA2 Chat 70B outperformed ChatGPT

tatsu-lab.github.io

131–135 of 135 posts

Re: LLaMA2 Chat 70B outperformed ChatGPT

#131
post #99

Earlier quoted context omitted.

Was this with or without fine-tuning?

That is with fine-tuning: https://stability.ai/blog/freewilly-large-instruction-fine-t...

I should be clear though, I didn't fine-tune on my specific tasks myself (as that is the end result of what I am doing with the pipelines outputs). I just tried the LLaMA2-chat model (which is fine-tuned), and the Freewilly 2 finetune.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#132

*When asked by GPT4 to compare the outputs. I'm a staunch believer that it would be foolish to rely on GPT4 for quality comparisons, and it has been mind boggling to see so many people do it and treat it as perfect proof of anything. It would be slightly more understandable if there was a study to see how human and gpt4 preferences compare, but I'm unaware of any such thing.

I agree for data creation, but evaluation seems to have little risk of contaminating the outcomes when used with human validation.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#133

*When asked by GPT4 to compare the outputs. I'm a staunch believer that it would be foolish to rely on GPT4 for quality comparisons, and it has been mind boggling to see so many people do it and treat it as perfect proof of anything. It would be slightly more understandable if there was a study to see how human and gpt4 preferences compare, but I'm unaware of any such thing.

[flagged]

Re: LLaMA2 Chat 70B outperformed ChatGPT

#134

Earlier quoted context omitted.

Ugh, when I was doing my PhD work we were studying creativity in an experiment, and we needed an assessment for how creative different solutions were, and trying to get inter-rater reliability on this quite simple thing was just agonizing. I wound up abandoning the experiment because getting enough reliability would have required screwing down the standards so tightly that it would have ruined the underlying point of…

It's surprising to me that you describe creativity as a simple thing. How were you defining and measuring it?

I don't (and didn't) think creativity was simple, but the task was super simple (alternate uses task). The fact that people couldn't agree on how creative the answers were to such a simple task was the insight, although maybe it was only an insight because I was naive.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#135

Earlier quoted context omitted.

Ugh, when I was doing my PhD work we were studying creativity in an experiment, and we needed an assessment for how creative different solutions were, and trying to get inter-rater reliability on this quite simple thing was just agonizing. I wound up abandoning the experiment because getting enough reliability would have required screwing down the standards so tightly that it would have ruined the underlying point of…

Kind of hard to consistently evaluate creativity when someone might pull a James T. Kirk ( https://en.wikipedia.org/wiki/Kobayashi_Maru ), which is only creative the first time and just a cheat thereafter.

Just one of many problems :)
Post reply on HN