Earlier quoted context omitted.
Was this with or without fine-tuning?
That is with fine-tuning: https://stability.ai/blog/freewilly-large-instruction-fine-t...
LLaMA2 Chat 70B outperformed ChatGPT
131–135 of 135 posts
Re: LLaMA2 Chat 70B outperformed ChatGPT
#132*When asked by GPT4 to compare the outputs. I'm a staunch believer that it would be foolish to rely on GPT4 for quality comparisons, and it has been mind boggling to see so many people do it and treat it as perfect proof of anything. It would be slightly more understandable if there was a study to see how human and gpt4 preferences compare, but I'm unaware of any such thing.
Re: LLaMA2 Chat 70B outperformed ChatGPT
#133*When asked by GPT4 to compare the outputs. I'm a staunch believer that it would be foolish to rely on GPT4 for quality comparisons, and it has been mind boggling to see so many people do it and treat it as perfect proof of anything. It would be slightly more understandable if there was a study to see how human and gpt4 preferences compare, but I'm unaware of any such thing.
Re: LLaMA2 Chat 70B outperformed ChatGPT
#134Earlier quoted context omitted.
Ugh, when I was doing my PhD work we were studying creativity in an experiment, and we needed an assessment for how creative different solutions were, and trying to get inter-rater reliability on this quite simple thing was just agonizing. I wound up abandoning the experiment because getting enough reliability would have required screwing down the standards so tightly that it would have ruined the underlying point of…
It's surprising to me that you describe creativity as a simple thing. How were you defining and measuring it?
Re: LLaMA2 Chat 70B outperformed ChatGPT
#135Earlier quoted context omitted.
Ugh, when I was doing my PhD work we were studying creativity in an experiment, and we needed an assessment for how creative different solutions were, and trying to get inter-rater reliability on this quite simple thing was just agonizing. I wound up abandoning the experiment because getting enough reliability would have required screwing down the standards so tightly that it would have ruined the underlying point of…
Kind of hard to consistently evaluate creativity when someone might pull a James T. Kirk ( https://en.wikipedia.org/wiki/Kobayashi_Maru ), which is only creative the first time and just a cheat thereafter.