Curious if anyone has done a side-by-side analysis of this offering vs just running LLaMA? I'm currently running a side-by-side comparison/evaluation of MSFT GPT via Cognitive Services vs LLaMA[7B/13B/70B] and intrigued by the possibility of a truly air-gapped offering not limited by external computer power (nor by metered fees racking up.) Any reads on comparisons would be nice to see. (yes, I realize we'll eventual…
GPT-4, Bard and Claude 2 came out on top.
Llama 2 70b chat scored similarly to GPT-3.5, though GPT-3.5 still seemed to perform a bit better overall.
My personal takeaway is I’m going to continue using GPT-4 for everything where the cost and response time are workable.
Related: A belief I have is that LLM benchmarks are all too research oriented. That made sense when LLMs were in the lab. It doesn't make sense now that LLMs have tens of millions of DAUs — i.e. ChatGPT. The biggest use cases for LLMs so far are chat assistants and programming assistants. We need benchmarks that are based on the way people use LLMs in chatbots and the type of questions that real users use LLM products, not hypothetical benchmarks and random academic tests.