Live data from Hacker News

GPT-4.5 or GPT-5 being tested on LMSYS?

rentry.co

341–350 of 380 posts

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#341

Earlier quoted context omitted.

Prove it...

Have you ever heard the term NP-complete ?

Yeah, I mean, that's the joke.

The comment I replied to, "a huge class of problems that's extremely difficult to solve but very easy to check", sounded to me like an assertion that P != NP, which everyone takes for granted but actually hasn't been proved. If, contrary to all expectations, P = NP, then that huge class of problems wouldn't exist, right? Since they'd be in P, they'd actually be easy to solve as well.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#342
post #304

Earlier quoted context omitted.

Llama3 8B is for all intents and purposes just as fast.

Mistral 7b inferences about 18% faster for me as a 4bit quantized version on an A100. Thats definitely relevant when running anything but chatbots.

Are you measuring tokens/sec or words per second?

The difference matters as generally in my experience, Llama 3, by virtue of its giant vocabulary, generally tokenizes text with 20-25% less tokens than something like Mistral. So even if its 18% slower in terms of tokens/second, it may, depending on the text content, actually output a given body of text faster.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#343

Earlier quoted context omitted.

I simply asked it "what are you" and it responded that it was GPT-4 based. > I'm ChatGPT, a virtual assistant powered by artificial intelligence, specifically designed by OpenAI based on the GPT-4 model. I can help answer questions, provide explanations, generate text based on prompts, and assist with a wide range of topics. Whether you need help with information, learning something new, solving problems, or just loo…

Why would the model be self aware? There is no mechanism for the llm to know the answer to “what are you” other than training data it was fed. So it’s going to spit out whatever it was trained with, regardless of the “truth”

I agree there's no reason to believe it's self-aware (or indeed aware at all) but capabilities and origins is probably among the questions they get most, especially as the format is so inviting for anthropomorphizing and those questions are popular starters in real human conversation. It's simply due diligence in interface design to add that task to the optimization. It would be easy to mislead about if the maker wished to do that of course, but it seems plausible that it would usually be have been put truthfully as a service to the user.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#344
post #113

Earlier quoted context omitted.

This wasn't a "metric." It was a test to see whether or not this LLM might actually be useful to me. Just like every other LLM, the answer is a hard no: using this chatbot for real work is at best a huge waste of time, and at worst unconscionably reckless. For my specific question, I would have been much better off with a plain Google Scholar search.

If your everyday work consists of looking up academic citations then yeah, LLMs are not going to be useful for that - you'll get hallucinations every time. That's absolutely not a task they are useful for. There are plenty of other tasks that they ARE useful for, but you have to actively seek those out.

[deleted]

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#345

gpt2-chatbot is not the only "mystery model" on LMSYS. Another is "deluxe-chat". When asked about it in October last year, LMSYS replied [0] "It is an experiment we are running currently. More details will be revealed later" One distinguishing feature of "deluxe-chat": although it gives high quality answers, it is very slow, so slow that the arena displays a warning whenever it is chosen as one of the competitors [0]…

I looked 3 times through the list, I can't find a "deluxe-chat" there, how do I select it?

It has only ever been available in battle mode. If you do the battle enough you will probably get it eventually. I believe it is still there but doesn’t come up often, it used to be more frequent. (Battle mode is not uniformly random, some models are weighted to compete more often than others.)

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#346

Earlier quoted context omitted.

You know at one point we wouldn't be able to benchmark them, due to the sheer complexity of the test required. I.e. if you are testing a model on maths, the problem will have to be extremely difficult to even consider a 'hustle' for the LLM; it would then take you a day to work out the solution yourself. See where it's getting at? When humans are no longer on the same spectrum as LLMs, that's probably the definition…

Me: 478700000000+99000000+580000+7000? GPT4: 478799650000 Me: Well? GPT4: Apologies for the confusion. The sum of 478700000000, 99000000, 580000 and 7000 is 478799058000. I will be patient. The answer is 478799587000 by the way. You just put the digits side by side.

I recently tried a Fermi estimation problem on a bunch of LLMs and they all failed spectacularly. It was crossing too many orders of magnitude, all the zeroes muddled them up.

E.g.: the right way to work with numbers like a “trillion trillion” is to concentrate on the powers of ten, not to write the number out in full.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#347

Earlier quoted context omitted.

Usually when I encounter sentiment like this it is because they only have used 3.5 (evidently not the case here) or that their prompting is terrible/misguided. When I show a lot of people GPT4 or Claude, some percentage of them jump right to "What year did Nixon get elected?" or "How tall is Barack Obama?" and then kind of shrug with a "Yeah, Siri could do that ten years ago" take. Beyond that you have people who pro…

> or that their prompting is terrible/misguided. This is the "You're not using it right" defense. It's an LLM, it's supposed to understand human language queries. I shouldn't have to speak LLM to speak to an LLM.

People don’t even know how to use traditional web search properly.

Here’s a real scenario: A Citrix virtual desktop crashed because a recent critical security fix forced an upgrade of a shared DLL. The output is a really specific set of errors in a stack trace. I watched with my own two eyes an IT professional typed the following phrase into Google: “Why did my PC crash?”

Then he sat there and started reading through each result… including blog posts by random kids complaining about Windows XP.

I wish I could say this kind of thing is an isolated incident.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#348
post #227

Earlier quoted context omitted.

The level of competition right now is so crazy, if it ISN’T a new version of ChatGPT then OpenAI is in trouble. GPT-4 isn’t significantly better than Llama 3. Sure, it’s better, but the margins are evaporating fast. (I’d mention Gemini…but why bother? IMHO, this is now OpenAI the company vs Open models)

llama3 on groq is just stunning in its accuracy and output performance. I've already switched out some gpt4-turbo calls with it.

llama3 on groq hits the sweet spot of being so fast that I now avoid going back to waiting on gpt4 unless I really need it, and being smart enough that for 95% of the cases I won't need to.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#350
post #289

Earlier quoted context omitted.

100% of the time when I post a critique someone replies with this. I tell them I've used literally every LLM under the sun quite a bit to find any use I can think of and then it's immediately crickets.

You see no difference between non-RLHFed GPT3 from early 2022 and GPT-4 in 2024? It's a very broad consensus that there is a huge difference so that's why I wanted to clarify and make sure you were comparing the right things. What type of usages are you testing? For general knowledge it hallucinates way less often, and for reasoning and coding and modifying its past code based on English instructions it is way, way b…

I always use GPT4 to write boiler plate code etc. It probably automates 50% of my tasks, pretty good.
Post reply on HN