Live data from Hacker News

GPT-4.5 or GPT-5 being tested on LMSYS?

rentry.co

231–240 of 380 posts

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#231
post #130

Plot twist! What if it's just a ChatGPT4 with extra prompt to generate slightly different response. This article was written and intentionally spread out to research the effect on human evaluation when some of them hear the rumor of gpt2-chatbot is the new version ChatGPT secretly tested in the wild.

I don't think a magical prompt is suddenly going to make any current public model draw an ASCII unicorn like this thing does. (besides, it already leaks a system prompt which seems very basic)

It works fine on the current chatgpt 4 turbo model https://chat.openai.com/share/d00a775d-aac7-4d14-a578-4e90af...

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#232

Earlier quoted context omitted.

I didn't ask what you do with LLMs, I asked how you see "fastest growing technology of all time".

It strikes me as unprecedented that a technology which takes arbitrary language-based commands can actually surface and synthesize useful information, and it gets better at doing it (even according to extensive impartial benchmarking) at a fairly rapid pace. It’s technology we haven’t really seen before recently, improving quite quickly. It’s also being adopted very rapidly. I’m not saying it’s certainly the fastest…

> unprecedented that a technology [...] It’s technology we haven’t really seen before recently

This is what frustrates me: First that it's not unprecedented, but second that you follow up with "haven't really" and "recently".

> fairly rapid pace ... decent case for it being a contender

Any evidence for this?

> extensive impartial benchmarking

Or this? The last two "benchmarks" I've seen that were heralded both contained an incredible gap between what was claimed and what was even proven (4 more required you to run the benchmarks even get the results!)

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#234

Earlier quoted context omitted.

I'm asking it about how to make turbine blades for a high bypass turbofan engine and it's giving very good answers, including math and some very esoteric material science knowledge. Way past the point where the knowledge can be easily checked for hallucinations without digging into literature including journal papers and using the math to build some simulations. I don't even have to prompt it much, I just keep saying…

You know at one point we wouldn't be able to benchmark them, due to the sheer complexity of the test required. I.e. if you are testing a model on maths, the problem will have to be extremely difficult to even consider a 'hustle' for the LLM; it would then take you a day to work out the solution yourself. See where it's getting at? When humans are no longer on the same spectrum as LLMs, that's probably the definition…

You know, people often complain about goal shifting in AI. We hit some target that was supposed to be AI (or even AGI), kind of go meh - and then change to a new goal. But the problem isn't goal shifting, the problem is that the goals were set to a level that had nothing whatsoever to do where we "really" want to go, precisely in order to make them achievable. So it's no surprise that when we hit these neutered goals we aren't then where we hope to actually be!

So here, with your example. Basic software programs can multiply million digit numbers near instantly with absolutely no problem. This would take a human years of dedicated effort to solve. Solving work, of any sort, that's difficult for a human has absolutely nothing to do with AGI. If we think about what we "really" mean by AGI, I think it's the exact opposite even. AGI will instead involve computers doing what's relatively easy for humans.

Go back not that long ago in our past and we were glorified monkeys. Now we're glorified monkeys with nukes and who've landed on the Moon! The point of this is that if you go back in time we basically knew nothing. State of the art technology was 'whack it with stick!', communication was limited to various grunts, and our collective knowledge was very limited, and many assumptions of fact were simply completely wrong.

Now imagine training an LLM on the state of human knowledge from this time, perhaps alongside a primitive sensory feed of the world. AGI would be able to take this and not only get to where we are today, but then go well beyond it. And this should all be able to happen at an exceptionally rapid rate, given historic human knowledge transfer and storage rates over time has always been some number really close to zero. AGI not only would not suffer such problems but would have perfect memory, orders of magnitude greater 'conscious' raw computational ability (as even a basic phone today has), and so on.

---

Is this goal achievable? No, not anytime in the foreseeable future, if ever. But people don't want this. They want to believe AGI is not only possible, but might even happen in their lifetime. But I think if we objectively think about what we "really" want to see, it's clear that it isn't coming anytime soon. Instead we're doomed to just goal shift our way endlessly towards creating what may one day be a really good natural language search engine. And hey, that's a heck of an accomplishment that will have immense utility, but it's nowhere near the goal that we "really" want.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#235
post #230

Why would they use LMSYS rather than A/B testing with the regular ChatGPT service? Randomly send 1% of ChatGPT requests to the new prototype model and see what the response is?

How do you measure the response? Also, it might be underaligned, so it is safer (from the legal point of view) to test it without formally associating it with OpenAI.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#236

Earlier quoted context omitted.

Unfortunately this is not reliable, many Non-GPT models happily claim to be GPT-4 e.g.

I simply asked it "what are you" and it responded that it was GPT-4 based. > I'm ChatGPT, a virtual assistant powered by artificial intelligence, specifically designed by OpenAI based on the GPT-4 model. I can help answer questions, provide explanations, generate text based on prompts, and assist with a wide range of topics. Whether you need help with information, learning something new, solving problems, or just loo…

It means its training data set has GPT4-generated text in it.

Yes, that's it.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#237
post #230

Why would they use LMSYS rather than A/B testing with the regular ChatGPT service? Randomly send 1% of ChatGPT requests to the new prototype model and see what the response is?

How do you measure the response? Also, it might be underaligned, so it is safer (from the legal point of view) to test it without formally associating it with OpenAI.

GPT regularly gives me A/B responses and asks me which one is better.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#238
post #96

Earlier quoted context omitted.

what's the right answer? my assumption is "not enough information"

What, you mean your fruit preferences don't form a total order?

Of course they do, but in this example there's no way to compare cherries to bananas.

Grapefruit is of course the best fruit.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#239
post #49

Prompt: my mother's sister has two brothers. each of her siblings have at least one child except for the sister that has 3 children. I have four siblings. How many grandchildren my grandfather has? Answer only with the result (the number) ChatGPT4: 13 Claude3 Opus: 10 (correct) GPT2-Chatbot: 15

It's impossible to answer. "at least one child" could mean much more than one.

Also there could be more sisters.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#240

Earlier quoted context omitted.

Prove it...

*Assuming you don't mean mathematically prove.* I can't test the bot right now, because it seems to have been hugged to death. But there's quite a lot of simple tests LLMs fail. Basically anything where the answer is both precise/discrete and unlikely to be directly in its training set. There's lots of examples in this [1] post, which oddly enough ended up flagged. In fact this guy [2] is offering $10k to anybody tha…

Multiple people found prompts to make LLM solve the problem, and the $10k has been awarded: https://twitter.com/VictorTaelin/status/1777049193489572064
Post reply on HN