Live data from Hacker News

GPT-4.5 or GPT-5 being tested on LMSYS?

rentry.co

131–140 of 380 posts

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#132

Man, its knowledge is insane. I run a dying forum. I first prompted with "Who is at ?" and it gave me a very endearing, weirdly knowledgeable bio of myself and my contributions to the forum including various innovations I made in the space back in the day. It summarized my role on my own forum better than I could have ever written it. And then I asked "who are other notable users at " and it gave me a list of some mo…

If this isn’t Google, google stock may really go down hard

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#135
post #109

I'm impressed. I gave the same prompt to opus, gpt-4, and this model. I'm very impressed with the quality. I feel like it addresses my ask better than the other 2 models. GPT2-Chatbot: https://pastebin.com/vpYvTf3T Claude: https://pastebin.com/SzNbAaKP GPT-4: https://pastebin.com/D60fjEVR Prompt: I am a senate aid, my political affliation does not matter. My goal is to once and for all fix the American healthcare sys…

They all did pretty well tbh. GPT2 didnt talk about moving away from fee for service which I think the evidence shows is the best idea, the other 2 did. GPT2 did have some other good ideas that the others didnt touch on though.

I agree, all 3 were great viable answers to my question. But Claude and GPT-4 felt a lot more like a regurgitation of the same suggestions people have proposed for years (ie, their training material). The GPT2, while similar felt more like it tried to approach the question from first principals. IE, it laid out a set of root causes, and reasoned a solution from there, which was subtly my primary ask.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#137
post #19

Prompt: code up an analog clock in html/js/css. make sure the clock is ticking exactly on the second change. second hand red. other hands black. all 12 hours marked with numbers. ChatGPT-4 Results: https://jsbin.com/giyurulajo/edit?html,css,js,output GPT2-Chatbot Results: https://jsbin.com/dacenalala/2/edit?html,css,js,output Claude3 Opus Results: https://jsbin.com/yifarinobo/edit?html,css,js,output None is correct.…

damn these bots must have dementia

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#139
post #130

Plot twist! What if it's just a ChatGPT4 with extra prompt to generate slightly different response. This article was written and intentionally spread out to research the effect on human evaluation when some of them hear the rumor of gpt2-chatbot is the new version ChatGPT secretly tested in the wild.

I don't think a magical prompt is suddenly going to make any current public model draw an ASCII unicorn like this thing does. (besides, it already leaks a system prompt which seems very basic)

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#140
gpt2-chatbot is not the only "mystery model" on LMSYS. Another is "deluxe-chat".

When asked about it in October last year, LMSYS replied [0] "It is an experiment we are running currently. More details will be revealed later"

One distinguishing feature of "deluxe-chat": although it gives high quality answers, it is very slow, so slow that the arena displays a warning whenever it is chosen as one of the competitors

[0] https://github.com/lm-sys/FastChat/issues/2527

Post reply on HN