Live data from Hacker News

GPT-4.5 or GPT-5 being tested on LMSYS?

rentry.co

281–290 of 380 posts

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#281

Still can't handle the question: I have three doors in front of me. behind one is a great prize. Behind the other two are bad prizes. I know which door contains the prize, and I choose that door. Before I open it the game show host eliminates one of the doors that contain the bad prize. He then asks if I'd like to switch to the other remaining door instead of the one I chose. Should I switch doors?` Big answer: This…

Your question doesn’t make sense if you read it directly, why are you asking which door to pick if you already know the door? It’s what you call a “trick” question, something humans are also bad at. It’s equally plausible (and arguably more useful for general purposes) for the model to assume that you mistyped and for eg. forgot to type “don’t” between I and know.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#282
post #6

Indeed, all my test prompts are giving much better results than gpt4-turbo and Claude Opus. And yet, the OpenAI style is clearly recognizable.

Agreed. I had it solve a little programming problem in a really obscure programming language and after some prompt tuning got results strongly superior to GPT4, Claude, Llama3 and Mixtral. As the language (which I won't name here) is acceptably documented but there are _really_ few examples available online, this seems to indicate very good generalization and reasoning capabilities.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#283
post #94

Earlier quoted context omitted.

i just tested this too, really cool. i own a yaris have used an online forum for yaris cars for the past decade and had a vague memory of a user who deleted some of the most helpful guides. i asked about it and sure enough it knew exactly who i meant: who's a user on yaris forums that deleted a ton of their helpful guides and how-to posts?: One notable user from the Yaris forums who deleted many of their helpful guid…

This answer looks eerily similar to the llama-3-sonar-large-32k-online model by Perplexity on labs.perplexity.ai

https://www.google.com/search?q=who%27s+a+user+on+yaris+foru...

This is searchable now.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#284

Earlier quoted context omitted.

Weird, it doesn't seem to have any info on reddit users or their writings. I tried asking about a bunch, also just about general "legendary users" from various subreddits and it seemingly just hallucinated.

Reddit may have told OpenAI to pay (probably a lot of) money to legally use Reddit content for training, which is something Reddit is doing with other AI labs ( https://www.cbsnews.com/news/google-reddit-60-million-deal-a... ); but GPTBot is not banned under the Reddit robots.txt ( https://www.reddit.com/robots.txt ). This is assuming that lmsys' GPT-2 is retained GPT-4t or a new GPT-4.5/5 though; I doubt that (one o…

Sam Altman was on the board of reddit until recently. I don't know how these things work in SV but I wouldn't think one would go from 'partly running a company' to 'being charged for something that is probably not enforceable'. It would maybe make sense if they did pay reddit for it, because it isn't Sam's money, anyway, but for reddit to demand payment and then OpenAI to just not use the text data from reddit -- one of the largest sources of good quality conversational training data available -- strikes me as odd. But nothing would surprise me when it comes to this market.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#285

Earlier quoted context omitted.

You know at one point we wouldn't be able to benchmark them, due to the sheer complexity of the test required. I.e. if you are testing a model on maths, the problem will have to be extremely difficult to even consider a 'hustle' for the LLM; it would then take you a day to work out the solution yourself. See where it's getting at? When humans are no longer on the same spectrum as LLMs, that's probably the definition…

Me: 478700000000+99000000+580000+7000? GPT4: 478799650000 Me: Well? GPT4: Apologies for the confusion. The sum of 478700000000, 99000000, 580000 and 7000 is 478799058000. I will be patient. The answer is 478799587000 by the way. You just put the digits side by side.

Works for me https://chat.openai.com/share/51c55c5e-9bb2-4b8c-afee-87c032...

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#286
post #94

Earlier quoted context omitted.

i just tested this too, really cool. i own a yaris have used an online forum for yaris cars for the past decade and had a vague memory of a user who deleted some of the most helpful guides. i asked about it and sure enough it knew exactly who i meant: who's a user on yaris forums that deleted a ton of their helpful guides and how-to posts?: One notable user from the Yaris forums who deleted many of their helpful guid…

This answer looks eerily similar to the llama-3-sonar-large-32k-online model by Perplexity on labs.perplexity.ai

Based on the 11-2023 knowledge cutoff date, I have to wonder if it might be Llama 3 400B rather than GPT-N. Llama 3 70B cutoff was 12-2023 (8B was 3-2023).

Claude seems unlikely (unless it's a potential 3.5 rather than 4), since Claude-3 cutoff was 8-2023, so 11-2023 seems too soon after for next gen model.

The other candidate would be Gemini, which has an early 2023 cutoff, similar to that of GPT-4.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#287

Earlier quoted context omitted.

I tried this with GPT-4 for NYC, from my address on the upper west side of Manhattan to the Brooklyn botanical gardens. It basically got the whole thing pretty much correct. I wouldn’t use it as directions, since it sometimes got left and right turns mixed up, stuff like that, but overall amazing.

That's wild. I don't understand how that's even possible with a "next token predictor" unless some weird emergence, or maybe i'm over complicating things? How does it know what the next street or neighbourhood it should traverse in each step without a pathfinding algo? Maybe there's some bus routes in the data it leans on?

One thing you could think about is the very simple idea of “I’m walking north on park avenue, how do I get into the park”?

The answer is always ‘turn left’ no matter where you are on park avenue. These kind of heuristics would allow you to build more general path finding based on ‘next direction’ prediction.

OpenAI may well have built a lot of synthetic direction data from some maps system like Google maps, which would then heavily train this ‘next direction’ prediction system. Google maps builds a list of smaller direction steps to follow to achieve the larger navigation goal.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#288

gpt2-chatbot is not the only "mystery model" on LMSYS. Another is "deluxe-chat". When asked about it in October last year, LMSYS replied [0] "It is an experiment we are running currently. More details will be revealed later" One distinguishing feature of "deluxe-chat": although it gives high quality answers, it is very slow, so slow that the arena displays a warning whenever it is chosen as one of the competitors [0]…

> One distinguishing feature of "deluxe-chat": although it gives high quality answers, it is very slow, so slow that the arena displays a warning whenever it is chosen as one of the competitors Beam search or weird attention/non-transformer architecture?

It has a room full of scientists typing out the answers by hand.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#289
post #160

Earlier quoted context omitted.

Are you trying the paid gpt or just free 3.5 chatgpt?

100% of the time when I post a critique someone replies with this. I tell them I've used literally every LLM under the sun quite a bit to find any use I can think of and then it's immediately crickets.

You see no difference between non-RLHFed GPT3 from early 2022 and GPT-4 in 2024? It's a very broad consensus that there is a huge difference so that's why I wanted to clarify and make sure you were comparing the right things.

What type of usages are you testing? For general knowledge it hallucinates way less often, and for reasoning and coding and modifying its past code based on English instructions it is way, way better than GPT-3 in my experience.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#290
post #7

It’s an impressive model, but why would OpenAI need to do that?

Altman said in the latest Lex Friedman podcast that OAI has consistently received feedback their releases "shock the world", and that they'd like to fix that. I think releasing to this 3rd party so the internet can start chattering about it and discovering new functionality several months before an official release aligns with that goal of drip-feeding society incremental updates instead of big new releases.

They did the same with GPT-4, they were sitting on it for months not knowing how to release. Ended up releasing GPT-3.5 and releasing 4 quietly after nerfing 3.5 into a turbo.

OpenAI sucks at naming though. GPT2 now? Their specific gpt-4-314 etc. model naming was also a mess.

Post reply on HN