Still can't handle the question: I have three doors in front of me. behind one is a great prize. Behind the other two are bad prizes. I know which door contains the prize, and I choose that door. Before I open it the game show host eliminates one of the doors that contain the bad prize. He then asks if I'd like to switch to the other remaining door instead of the one I chose. Should I switch doors?` Big answer: This…
GPT-4.5 or GPT-5 being tested on LMSYS?
281–290 of 380 posts
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#282Indeed, all my test prompts are giving much better results than gpt4-turbo and Claude Opus. And yet, the OpenAI style is clearly recognizable.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#283Earlier quoted context omitted.
i just tested this too, really cool. i own a yaris have used an online forum for yaris cars for the past decade and had a vague memory of a user who deleted some of the most helpful guides. i asked about it and sure enough it knew exactly who i meant: who's a user on yaris forums that deleted a ton of their helpful guides and how-to posts?: One notable user from the Yaris forums who deleted many of their helpful guid…
This answer looks eerily similar to the llama-3-sonar-large-32k-online model by Perplexity on labs.perplexity.ai
This is searchable now.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#284Earlier quoted context omitted.
Weird, it doesn't seem to have any info on reddit users or their writings. I tried asking about a bunch, also just about general "legendary users" from various subreddits and it seemingly just hallucinated.
Reddit may have told OpenAI to pay (probably a lot of) money to legally use Reddit content for training, which is something Reddit is doing with other AI labs ( https://www.cbsnews.com/news/google-reddit-60-million-deal-a... ); but GPTBot is not banned under the Reddit robots.txt ( https://www.reddit.com/robots.txt ). This is assuming that lmsys' GPT-2 is retained GPT-4t or a new GPT-4.5/5 though; I doubt that (one o…
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#285Earlier quoted context omitted.
You know at one point we wouldn't be able to benchmark them, due to the sheer complexity of the test required. I.e. if you are testing a model on maths, the problem will have to be extremely difficult to even consider a 'hustle' for the LLM; it would then take you a day to work out the solution yourself. See where it's getting at? When humans are no longer on the same spectrum as LLMs, that's probably the definition…
Me: 478700000000+99000000+580000+7000? GPT4: 478799650000 Me: Well? GPT4: Apologies for the confusion. The sum of 478700000000, 99000000, 580000 and 7000 is 478799058000. I will be patient. The answer is 478799587000 by the way. You just put the digits side by side.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#286Earlier quoted context omitted.
i just tested this too, really cool. i own a yaris have used an online forum for yaris cars for the past decade and had a vague memory of a user who deleted some of the most helpful guides. i asked about it and sure enough it knew exactly who i meant: who's a user on yaris forums that deleted a ton of their helpful guides and how-to posts?: One notable user from the Yaris forums who deleted many of their helpful guid…
This answer looks eerily similar to the llama-3-sonar-large-32k-online model by Perplexity on labs.perplexity.ai
Claude seems unlikely (unless it's a potential 3.5 rather than 4), since Claude-3 cutoff was 8-2023, so 11-2023 seems too soon after for next gen model.
The other candidate would be Gemini, which has an early 2023 cutoff, similar to that of GPT-4.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#287Earlier quoted context omitted.
I tried this with GPT-4 for NYC, from my address on the upper west side of Manhattan to the Brooklyn botanical gardens. It basically got the whole thing pretty much correct. I wouldn’t use it as directions, since it sometimes got left and right turns mixed up, stuff like that, but overall amazing.
That's wild. I don't understand how that's even possible with a "next token predictor" unless some weird emergence, or maybe i'm over complicating things? How does it know what the next street or neighbourhood it should traverse in each step without a pathfinding algo? Maybe there's some bus routes in the data it leans on?
The answer is always ‘turn left’ no matter where you are on park avenue. These kind of heuristics would allow you to build more general path finding based on ‘next direction’ prediction.
OpenAI may well have built a lot of synthetic direction data from some maps system like Google maps, which would then heavily train this ‘next direction’ prediction system. Google maps builds a list of smaller direction steps to follow to achieve the larger navigation goal.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#288gpt2-chatbot is not the only "mystery model" on LMSYS. Another is "deluxe-chat". When asked about it in October last year, LMSYS replied [0] "It is an experiment we are running currently. More details will be revealed later" One distinguishing feature of "deluxe-chat": although it gives high quality answers, it is very slow, so slow that the arena displays a warning whenever it is chosen as one of the competitors [0]…
> One distinguishing feature of "deluxe-chat": although it gives high quality answers, it is very slow, so slow that the arena displays a warning whenever it is chosen as one of the competitors Beam search or weird attention/non-transformer architecture?
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#289Earlier quoted context omitted.
Are you trying the paid gpt or just free 3.5 chatgpt?
100% of the time when I post a critique someone replies with this. I tell them I've used literally every LLM under the sun quite a bit to find any use I can think of and then it's immediately crickets.
What type of usages are you testing? For general knowledge it hallucinates way less often, and for reasoning and coding and modifying its past code based on English instructions it is way, way better than GPT-3 in my experience.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#290It’s an impressive model, but why would OpenAI need to do that?
Altman said in the latest Lex Friedman podcast that OAI has consistently received feedback their releases "shock the world", and that they'd like to fix that. I think releasing to this 3rd party so the internet can start chattering about it and discovering new functionality several months before an official release aligns with that goal of drip-feeding society incremental updates instead of big new releases.
OpenAI sucks at naming though. GPT2 now? Their specific gpt-4-314 etc. model naming was also a mess.