Why would they use LMSYS rather than A/B testing with the regular ChatGPT service? Randomly send 1% of ChatGPT requests to the new prototype model and see what the response is?
GPT-4.5 or GPT-5 being tested on LMSYS?
261–270 of 380 posts
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#262Prompt: code up an analog clock in html/js/css. make sure the clock is ticking exactly on the second change. second hand red. other hands black. all 12 hours marked with numbers. ChatGPT-4 Results: https://jsbin.com/giyurulajo/edit?html,css,js,output GPT2-Chatbot Results: https://jsbin.com/dacenalala/2/edit?html,css,js,output Claude3 Opus Results: https://jsbin.com/yifarinobo/edit?html,css,js,output None is correct.…
To be fair, this is pretty hard. Imagine you had to do to sit down and write this without being able to test it.
Here is the fix for the "GPT-2" version:
.hand {
top: 48.5%;
}
.hour-hand {
left: 20%;
}
.minute-hand {
left: 10%;
}
.second-hand {
left: 5%;
}Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#263I certainly hope it's not GPT-5. This model struggles with reasoning tasks Opus does wonderfully with. A cheaper GPT-4 that's this good? Neat, I guess. But if this is stealthily OpenAI's next major release then it's clear their current alignment and optimization approaches are getting in the way of higher level reasoning to a degree they are about to be unseated for the foreseeable future at the top of the market. (T…
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#264Earlier quoted context omitted.
At this moment, there's no real world benchmark at scale other than lmsys. All other "benchmarks" are merely sanity checks.
OpenAI could either hire private testers or use AB testing on ChatGPT Plus users (for example, oftentimes, when using ChatGPT, I have to select between 2 different responses to continue a conversation); both are probably much more better (in many aspects: not leaking GPT-4.5/5 generations (or the existence of a GPT-4.5/5) to the public at scale and avoiding bias* (because people probably rate GPT-4 generations better…
The only case where detecting a model makes any difference is for vendors who want to boost their own model by hiring people and paying them every time they select the vendor's model.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#265My very first response from gpt2-chatbot included a fictional source :( > A study by Lucon-Xiccato et al. (2020) tested African clawed frogs (Xenopus laevis) and found that they could discriminate between two groups of objects differing in number (1 vs. 2, 2 vs. 3, and 3 vs. 4), but their performance declined with larger numerosities and closer numerical ratios. It appears to be referring to this[1] 2018 study from t…
Its been known that most of these models hallucinate research articles frequently, perplexity.ai seems to do quite well in that regard. Not sure why that is your specific metric though when LLMs seem to be improving across a large class of other metrics.
The data from the "AI Search Engine Multilingual Evaluation Report (v1.0) | Search.Glarity.ai" indicates that generative search engines have a long road ahead in terms of exploration, which I find to be of significant importance.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#266Earlier quoted context omitted.
Weird, it doesn't seem to have any info on reddit users or their writings. I tried asking about a bunch, also just about general "legendary users" from various subreddits and it seemingly just hallucinated.
Reddit may have told OpenAI to pay (probably a lot of) money to legally use Reddit content for training, which is something Reddit is doing with other AI labs ( https://www.cbsnews.com/news/google-reddit-60-million-deal-a... ); but GPTBot is not banned under the Reddit robots.txt ( https://www.reddit.com/robots.txt ). This is assuming that lmsys' GPT-2 is retained GPT-4t or a new GPT-4.5/5 though; I doubt that (one o…
One note, its name is not gpt-2 it is gpt2 which could indicate its a "second version" of the previous gpt architecture, gpt-3, gpt-4 being gpt1-3, gpt1-4. I am just speculating and am not an expert whatsoever this could be total bullshit.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#267Earlier quoted context omitted.
How do you measure the response? Also, it might be underaligned, so it is safer (from the legal point of view) to test it without formally associating it with OpenAI.
GPT regularly gives me A/B responses and asks me which one is better.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#268Earlier quoted context omitted.
It strikes me as unprecedented that a technology which takes arbitrary language-based commands can actually surface and synthesize useful information, and it gets better at doing it (even according to extensive impartial benchmarking) at a fairly rapid pace. It’s technology we haven’t really seen before recently, improving quite quickly. It’s also being adopted very rapidly. I’m not saying it’s certainly the fastest…
But humans aren't 'original' ourselves. How do you do 3*9? You memorized it. It's striking how humans could reason at all.
I put my hands out, count to the third finger from the left, and put that finger down. I then count the fingers to the left (2) and count the fingers to the right (2 + hand aka 5) and conclude 27.
I have memorised the technique, but I definitely never memorised my nine times table. If you’d said ‘6’, then the answer would be different, as I’d actually have to sing a song to get to the answer.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#269Man, its knowledge is insane. I run a dying forum. I first prompted with "Who is at ?" and it gave me a very endearing, weirdly knowledgeable bio of myself and my contributions to the forum including various innovations I made in the space back in the day. It summarized my role on my own forum better than I could have ever written it. And then I asked "who are other notable users at " and it gave me a list of some mo…
OpenAI has been crawling the web for quite a while, but how much of that data have they actually used during training? It seems like this might include all that data?
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#270Earlier quoted context omitted.
There are different shades of AGI, but we don’t know if they will happen all at once or not. For example, if an AI can replace the average white collar worker and therefore cause massive economic disruption, that would be a shade of AGI. Another shade of AGI would be an AI that can effectively do research level mathematics and theoretical physics and is therefore capable of very high-level logical reasoning. We don’t…
At one time it was thought that software that could beat a human at chess would be, in your lingo, "a shade of AGI." And for the same reason you're listing your milestones - because it sounded extremely difficult and complex. Of course now we realize that was quite silly. You can develop software that can crush even the strongest humans through relatively simple algorithmic processes. And I think this is the trap we…
It’s hard for people to define AGI because Earth only has one generally intelligent family: Homo. So there is a tendency to identify Human intelligence or capabilities with General intelligence.
Imagine if dolphins were much more intelligent and could write research-level mathematics papers on par with humans, communicating with clicks. Even though dolphins can’t play the cello or do origami, lacking the requisite digits, UCLA still has a dolphin tank to house some of their mathematics professors, who work hand-in-flipper with their human counterparts. That’s General intelligence.
Artificial General Intelligence is the same but with a computer instead of a dolphin.