Live data from Hacker News

GPT-4.5 or GPT-5 being tested on LMSYS?

rentry.co

251–260 of 380 posts

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#251

Earlier quoted context omitted.

You can still use the original gpt-4 model via the API, no? Or have they shut that down? I haven't checked lately.

If you had used GPT-4 from the beginning, the quality of the responses would have been incredibly high. It also took 3 minutes to receive a full response. And prompt engineering tricks could get you wildly different outputs to a prompt. Using the 3xx ChatGPT4 model from the API doesn't hold a candle to the responses from back then.

Ah, those were the 1 tok/sec days...

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#252

An interesting thing i've been trying is to ask for a route from A to B in some city. Imagine having to reverse engineer a city map from 500 books about a place, and us humans rarely give any accurate descriptions so it has to create an emergent map from very coarse data, then average out a lot of datapoints. I tried for various scandinavian capitals and it seems to be able to, very crudely traverse various neighbour…

Maybe related: LLM can actually learn the world map from texts https://medium.com/@fleszarjacek/large-language-models-repre...

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#253
post #175

I certainly hope it's not GPT-5. This model struggles with reasoning tasks Opus does wonderfully with. A cheaper GPT-4 that's this good? Neat, I guess. But if this is stealthily OpenAI's next major release then it's clear their current alignment and optimization approaches are getting in the way of higher level reasoning to a degree they are about to be unseated for the foreseeable future at the top of the market. (T…

Reasoning tasks, or riddles.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#255

Earlier quoted context omitted.

You can still use the original gpt-4 model via the API, no? Or have they shut that down? I haven't checked lately.

If you had used GPT-4 from the beginning, the quality of the responses would have been incredibly high. It also took 3 minutes to receive a full response. And prompt engineering tricks could get you wildly different outputs to a prompt. Using the 3xx ChatGPT4 model from the API doesn't hold a candle to the responses from back then.

Can you give an example that I can test with gpt-4-0314?

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#256
post #160

Earlier quoted context omitted.

Are you trying the paid gpt or just free 3.5 chatgpt?

100% of the time when I post a critique someone replies with this. I tell them I've used literally every LLM under the sun quite a bit to find any use I can think of and then it's immediately crickets.

RT-2 is a vision language model fine tuned on the current vision input and actuator positions as the output. Google uses a bunch of TPUs to produce a full response at a cycle rate of 3 Hz and the VLM has learned the kinematics of the robot and knows how to pick up objects according to given instructions.

Given the current rate of progress, we will have robots that can learn simple manual labor from human demonstrations (e.g. Youtube as a dataset, no I do not mean bimanual teleoperation) by the end of the decade.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#257

Fascinating results: "Write a three sentence short story in the style of Nabokov about a broken motorcycle." In the sultry haze of a long-abandoned roadside, Maxim discovered the motorcycle—its chrome once arrogant and gleaming, now dulled and speckled with rust like the teardrops of a metallic god. He caressed the handle, the cold touch returning him to that electric dawn when, with Amalia clinging breathlessly to h…

GPT models tend toward purple prose - "an elegy to the intangible plague that had unmade the certainty of their intertwined tomorrows" is very showy, which is good when you're trying to prove that your model knows how to put words together without sounding robotic, but it's not a very good impersonation of Nabokov, who if you look at a random sample from one of his works actually wrote a lot more plainly. The same wi…

Apparently much of ChatGPT's purple prose and occasional rare word usage is because it's speaking African-accented English because they used Kenyan/Nigerian workers for training.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#258
post #135

Earlier quoted context omitted.

They all did pretty well tbh. GPT2 didnt talk about moving away from fee for service which I think the evidence shows is the best idea, the other 2 did. GPT2 did have some other good ideas that the others didnt touch on though.

I agree, all 3 were great viable answers to my question. But Claude and GPT-4 felt a lot more like a regurgitation of the same suggestions people have proposed for years (ie, their training material). The GPT2, while similar felt more like it tried to approach the question from first principals. IE, it laid out a set of root causes, and reasoned a solution from there, which was subtly my primary ask.

  "- Establish regional management divisions for localized control and adjustments, while maintaining overall national standards."
This was a nice detail, applying an understanding of the local knowledge problem.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#259
post #94

Man, its knowledge is insane. I run a dying forum. I first prompted with "Who is at ?" and it gave me a very endearing, weirdly knowledgeable bio of myself and my contributions to the forum including various innovations I made in the space back in the day. It summarized my role on my own forum better than I could have ever written it. And then I asked "who are other notable users at " and it gave me a list of some mo…

i just tested this too, really cool. i own a yaris have used an online forum for yaris cars for the past decade and had a vague memory of a user who deleted some of the most helpful guides. i asked about it and sure enough it knew exactly who i meant: who's a user on yaris forums that deleted a ton of their helpful guides and how-to posts?: One notable user from the Yaris forums who deleted many of their helpful guid…

Holy crap. Even if this is RAG-based, this is insanely good.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#260

Earlier quoted context omitted.

You know, people often complain about goal shifting in AI. We hit some target that was supposed to be AI (or even AGI), kind of go meh - and then change to a new goal. But the problem isn't goal shifting, the problem is that the goals were set to a level that had nothing whatsoever to do where we "really" want to go, precisely in order to make them achievable. So it's no surprise that when we hit these neutered goals…

There are different shades of AGI, but we don’t know if they will happen all at once or not. For example, if an AI can replace the average white collar worker and therefore cause massive economic disruption, that would be a shade of AGI. Another shade of AGI would be an AI that can effectively do research level mathematics and theoretical physics and is therefore capable of very high-level logical reasoning. We don’t…

At one time it was thought that software that could beat a human at chess would be, in your lingo, "a shade of AGI." And for the same reason you're listing your milestones - because it sounded extremely difficult and complex. Of course now we realize that was quite silly. You can develop software that can crush even the strongest humans through relatively simple algorithmic processes.

And I think this is the trap we need to avoid falling into. Complexity and intelligence are not inherently linked in any way. Primitive humans did not solve complex problems, yet obviously were highly intelligent. And so, to me, the great milestones are not some complex problem or another, but instead achieving success in things that have no clear path towards them. For instance, many (if not most) primitive tribes today don't even have the concept of numbers. Instead they rely on, if anything, broad concepts like a few, a lot, and more than a lot.

Think about what an unprecedented and giant leap is to go from that to actually quantifying things and imagining relationships and operations. If somebody did try to do this, he would initially just look like a fool. Yes here is one rock, and here is another. Yes you have "two" now. So what? That's a leap that has no clear guidance or path towards it. All of the problems that mathematics solve don't even exist until you discover it! So you're left with something that is not just a recombination or stair step from where you currently are, but something entirely outside what you know. That we are not only capable of such achievements, but repeatedly achieve such is, to me, perhaps the purest benchmark for general intelligence.

So if we were actually interested in pursuing AGI, it would seem that such achievements would also be dramatically easier (and cheaper) to test for. Because you need not train on petabytes of data, because the quantifiable knowledge of these peoples is nowhere even remotely close to that. And the goal is to create systems that get from that extremely limited domain of input, to what comes next, without expressly being directed to do so.

Post reply on HN