Live data from Hacker News

GPT-4.5 or GPT-5 being tested on LMSYS?

rentry.co

221–230 of 380 posts

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#221
post #175

I certainly hope it's not GPT-5. This model struggles with reasoning tasks Opus does wonderfully with. A cheaper GPT-4 that's this good? Neat, I guess. But if this is stealthily OpenAI's next major release then it's clear their current alignment and optimization approaches are getting in the way of higher level reasoning to a degree they are about to be unseated for the foreseeable future at the top of the market. (T…

To me, it seemed a bit better than GPT-4 at some coding task, or at least less inclined to just give the skeleton and leave out all the gnarly details, like GPT-4 likes to do these days. What frustrates me a bit is that I cannot really say if GPT-4, as it was in the very beginning when it happily executed even complicated and/or large requests for code, wasn't on the same level as this model actually, maybe not in terms of raw knowledge, but at least in term of usefulness/cooperativeness.

This aside, I agree with you that it does not feel like a leap, more like 4.x.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#222

Earlier quoted context omitted.

I didn't ask what you do with LLMs, I asked how you see "fastest growing technology of all time".

It strikes me as unprecedented that a technology which takes arbitrary language-based commands can actually surface and synthesize useful information, and it gets better at doing it (even according to extensive impartial benchmarking) at a fairly rapid pace. It’s technology we haven’t really seen before recently, improving quite quickly. It’s also being adopted very rapidly. I’m not saying it’s certainly the fastest…

But humans aren't 'original' ourselves. How do you do 3*9? You memorized it. It's striking how humans could reason at all.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#223

Wow. I did the arena and kept asking this: Has Anyone Really Been Far Even as Decided to Use Even Go Want to do Look More Like? All of them thought it was gibberish except gpt2-chatbot. It said: The phrase you're asking about, "Has anyone really been far even as decided to use even go want to do look more like?" is a famous example of internet gibberish that became a meme. It originated from a post on the 4chan board…

An engine that could actually translate the "harbfeadtuegwtdlml" expression into a clear one could be a good indicator. One poster in that original chain stated he understood what the original poster meant - I surely never could.

The chatbot replied to you with marginal remarks anyone could have produced (and discarded immediately through filtering), but the challenge is getting its intended meaning...

"Sorry, where am I?" // "In a car."

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#226

Earlier quoted context omitted.

There is a huge class of problems that's extremely difficult to solve but very easy to check.

Prove it...

*Assuming you don't mean mathematically prove.*

I can't test the bot right now, because it seems to have been hugged to death. But there's quite a lot of simple tests LLMs fail. Basically anything where the answer is both precise/discrete and unlikely to be directly in its training set. There's lots of examples in this [1] post, which oddly enough ended up flagged. In fact this guy [2] is offering $10k to anybody that create a prompt to get an LLM to solve a simple replacement problem he's found they fail at.

They also tend to be incapable of playing even basic level chess, in spite of there being undoubtedly millions of pages of material on the topic in their training base. If you do play, take the game out of theory ASAP (1. a3!? 2. a4!!) such that the bot can't just recite 30 moves of the ruy lopez or whatever.

[1] - https://news.ycombinator.com/item?id=39959589

[2] - https://twitter.com/VictorTaelin/status/1776677635491344744

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#227

The results of this LLM are consistently far better than any other that I choose. I asked ‘what is the most efficient approach to building a led grow light with off-the-shelf parts?’ and its response was incredible. Very much in line with how I’ve done it in the past after weeks of research, trial and error, and feedback from people. The other LLMs gave mostly reasonable yet sparse and incomplete answers. It also opt…

The level of competition right now is so crazy, if it ISN’T a new version of ChatGPT then OpenAI is in trouble. GPT-4 isn’t significantly better than Llama 3. Sure, it’s better, but the margins are evaporating fast. (I’d mention Gemini…but why bother? IMHO, this is now OpenAI the company vs Open models)

llama3 on groq is just stunning in its accuracy and output performance. I've already switched out some gpt4-turbo calls with it.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#229
It still takes the goat first, doesn't seem smarter than GPT4.

> Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three across?

Post reply on HN