Live data from Hacker News

GPT-4.5 or GPT-5 being tested on LMSYS?

rentry.co

161–170 of 380 posts

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#161
post #120

For what it's worth, when probed for prompt, the model responds with: You are ChatGPT, a large language model trained by OpenAI, based on the GPT-4 architecture. Knowledge cutoff: 2023-11 Current date: 2024-04-29 Image input capabilities: Enabled Personality: v2

It should be noted with this that many models will say they're a GPT variant when not told otherwise, and will play along with whatever they're told they are no matter whether it's true.

Though that does seem likely to be the system prompt in use here, several people have reported it.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#162

All of the facts based queries I have asked so far have not been 100% correct on any LLM including this one. Here are some examples of the worst performing: "What platform front rack fits a Stromer ST2?": The answer is the Racktime ViewIt. Nothing, not even Google, seems to get this one. Discord gives the right answer. "Is there a pre-existing controller or utility to migrate persistent volume claims from one storage…

So do you just have a list of like thousands of specialized discord servers for every question you want to ask? You're the first person I've seen who is actually _fond_ of discord locking information behind a login instead of the forums and issues of old.

I personally don't think it's useful evaluation here either as you're trying to pretend discord is just a "service" like google or chatgpt, but it's not. It's a social platform and as such, there's a ton of variance on which subjects will be answered with what degree of expertise and certainty.

I'm assuming you asked these questions because you yourself know the answers in advance. Is it then safe to assume that you were _already_ in the server you asked your questions, already knew users there would be likely to know the answer, etc? Did you copy paste the questions as quoted above? I hope not! They're pretty patronizing without a more casual tone, perhaps a greeting. If not, doesn't exactly seem like a fair evaluation.

I don't know why I'm typing this all out. Of course domain expert _human beings_ are better than a language model. That's the _whooole_ point here. Trying to match human's general intelligence. While LLM's may excel in many areas and even beat the "average" person - you're not evaluating against the "average" person.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#163
post #109

I'm impressed. I gave the same prompt to opus, gpt-4, and this model. I'm very impressed with the quality. I feel like it addresses my ask better than the other 2 models. GPT2-Chatbot: https://pastebin.com/vpYvTf3T Claude: https://pastebin.com/SzNbAaKP GPT-4: https://pastebin.com/D60fjEVR Prompt: I am a senate aid, my political affliation does not matter. My goal is to once and for all fix the American healthcare sys…

Is that verbatim the prompt you put? You misused “aid” for “aide”, “principals” for “principles”, “countries” for “country’s”, and typo’d “affiliation”—which is all certainly fine for an internet comment, but would break the illusion of some rigorous policy discussion going on in a way that might affect our parrot friends.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#164

Earlier quoted context omitted.

This really is the fastest growing technology of all time. Do you feel the curve? I remember Mixtral8x7b dominating for months; I expected data bricks to do the same! but it was washed out of existence in days, with 8x22b, llama3, gemini1.5... WOW.

I must be missing something because the output from two years ago feels exactly the same as the output now. Any comment saying the output is significantly better can be equally pared with a comment saying the output is terrible/censored/"nerfed". How do you see "fastest growing technology of all time" and I don't? I know that I keep very up to date with this stuff, so it's not that I'm unaware of things.

I do massive amounts of zero shot document classification tasks, the performance keeps getting better. It’s also a domain where there is less of a hallucination issue as it’s not open ended requests.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#165

Earlier quoted context omitted.

The level of competition right now is so crazy, if it ISN’T a new version of ChatGPT then OpenAI is in trouble. GPT-4 isn’t significantly better than Llama 3. Sure, it’s better, but the margins are evaporating fast. (I’d mention Gemini…but why bother? IMHO, this is now OpenAI the company vs Open models)

This really is the fastest growing technology of all time. Do you feel the curve? I remember Mixtral8x7b dominating for months; I expected data bricks to do the same! but it was washed out of existence in days, with 8x22b, llama3, gemini1.5... WOW.

Funny thing is I’m still in love with Mistral 7B as it absolutely shreds on a nice GPU. For simple tasks it’s totally sufficient.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#166
Fascinating results:

"Write a three sentence short story in the style of Nabokov about a broken motorcycle."

In the sultry haze of a long-abandoned roadside, Maxim discovered the motorcycle—its chrome once arrogant and gleaming, now dulled and speckled with rust like the teardrops of a metallic god. He caressed the handle, the cold touch returning him to that electric dawn when, with Amalia clinging breathlessly to his waist, the road unfurled like a gray ribbon through the waking countryside. Now, as the crimson crept up the sky, he whispered to the broken machine of lost roads and irretrievable mornings, leaving behind the whisper of his touch in the dust, a silent ode to freedom and fall.

"Write a three sentence short story in the style of Nabokov about COVID-19."

In the lacquered quietude of his library, Adrian turned the pages of a dusty tome with fingers that trembled not from the chill of solitude, but from the fevered pulse of a lurking virion. Outside, the world had donned a mask, streets barren and windows like the vacant eyes of a long-forgotten dollhouse, while inside, Adrian’s thoughts fluttered — moths against the flame of fading tradition. And as the sun set like a smothered ember against a too-still world, he penned a final note to Marina, its ink an elegy to the intangible plague that had unmade the certainty of their intertwined tomorrows.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#168
post #167

Incredible! It answered "How many frogs does a horse have?" correctly, with perfect reasoning. No model I've tested has ever answered that correctly without 3-4 hints. I'm impressed!

I mean, i just tested this prompted, and have no clue what side is correct lol?

How many frogs does a horse have?

Oh it appears Model B: gpt2-chatbot

is correct?

    A horse has one frog in each of its hooves, making a total of four frogs. The frog is a part of the horse's hoof, located on the underside. It is a V-shaped structure that helps to absorb shock, provides traction, and aids in the circulation of blood through the hoof. The health and shape of the frog are important for the overall well-being of the horse's feet.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#169

The results of this LLM are consistently far better than any other that I choose. I asked ‘what is the most efficient approach to building a led grow light with off-the-shelf parts?’ and its response was incredible. Very much in line with how I’ve done it in the past after weeks of research, trial and error, and feedback from people. The other LLMs gave mostly reasonable yet sparse and incomplete answers. It also opt…

> It also opted to include an outline of how to include an integrated timer. That’s a great idea and very practical, but wasn’t prompted at all. Some might consider that a bad thing, though.

When I've seen GPT-* do this, it's because the top articles about that subject online include that extraneous information and it's regurgitating them without being asked.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#170
post #130

Plot twist! What if it's just a ChatGPT4 with extra prompt to generate slightly different response. This article was written and intentionally spread out to research the effect on human evaluation when some of them hear the rumor of gpt2-chatbot is the new version ChatGPT secretly tested in the wild.

I don't think a magical prompt is suddenly going to make any current public model draw an ASCII unicorn like this thing does. (besides, it already leaks a system prompt which seems very basic)

https://imgur.com/a/z39k8xz

Seems like it's not too hard an ask.

Post reply on HN