Live data from Hacker News

GPT-4.5 or GPT-5 being tested on LMSYS?

rentry.co

141–150 of 380 posts

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#142
Has anyone found that GPT3.5 was better in many ways to GPT4? I have had consistent issues with GPT4, I had it search a few spreadsheets recently looking unique numbers, not only did it not find all the numbers, but it also hallucinated numbers. Which is obviously quite bad. It also seems worse at helping to solve/fix coding issues. Only giving you vague suggestions where 3.5 would just jump right into it.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#143

The results of this LLM are consistently far better than any other that I choose. I asked ‘what is the most efficient approach to building a led grow light with off-the-shelf parts?’ and its response was incredible. Very much in line with how I’ve done it in the past after weeks of research, trial and error, and feedback from people. The other LLMs gave mostly reasonable yet sparse and incomplete answers. It also opt…

The level of competition right now is so crazy, if it ISN’T a new version of ChatGPT then OpenAI is in trouble.

GPT-4 isn’t significantly better than Llama 3. Sure, it’s better, but the margins are evaporating fast.

(I’d mention Gemini…but why bother? IMHO, this is now OpenAI the company vs Open models)

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#144
post #130

Plot twist! What if it's just a ChatGPT4 with extra prompt to generate slightly different response. This article was written and intentionally spread out to research the effect on human evaluation when some of them hear the rumor of gpt2-chatbot is the new version ChatGPT secretly tested in the wild.

I don't think a magical prompt is suddenly going to make any current public model draw an ASCII unicorn like this thing does. (besides, it already leaks a system prompt which seems very basic)

On GPT-4 release I was able to get it to draw Factorio factory layouts in ASCII which were tailored to my general instructions.

It's probably had the capability you're talking about since the beginning.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#148
post #48

I'm surprised by people's impression. I tried it in my own language and much worse than GPT-4. Of the open source LLMs I've tried, all suck in non-English. I imagine it's difficult to make an LLM work in tens of languages on a consumer computer.

There's a core problem with LLMs: they learn sentences, not facts. So an LLM may learn a ton of English-language sentences about cats, and much fewer Spanish sentences about gatos. And it even learns that cat-gato is a correct translation. But it does not ever figure out that cats and gatos are the same thing. And in particular, a true English-language fact about cats is still true if you translate it into Spanish. S…

>> These machines are just unfathomably dumb.

I agree with you, and we seem to hold a minority opinion. LLMs contain a LOT of information and are very articulate - they are language models after all. So they seem answer questions well, but fall down on thinking/reasoning about the information they contain.

But then they can play chess. I'm not sure what to make of that. Such an odd mix of capability and uselessness, but the distinction is always related to something like understanding.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#150

The results of this LLM are consistently far better than any other that I choose. I asked ‘what is the most efficient approach to building a led grow light with off-the-shelf parts?’ and its response was incredible. Very much in line with how I’ve done it in the past after weeks of research, trial and error, and feedback from people. The other LLMs gave mostly reasonable yet sparse and incomplete answers. It also opt…

The level of competition right now is so crazy, if it ISN’T a new version of ChatGPT then OpenAI is in trouble. GPT-4 isn’t significantly better than Llama 3. Sure, it’s better, but the margins are evaporating fast. (I’d mention Gemini…but why bother? IMHO, this is now OpenAI the company vs Open models)

This really is the fastest growing technology of all time. Do you feel the curve? I remember Mixtral8x7b dominating for months; I expected data bricks to do the same! but it was washed out of existence in days, with 8x22b, llama3, gemini1.5... WOW.
Post reply on HN