GPT-4.5 or GPT-5 being tested on LMSYS?
141–150 of 380 posts
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#142Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#143The results of this LLM are consistently far better than any other that I choose. I asked ‘what is the most efficient approach to building a led grow light with off-the-shelf parts?’ and its response was incredible. Very much in line with how I’ve done it in the past after weeks of research, trial and error, and feedback from people. The other LLMs gave mostly reasonable yet sparse and incomplete answers. It also opt…
GPT-4 isn’t significantly better than Llama 3. Sure, it’s better, but the margins are evaporating fast.
(I’d mention Gemini…but why bother? IMHO, this is now OpenAI the company vs Open models)
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#144Plot twist! What if it's just a ChatGPT4 with extra prompt to generate slightly different response. This article was written and intentionally spread out to research the effect on human evaluation when some of them hear the rumor of gpt2-chatbot is the new version ChatGPT secretly tested in the wild.
I don't think a magical prompt is suddenly going to make any current public model draw an ASCII unicorn like this thing does. (besides, it already leaks a system prompt which seems very basic)
It's probably had the capability you're talking about since the beginning.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#145Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#146Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#147What is it about LLMs that brings otherwise rational people to become bumbling sycophants??
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#148I'm surprised by people's impression. I tried it in my own language and much worse than GPT-4. Of the open source LLMs I've tried, all suck in non-English. I imagine it's difficult to make an LLM work in tens of languages on a consumer computer.
There's a core problem with LLMs: they learn sentences, not facts. So an LLM may learn a ton of English-language sentences about cats, and much fewer Spanish sentences about gatos. And it even learns that cat-gato is a correct translation. But it does not ever figure out that cats and gatos are the same thing. And in particular, a true English-language fact about cats is still true if you translate it into Spanish. S…
I agree with you, and we seem to hold a minority opinion. LLMs contain a LOT of information and are very articulate - they are language models after all. So they seem answer questions well, but fall down on thinking/reasoning about the information they contain.
But then they can play chess. I'm not sure what to make of that. Such an odd mix of capability and uselessness, but the distinction is always related to something like understanding.
Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#149Re: GPT-4.5 or GPT-5 being tested on LMSYS?
#150The results of this LLM are consistently far better than any other that I choose. I asked ‘what is the most efficient approach to building a led grow light with off-the-shelf parts?’ and its response was incredible. Very much in line with how I’ve done it in the past after weeks of research, trial and error, and feedback from people. The other LLMs gave mostly reasonable yet sparse and incomplete answers. It also opt…
The level of competition right now is so crazy, if it ISN’T a new version of ChatGPT then OpenAI is in trouble. GPT-4 isn’t significantly better than Llama 3. Sure, it’s better, but the margins are evaporating fast. (I’d mention Gemini…but why bother? IMHO, this is now OpenAI the company vs Open models)