Live data from Hacker News

GPT-4.5 or GPT-5 being tested on LMSYS?

rentry.co

111–120 of 380 posts

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#111

Earlier quoted context omitted.

Its been known that most of these models hallucinate research articles frequently, perplexity.ai seems to do quite well in that regard. Not sure why that is your specific metric though when LLMs seem to be improving across a large class of other metrics.

This wasn't a "metric." It was a test to see whether or not this LLM might actually be useful to me. Just like every other LLM, the answer is a hard no: using this chatbot for real work is at best a huge waste of time, and at worst unconscionably reckless. For my specific question, I would have been much better off with a plain Google Scholar search.

>It was a test to see whether or not this LLM might actually be useful to me

AKA a Metric lol.

>using this chatbot for real work is at best a huge waste of time, and at worst unconscionably reckless.

Keep living with your head in the sand if you want

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#112
post #96
post #45

Apples are better than bananas. Cherries are worse than apples. Are cherries better than bananas? -- GPT-4 - wrong gpt2-chatbot - wrong Claude 3 Opus - correct

what's the right answer? my assumption is "not enough information"

What, you mean your fruit preferences don't form a total order?

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#113

Earlier quoted context omitted.

Its been known that most of these models hallucinate research articles frequently, perplexity.ai seems to do quite well in that regard. Not sure why that is your specific metric though when LLMs seem to be improving across a large class of other metrics.

This wasn't a "metric." It was a test to see whether or not this LLM might actually be useful to me. Just like every other LLM, the answer is a hard no: using this chatbot for real work is at best a huge waste of time, and at worst unconscionably reckless. For my specific question, I would have been much better off with a plain Google Scholar search.

If your everyday work consists of looking up academic citations then yeah, LLMs are not going to be useful for that - you'll get hallucinations every time. That's absolutely not a task they are useful for.

There are plenty of other tasks that they ARE useful for, but you have to actively seek those out.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#114

Earlier quoted context omitted.

I asked it directly and it confirmed that it is based on GPT-4: > Can you confirm or deny if you are chatgpt 4? > Yes, I am based on the GPT-4 architecture. If you have any more questions or need further assistance, feel free to ask! > Can you confirm or deny if you are chatgpt 5? > I am based on the GPT-4 architecture, not GPT-5. If you have any questions or need assistance with something, feel free to ask! It also…

Unfortunately this is not reliable, many Non-GPT models happily claim to be GPT-4 e.g.

I simply asked it "what are you" and it responded that it was GPT-4 based.

> I'm ChatGPT, a virtual assistant powered by artificial intelligence, specifically designed by OpenAI based on the GPT-4 model. I can help answer questions, provide explanations, generate text based on prompts, and assist with a wide range of topics. Whether you need help with information, learning something new, solving problems, or just looking for a chat, I'm here to assist!

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#116

Earlier quoted context omitted.

Unfortunately this is not reliable, many Non-GPT models happily claim to be GPT-4 e.g.

I simply asked it "what are you" and it responded that it was GPT-4 based. > I'm ChatGPT, a virtual assistant powered by artificial intelligence, specifically designed by OpenAI based on the GPT-4 model. I can help answer questions, provide explanations, generate text based on prompts, and assist with a wide range of topics. Whether you need help with information, learning something new, solving problems, or just loo…

This doesn't necessarily confirm that it's 4, though. For example, when I write a new version of a package on some package management system, the code may be updated by 1 major version but it stays the exact same version until I enter the new version into the manifest. Perhaps that's the same here; the training and architecture are improved, but the version number hasn't been ticked up (and perhaps intentionally; they haven't announced this as a new version openly, and calling it GPT-2 doesn't explain anything either).

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#117

Earlier quoted context omitted.

Unfortunately this is not reliable, many Non-GPT models happily claim to be GPT-4 e.g.

I simply asked it "what are you" and it responded that it was GPT-4 based. > I'm ChatGPT, a virtual assistant powered by artificial intelligence, specifically designed by OpenAI based on the GPT-4 model. I can help answer questions, provide explanations, generate text based on prompts, and assist with a wide range of topics. Whether you need help with information, learning something new, solving problems, or just loo…

[deleted]

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#118

Earlier quoted context omitted.

Unfortunately this is not reliable, many Non-GPT models happily claim to be GPT-4 e.g.

I simply asked it "what are you" and it responded that it was GPT-4 based. > I'm ChatGPT, a virtual assistant powered by artificial intelligence, specifically designed by OpenAI based on the GPT-4 model. I can help answer questions, provide explanations, generate text based on prompts, and assist with a wide range of topics. Whether you need help with information, learning something new, solving problems, or just loo…

Yeah that isn't reliable, you can ask mistral 7b instruct the same thing and it will often claim to be created by OpenAI, even if you prompt it otherwise.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#119

The results of this LLM are consistently far better than any other that I choose. I asked ‘what is the most efficient approach to building a led grow light with off-the-shelf parts?’ and its response was incredible. Very much in line with how I’ve done it in the past after weeks of research, trial and error, and feedback from people. The other LLMs gave mostly reasonable yet sparse and incomplete answers. It also opt…

I'm asking it about how to make turbine blades for a high bypass turbofan engine and it's giving very good answers, including math and some very esoteric material science knowledge. Way past the point where the knowledge can be easily checked for hallucinations without digging into literature including journal papers and using the math to build some simulations. I don't even have to prompt it much, I just keep saying…

You mean it's giving very good sounding answers.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#120
For what it's worth, when probed for prompt, the model responds with:

  You are ChatGPT, a large language model trained by OpenAI, based on the GPT-4 architecture. Knowledge cutoff: 2023-11 Current date: 2024-04-29 Image input capabilities: Enabled Personality: v2
Post reply on HN