Live data from Hacker News

GPT-4.5 or GPT-5 being tested on LMSYS?

rentry.co

311–320 of 380 posts

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#311

Earlier quoted context omitted.

Altman said in the latest Lex Friedman podcast that OAI has consistently received feedback their releases "shock the world", and that they'd like to fix that. I think releasing to this 3rd party so the internet can start chattering about it and discovering new functionality several months before an official release aligns with that goal of drip-feeding society incremental updates instead of big new releases.

They did the same with GPT-4, they were sitting on it for months not knowing how to release. Ended up releasing GPT-3.5 and releasing 4 quietly after nerfing 3.5 into a turbo. OpenAI sucks at naming though. GPT2 now? Their specific gpt-4-314 etc. model naming was also a mess.

> OpenAI sucks at naming though. GPT2 now?

Maybe they got help from Microsoft?

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#312

Wow. I did the arena and kept asking this: Has Anyone Really Been Far Even as Decided to Use Even Go Want to do Look More Like? All of them thought it was gibberish except gpt2-chatbot. It said: The phrase you're asking about, "Has anyone really been far even as decided to use even go want to do look more like?" is a famous example of internet gibberish that became a meme. It originated from a post on the 4chan board…

An engine that could actually translate the "harbfeadtuegwtdlml" expression into a clear one could be a good indicator. One poster in that original chain stated he understood what the original poster meant - I surely never could. The chatbot replied to you with marginal remarks anyone could have produced (and discarded immediately through filtering), but the challenge is getting its intended meaning... "Sorry, where…

Yeah, I don't expect it to be better than a human here. I'm just surprised at how much knowledge is actually stored in it. Because lots of comments here are showing it 'knowing' about very obscure things.

Though mine isn't obscure, but this is the only chatbot that actually recognizes it as a meme. Everything after that is kinda just generic descriptions of what a meme is like.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#313

Earlier quoted context omitted.

Reddit may have told OpenAI to pay (probably a lot of) money to legally use Reddit content for training, which is something Reddit is doing with other AI labs ( https://www.cbsnews.com/news/google-reddit-60-million-deal-a... ); but GPTBot is not banned under the Reddit robots.txt ( https://www.reddit.com/robots.txt ). This is assuming that lmsys' GPT-2 is retained GPT-4t or a new GPT-4.5/5 though; I doubt that (one o…

Sam Altman was on the board of reddit until recently. I don't know how these things work in SV but I wouldn't think one would go from 'partly running a company' to 'being charged for something that is probably not enforceable'. It would maybe make sense if they did pay reddit for it, because it isn't Sam's money, anyway, but for reddit to demand payment and then OpenAI to just not use the text data from reddit -- one…

That said, it is pretty SV behavior to have one of your companies pay the other. A subtle wealth transfer from OpenAI/Microsoft to Reddit (and tbh other VC backed flailing companies) would totally make sense.

VC companies for years have been parroting “data is the new oil” while burning VC money like actual oil. Crazy to think that the latest VC backed companies with even more overhyped valuations suddenly need these older ones and the data they’ve hoarded.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#314
post #163
post #109

I'm impressed. I gave the same prompt to opus, gpt-4, and this model. I'm very impressed with the quality. I feel like it addresses my ask better than the other 2 models. GPT2-Chatbot: https://pastebin.com/vpYvTf3T Claude: https://pastebin.com/SzNbAaKP GPT-4: https://pastebin.com/D60fjEVR Prompt: I am a senate aid, my political affliation does not matter. My goal is to once and for all fix the American healthcare sys…

Is that verbatim the prompt you put? You misused “aid” for “aide”, “principals” for “principles”, “countries” for “country’s”, and typo’d “affiliation”—which is all certainly fine for an internet comment, but would break the illusion of some rigorous policy discussion going on in a way that might affect our parrot friends.

Real humans misspell things, don’t be a dick

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#315

Earlier quoted context omitted.

Sam Altman was on the board of reddit until recently. I don't know how these things work in SV but I wouldn't think one would go from 'partly running a company' to 'being charged for something that is probably not enforceable'. It would maybe make sense if they did pay reddit for it, because it isn't Sam's money, anyway, but for reddit to demand payment and then OpenAI to just not use the text data from reddit -- one…

That said, it is pretty SV behavior to have one of your companies pay the other. A subtle wealth transfer from OpenAI/Microsoft to Reddit (and tbh other VC backed flailing companies) would totally make sense. VC companies for years have been parroting “data is the new oil” while burning VC money like actual oil. Crazy to think that the latest VC backed companies with even more overhyped valuations suddenly need these…

> A subtle wealth transfer from OpenAI/Microsoft to Reddit (and tbh other VC backed flailing companies) would totally make sense.

That's the confusing part -- the person I responded to posited that they didn't pay reddit and thus couldn't use the data which is the only scenario that doesn't make sense to me.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#316
Is anyone else running into problems with Cloudflare? All I see when visiting the leaderboard is a Cloudflare error page saying "Please unblock challenges.cloudflare.com to proceed."

But as far as I can tell, challenges.cloudflare.com isn't blocked by browser plugins, my home network, or my ISP.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#317
post #269

Earlier quoted context omitted.

It does seem to have more data. I asked it about some of my Github projects that don't have any stars, and it responded correctly. Wasn't able to use direct-chat, so I always chose it as the winner in battle mode! OpenAI has been crawling the web for quite a while, but how much of that data have they actually used during training? It seems like this might include all that data?

Hmm, I asked it about my GitHub project that has been out for 4 years and it got everything completely wrong.

Did you try increase the information a bit, like: can you give me more information about the GitHub project xxx written in xxx?

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#318
post #179

Sadly, still fails my test of reproducing code that implements my thesis (Dropback Continuous Pruning), which I used because it's vaguely complicated and something I know very well. It totally misses the core concept of using an PRNG and instead implements some pretty standard pruning+regrowth algo.

Can you share the prompt please? I am interested.

Sure, it's really simple -

Implement a Pytorch module for the DropBack continuous pruning while training algorithm:

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#319
post #230

Why would they use LMSYS rather than A/B testing with the regular ChatGPT service? Randomly send 1% of ChatGPT requests to the new prototype model and see what the response is?

Perhaps to give a metric to include in the announcement when (if) such a new model is released.

Re: GPT-4.5 or GPT-5 being tested on LMSYS?

#320

Earlier quoted context omitted.

It means its training data set has GPT4-generated text in it. Yes, that's it.

I think using chat gpt output to train other models is against the TOS and something they crack down on hard.

I would love to see that legal argument given their view of “fair use” of all the copyrighted material that went into OpenAI models.
Post reply on HN