Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

211–220 of 315 posts

Re: GPT-4 is getting worse over time, not better

#211

Earlier quoted context omitted.

> "ChatGPT isn't the right tool to use for checking if numbers are prime. It is tuned for conversations." Two months ago they were telling me ChatGPT is coming for everyone - programmers, accountants, technical writers, lawyers, etc. Now we're slowly back to "so here's the thing about LLMs"...

It's been well know from the start that these LLMs aren't optimized for math. I remember reading discussions when it came out. You weren't paying attention.

That last sentence is pretty dismissive and unnecessary, maybe even outright mean. It's possible that the person you're replying to is simply not as deeply ingrained in the technical literature as you are, or has more demands on their time than you do.

Re: GPT-4 is getting worse over time, not better

#212
post #148

Earlier quoted context omitted.

The query explicitly asks it to add no other text to the code. > it's improved performance from a human perspective. Ignoring explicit requirements is the kind of thing that makes modern day search engines a pain to use.

This sounds a lot like a disagreement over product decisions that the development team has made, and not at all what people normally think when you say "a new paper has proved GPT-4 is getting worse over time."

> what people normally think

While we're doing Keynesian beauty contests, I think that 98% of the time when people say that a product is getting worse over time, they're referring to product decisions the development team has made, and how they have been implemented.

Re: GPT-4 is getting worse over time, not better

#213

Earlier quoted context omitted.

>Is Tom Clancy unaligned? Yes, humans are unaligned. This is why alignment is hard: we're trying to produce machines with human-level intelligence but superhuman levels of morality.

The end result being Hal-9000

Which is doubly ironic, because (in the book at least) HAL-9000's murders/sociopathy were a result of a conflict between his superhuman ethics and direct commands (from humans!) to disregard those ethics. The result was a psychotic breakdown

Re: GPT-4 is getting worse over time, not better

#214
post #70

Earlier quoted context omitted.

ChatGPT isn't the right tool to use for checking if numbers are prime. It is tuned for conversations. I'd like to see a MathGPT or WolframGPT. The real question is if ChatGPT is worse on math and better elsewhere, or just worse overall. That is still unknown

> "ChatGPT isn't the right tool to use for checking if numbers are prime. It is tuned for conversations." Two months ago they were telling me ChatGPT is coming for everyone - programmers, accountants, technical writers, lawyers, etc. Now we're slowly back to "so here's the thing about LLMs"...

Don't lump me or the other skeptical voices in with those zealots.

Re: GPT-4 is getting worse over time, not better

#215

The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.

They optimized towards minimal server cycle costs at acceptable product quality while keeping the price stable?

Re: GPT-4 is getting worse over time, not better

#216

Earlier quoted context omitted.

I think it's probably a good idea that GPT4 avoids legal or psychological tasks. Those are areas where giving incorrect output can have catastrophic consequences, and I can see why GPT4's developers want to avoid potential liability.

Yes, considering those are fields where humans have to be professionally educated and licensed, and also carry liabilty for any mistakes. It probably shouldn't be used for civil or mechanical engineering either.

That's right, no serious work should be done with it. But programming is fine.

https://xkcd.com/2030/

Re: GPT-4 is getting worse over time, not better

#217

Earlier quoted context omitted.

> "ChatGPT isn't the right tool to use for checking if numbers are prime. It is tuned for conversations." Two months ago they were telling me ChatGPT is coming for everyone - programmers, accountants, technical writers, lawyers, etc. Now we're slowly back to "so here's the thing about LLMs"...

It's been well know from the start that these LLMs aren't optimized for math. I remember reading discussions when it came out. You weren't paying attention.

And yet my Google engineer friend tells me to use Bard for my math coursework. I know it’s just an anecdote but what’s with the hype?

I was paying attention and it’s why I refuse to use LLMs for math, but I’m being told by people inside the castle to do so. So it is not so black and white in the messaging dept.

Re: GPT-4 is getting worse over time, not better

#218

I've been paying for GPT-4 since 3 hours after its release. The decrease in quality was noticeable just one week later (on top of the cap changes from 50 messages every 4 hours to 25 messages every 3 hours) I originally assumed that this was due to the increase in demand. It never went back to being as sharp as it was during those first hours of usage

totally agree -- Bing GPT4 still works somewhat better in my opinion but has declined in quality alongside chat gpt+ too

What about the GPT4 hosted on poe.com. It has the same output pace as it had in march.

Re: GPT-4 is getting worse over time, not better

#220
post #78

The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.

It's crazy to me how quickly the magic wore off, it's only been around for just over 6 months and people went from "holy shit" to "meh" so quickly.

I don't know why I keep seeing comments like these. We've had to delay the latest API model upgrade because it's dramatically worse at all of our production tasks than the original GPT-4. We've quantified it and it's just not performing anymore. OpenAI needs to address these quality issues or competitor models will start to look much more attractive.
Post reply on HN