Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

111–120 of 315 posts

Re: GPT-4 is getting worse over time, not better

#111
post #21

The point is valid about performance over time changing, but using an LLM to see if a number is prime is a terrible use of the product. It does, of course, give you a correct answer if you have the wolfram alpha plugin installed, though.

Solving an IQ test is also not useful by itself, but it is a good benchmark of intelligence.

> it is a good benchmark of intelligence.

This is not well-established, and is subject to a great deal of dispute amongst experts in the field.

Re: GPT-4 is getting worse over time, not better

#113
I don't understand. I literally just copy-pasted the "Is 17077 a prime number? Think step by step." question in my GPT-4 and it wrote me a full-page response with step-by-step explanation.

The author is claiming that "the latest version of GPT-4 did not generate intermediate steps and instead answered incorrectly with a simple "No." but that is not the case

Re: GPT-4 is getting worse over time, not better

#114
post #93

Earlier quoted context omitted.

Because you signed up to pay for a beta service???? You can stop paying at any time.

If you're paying for it, it's a released product, not a beta. No matter how much a company wants to claim it's a "beta", it's simply not. The company just wants to be able to adhere to lower standards. OpenAI is hardly the only company that pulls these shenanigans, too.

this is just not at all true -- companies have private betas that they monetize all the time. yes, the company provides lower guarantees on the product being supported. but the customer gets to use a product far before they normally would if they waited for general availability. that's the core of the idea behind a beta, not related to if it's paid or not

Re: GPT-4 is getting worse over time, not better

#115
post #34

Earlier quoted context omitted.

Are you sure it wasn't just that the novelty wore off after a few hours of usage? I never really got into LLMs, but I must say at first it seemed like pretty cool stuff. OpenAI have repeatedly stated the model hasn't changed so how could this happen otherwise?

I mean, they could be lying. Or only talking about the API version and not the front-facing ChatGPT.

Not saying it's so in this case (or even in the general case), but "marketing" is sometimes a synonym for lying.

Re: GPT-4 is getting worse over time, not better

#116
post #79
post #59

Earlier quoted context omitted.

[flagged]

Then why comment? I'm not trying to be an ass, but this is literally the only comment you've made in this thread. The linked article is a Twitter thread. If you don't have Twitter, fine, move on to the next thing you can actually comment on. Person A offers a reason why we may want to be slightly more skeptical than normal about this information. Person B suggests we can pretty easily look past that. Person C (you) i…

I don't think the comment is unwarranted - it happened to me the other day: someone refers me to a particular Twitter post and I can't access it without a Twitter account. I don't know if it's a glitch or deliberate but it makes referring people to Twitter similar to referring them to Facebook posts (which is not really practiced on HN). Times change, I guess.

Re: GPT-4 is getting worse over time, not better

#117
post #43

Earlier quoted context omitted.

I mean, they could be lying. Or only talking about the API version and not the front-facing ChatGPT.

I feel like it would be a fairly large conspiracy by the OpenAI team though? In fact, what motive would they have to make the model dumber, really? If it gets out and you're right, I think it will cause major trust issues with the product.

They want to free up GPU's for other projects. And to do that, they need to make the model dumber.

Re: GPT-4 is getting worse over time, not better

#119
post #92

Earlier quoted context omitted.

Fine tuning or not, its definitelly a proof one should not rely on it apart from very specific use cases (like lorem ipsum generator or something).

Just to be clear, you're saying that because they're tweaking GPT-4 to give more explanations of code, you shouldn't rely on it for coding? Obviously if that's your own preference, I'm not going to tell you that you're wrong; but I think in general , most people wouldn't agree with that statement.

I'm still wondering, why should anyone rely on AI generated answers? They are logically no better than search engine results. By that I mean, you can't tell if it's returning absolute trash or spot on correct. Building trust into it all is going to be either a) expensive or b) driven by all the wrong incentives.

Re: GPT-4 is getting worse over time, not better

#120
We are just beginning to understand that we entered the age of software taming: tune the search engine a little, more SEO sites will show first; tune the spam filter a little, more real users get banned. Same happens to physics in videogames and now LLM's. It's all about trying to control complexity with a few knobs.
Post reply on HN