Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

41–50 of 315 posts

Re: GPT-4 is getting worse over time, not better

#41

I also feel like I'm getting dumber over time, not smarter, so perhaps GPT-4 is humanlike in that regard.

How’s life in Japan?

Life is good, Japan is a wonderful place to live. Was just in Kyoto last weekend for the Gion Matsuri, a big parade festival. If you're interested to come, check out my company TableCheck: https://careers.tablecheck.com/ We have a really talented and motivated team, great clients, and have a lot of fun.

Re: GPT-4 is getting worse over time, not better

#42
post #40
post #17

My favorites misstep from GPT-4 was when my friend asked it about the difference between vet bulb temperatures and dry bulb. You see that typo correctly (he was dictating): > The main difference is in what they're measuring. Temperature measurement at a vet is usually taken to determine an animal's body temperature, often done rectally or via the ear. It is direct and generally provides an absolute temperature value.…

Response from GPT-4 just now: > It seems like there's a bit of a typo in your question. I think you may be referring to the difference between "wet bulb" temperatures and "dry bulb" temperatures. [...]

Also tried with GPT-3.5 which also catches the issue.

Re: GPT-4 is getting worse over time, not better

#43
post #34

Earlier quoted context omitted.

Are you sure it wasn't just that the novelty wore off after a few hours of usage? I never really got into LLMs, but I must say at first it seemed like pretty cool stuff. OpenAI have repeatedly stated the model hasn't changed so how could this happen otherwise?

I mean, they could be lying. Or only talking about the API version and not the front-facing ChatGPT.

I feel like it would be a fairly large conspiracy by the OpenAI team though? In fact, what motive would they have to make the model dumber, really?

If it gets out and you're right, I think it will cause major trust issues with the product.

Re: GPT-4 is getting worse over time, not better

#46
post #35

I have not read the paper yet (in my backlog, here's the paper: https://arxiv.org/pdf/2307.09009.pdf ), but it's important note that the paper is entitled "How Is ChatGPT’s Behavior Changing over Time?" not that it's necessarily "getting worse." Here's a more nuanced (not an AI clout chasing account) discussion by Arvind Narayanan (Princeton CS prof) about the results: https://twitter.com/random_walker/status/1681489…

How does Code Interpreter work vs base GPT-4 for code snippets? I'm writing questions and pasting context code into GPT-4 right now, and it works pretty well.

Re: GPT-4 is getting worse over time, not better

#47
If you have the API, you can track if the results change by setting temperature to 0, over time the difference to the same questions should not change drastically. I think this is the gold test, and not from using ChatGPT by feeling it out. The best place to track this would be to use GitHub

Re: GPT-4 is getting worse over time, not better

#48
Title: "GPT-4 is getting worse over time, not better"

Paper title: "How Is ChatGPT’s Behavior Changing over Time?"

Paper Abstract: "GPT-3.5 and GPT-4 are the two most widely used large language model (LLM) services."

When are people gonna realize that GPT-4/3.5 != ChatGPT

As far as I can tell, the paper doesn't explain the methodology either, so hard to know if they're actually using "raw" GPT-4 or GPT-4 via ChatGPT...

I hoped that eventually people would realize they are vastly different, and your experience/results with be vastly different depending on which you use too. But that hope is slowly fading away, and OpenAI isn't exactly seeming to want to help resolve the confusion either.

Re: GPT-4 is getting worse over time, not better

#49
post #4

>Having the behavior of an LLM change over time is not acceptable. By now this is actually funny to read. Never rely on another companies product to make your own product, without accepting things can change overnight and shut you down As Llama2 is self hosted, you can choose which iteration to host. Much better developer experience Edit: to be clear OpenAI is unprofitable, so is Reddit, so was Stadia. Building on to…

That's a pretty defeatist take. Surely if you pay for a service you should expect the provider to be making good faith efforts to provide the same quality of service over time? Natural degradation would be fine, but purposefully sandbagging the service so it gets worse because cheaper is unacceptable. That we have become numb to the point that we collectively accept such poor behavior on the part of vendors in concre…

The millions it takes to train OpenAI models is given by people who would like to get it back + some profits. Once the VCs start forcing the roadmap onto the employees, thats when customers get screwed over.

So to be more precise, building a product on top of an unprofitable product means things will usually change for the worse (cost cutting) sooner or later

Re: GPT-4 is getting worse over time, not better

#50
post #24

Earlier quoted context omitted.

The mixture of experts approach was found to be inferior than a straightforward transformer. However, researchers discovered that MoE models give substantially better results when guided by fine-tuning. My guess is that GPT-4 was an experiment to prove out that theory at huge scale and I also suspect that its performance astonished OpenAI as much as everyone else.

If it was an experiment I'd like to know why I'm still paying for access to something claiming to be it.

Because you are getting value from it. Why else would you pay?
Post reply on HN