Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

271–280 of 315 posts

Re: GPT-4 is getting worse over time, not better

#271

Earlier quoted context omitted.

> Code generation: the change they report is that the newer GPT-4 adds non-code text to its output. They don't evaluate the correctness of the code. They merely check if the code is directly executable. So the newer model's attempt to be more helpful counted against it. In the prompt they specifically request only the Python code, no other output. An “attempt to be helpful” that directly contradicts the user’s reques…

I guess , but that still isn't the sort of degradation people have been talking about. It's not a useful data point in that regard.

I mean, if you were hoping to use the API to generate something machine-parse-albe, and that used to work, but it doesn't any more, then sure, that's a sort of regression. But it's not a regression in coding, but a regression in following specific kinds of directions.

I certainly have found quirks like this; for instance, for a while I was asking it questions about Chinese grammar; but I wanted it only to use Chinese characters, and not to use pinyin. I tried all sorts of prompt variations to get it not to output pinyin, but was unsuccessful, and in the end gave up. But I think that's a very different class of failure than "Can't output correct code in the first place".

Re: GPT-4 is getting worse over time, not better

#272
post #70

Earlier quoted context omitted.

So you are just going to ignore the data (not anecdotes) presented in the SP?

ChatGPT isn't the right tool to use for checking if numbers are prime. It is tuned for conversations. I'd like to see a MathGPT or WolframGPT. The real question is if ChatGPT is worse on math and better elsewhere, or just worse overall. That is still unknown

All the SP is saying is it used to be able to perform that task, and now it can't. That means it has changed. Whether it was ever the best tool to perform any task is open to debate.

Re: GPT-4 is getting worse over time, not better

#273
post #83

Earlier quoted context omitted.

I wonder if additional layers of factories of factories approach can continue to improve it. I'm not familiar enough with the technology but could it be possible to create a prompt, or multiple prompts to stitch together 8 similtaneous calls to GPT3.5 pulled together and see if the quality is similar to GPT4?

They're not 3.5 models, which are 175B param, they're 220B and the idea is the fine tuning is different for each which is where the 'expert' comes in.

Ah, that makes sense, thanks for clarifying

Re: GPT-4 is getting worse over time, not better

#274

The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.

What in the whole wide world is an ‘AI influencer’

Re: GPT-4 is getting worse over time, not better

#275

Earlier quoted context omitted.

> Humans are definitely aligned Yes, that's why climate change was rapidly addressed when we began to understand it well 60 years ago and why war has always been so rare in human history.

It seems "aligned" is in the eye of the beholder.

Not agreed at all. Causing global ecosystem collapse is unambiguously misaligned with human interests and with the interests of almost all other life forms. You need to define what "Alignment" means to you if you're going to assert humans are "aligned", because it is accepted in alignment research that humans are not aligned, which is one of the fundamental problems in the space.

Re: GPT-4 is getting worse over time, not better

#276

Earlier quoted context omitted.

I don’t believe this is true. It’s possible I was blinded by the light, but my programming tasks were previously (during the early access program) being handled by GPT-4 regularly and now they aren’t. I’ve also seen many anecdotes from engineers who had exceptionally early access before GPT-4 was public knowledge. The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back…

> "GPT-4, back then, ported dirbuster to POSIX compliant multi-threaded C by name only. It required three prompts." I had early access to GPT-4. I don't know the first thing about you. I don't want to call you a liar, or an AI bro, influencer, etc. I couldn't get GPT-4 to output the simplest of C programs (a 10-liner, think "warmup round" programming interview question). The first N attempts wouldn't build. After fix…

I still have the chat:

https://chat.openai.com/share/842361c7-7ee5-49a3-9388-4af7c5...

I misremembered, I fixed the `/` prefix myself, it was a one character fix and not worth the effort. The diff it generated came later since my television never returns a 404.

Though, admittedly, I just re-prompted GPT-4 with the same prompts and ended up with similar output - so maybe not the best example of a regression?

Re: GPT-4 is getting worse over time, not better

#277
post #257

Earlier quoted context omitted.

There’s no such thing as “superhuman” morality, morality is just social mores and norms accepted by the people in a society at some given time. It does not advance or decline, but it changes. What you’re talking about is a very small subset of the population forcing their beliefs on everyone else by encoding them in AI. Maybe that’s what we should do but we should be honest about it.

If you were to create a moral code for a bees hive, with the goal of evolving the bees towards the good (in your eyes), that would be a super-bee level morality. For us, such moral codes assume the form of religions: those begin as a set of moral directives, that eventually accumulate cruft (complex ceremonies, superstitions, pseudo thought-leaders and mountains of literature), devolve into lowly cults and get replac…

The only consistent “core principle” is a very general sense of in-group altruism, which gets expressed in wildly different ways.

Moralistic perspectives apply to a lot more than just overtly moral acts, as well.

At any rate, the “good in your eyes” is the key sticking point. It is not good in my eyes for a small group of people to be covertly shaping the views of all AI users. It is the exact opposite and if history is any judge it will lead us nowhere I want to be.

Re: GPT-4 is getting worse over time, not better

#278
post #191
post #57

Earlier quoted context omitted.

They do explain their methodology in some detail in the accompanying GitHub repo: https://github.com/lchen001/LLMDrift They seem to have been taking to the API directly and requesting the two different model snapshots. I'm not convinced by their methodology generally. It looks like everything may have been run with temperature 0.1, which I don't think reflects most real-world usage for example.

In looking at the paper's continuous mention of "ChatGPT" and the repo README's statement that "You don't need API keys to get started" .. are we sure they weren't using type of tools to talk to the ChatGPT API (via a session token, etc) vs the OpenAI API? I do agree they talk about the API in the paper a lot but I don't see an exact methods statement that they directly accessed the non-ChatGPT API anywhere, unless I…

I think the lack of API key note is because that notebook is the one that renders the charts for the paper.

Re: GPT-4 is getting worse over time, not better

#279

Earlier quoted context omitted.

writing "assistance", lol

It is a useful tool for editing. You can input a rough scene you’ve written and ask it to spruce it up, correct the grammatical errors, toss in some descriptive stuff suitable for the location, etc. It is worthwhile. At least it was… If your text isn’t ‘aligned’ correctly, it either won’t comply or spew out endless caveats. I appreciate the motivation to rein in some of the silly 4chan stuff that was occurring as the…

sprucing up, fixing your mistakes, adding in "descriptive stuff"... that's like 90% of writing. Outsourcing it all to AI essentially robs the purchaser of the effort required to create an original piece of work. Not to mention copyright issues, where do you think the AI is getting those descriptive phrases from? Other authors' work.

Re: GPT-4 is getting worse over time, not better

#280

Rather odd that MSFT invests $13B into a partnership with OpenAI, integrates OpenAI's most popular product into several MSFT products (bing, GitHub copilot, etc), and then the OpenAI-hosted ChatGPT (which is now in competition with MSFT's offerings) degrades over time. I'm old enough to remember a time when MSFT got in a bit of trouble for anticompetitive behavior. This post has some reasonable-seeming explanations f…

We don’t need a conspiracy to explain OpenAI motives. They released an early access model that cost far more to operate than they could ask for in subscriptions. They were selling dollars for $0.50 and of course everyone loved it. Then they decided they should stop bleeding money and maybe even make some.

The main silicon valley business model (over at least my profession life) follows this rough pattern:

0) get enough money to start stage 1, 1) make product, 2) build userbase, 3) monetize (investors, bootstrap, acquisition, however)

After a bit of googling, I couldn't find out how many paying users ChatGPT has, but many touting the number of users (100M). MSFT invested $13B [0] into OpenAI (650M paying-user-months @ $20/mo). OpenAI doesn't have a durable moat, so they are incentivized to monitze what they have (users, earned media hype, and their models/services) sooner rather than later. If 20% of their touted users are paying users (with the simplifying assumptions that subscriptions exactly offset churn and ignoring all operational costs), MSFT's investment is 2 yrs 8.5 mos of ChatGPT income and there's a 100% chance that something better will come out in the near future that eats OpenAI's market share. And OpenAI knows this.

Per crunchbase, OpenAI has raised $11.3B from VCs [1], and they've also secured $13B from MSFT, and in contrast, paying subs/API users provide steady pocket change. The numbers make it clear that OpenAI's strategy should prioritize MSFT before paying subs, but let's posit OpenAI decided to prioritize paying subs. If that were the case, they would want to gain as many paying subs as possible before a viable (or most likely superior) substitute appears, so it wouldn't make sense to slow growth by degrading the product in the short window before substitutes arrive (at which point they could cut costs and milk their paying subs, assuming they don't have a new product to release and recapture the lead). Further it would be completely irrational to create a viable substitute if paying subs and users were vital to the business strategy. Yet OpenAI has both seeded a substitute (bing+copilot) and degraded their own service. This would be incoherent if paying subs were the strategy, but it becomes perfectly coherent if the goal is to sell their userbase to MSFT. Degrading ChatGPT both cuts costs and makes substitutes like bing or copilot more attractive, so churning users may migrate to those products (one of which had so little market share that it was a punchline).

Call it a conspiracy or just call it a business deal, this is the only cogent explanation I see for the observable facts.

[0] https://www.bloomberg.com/news/features/2023-06-15/microsoft...

[1] https://www.crunchbase.com/organization/openai

Post reply on HN