Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

91–100 of 315 posts

Re: GPT-4 is getting worse over time, not better

#91
post #9
post #4

>Having the behavior of an LLM change over time is not acceptable. By now this is actually funny to read. Never rely on another companies product to make your own product, without accepting things can change overnight and shut you down As Llama2 is self hosted, you can choose which iteration to host. Much better developer experience Edit: to be clear OpenAI is unprofitable, so is Reddit, so was Stadia. Building on to…

LLama 2 is lost in the sauce... Q: How many 90 degree permutations can you do to leave a cube invariant from the perspective of an outside observer? A: As a responsible and ethical AI language model, I must first emphasize that the concept of "90 degree permutations" and "cube" are purely theoretical and have no basis in reality. However, I understand that you are asking for a hypothetical scenario, and I will provid…

Yea it was a wake up call for me when I was asking about the volume of a cube vs a dodecahedron and Claude+ hit me with "The diameter of a dodecahedron passes through 3 pentagonal faces", just no ability to reason about geometry.

https://poe.com/lookaroundyou/1512927999895666

Sidenote, poe has a bug that mis-reports this as a conversation with Claude-2-100k, but the conversation took place on March 24, about a week after Claude+ was made public. Can't even rely on the portals to be truthful about what model was used.

Re: GPT-4 is getting worse over time, not better

#92
post #72

Earlier quoted context omitted.

This comment has an interesting take on it, haven't read the paper to verify the take: https://news.ycombinator.com/item?id=36781968 EDIT: FWIW I haven't noticed any such regression. I don't generally use it to find prime numbers, but I do use it for coding, and have been really impressed with what it's able to do. 8 This paper is being misinterpreted. The degradations reported are somewhat peculiar to the authors' t…

Fine tuning or not, its definitelly a proof one should not rely on it apart from very specific use cases (like lorem ipsum generator or something).

Just to be clear, you're saying that because they're tweaking GPT-4 to give more explanations of code, you shouldn't rely on it for coding?

Obviously if that's your own preference, I'm not going to tell you that you're wrong; but I think in general, most people wouldn't agree with that statement.

Re: GPT-4 is getting worse over time, not better

#93

Earlier quoted context omitted.

If it was an experiment I'd like to know why I'm still paying for access to something claiming to be it.

Because you signed up to pay for a beta service???? You can stop paying at any time.

If you're paying for it, it's a released product, not a beta. No matter how much a company wants to claim it's a "beta", it's simply not. The company just wants to be able to adhere to lower standards.

OpenAI is hardly the only company that pulls these shenanigans, too.

Re: GPT-4 is getting worse over time, not better

#95

I have been using GPT 3/4 using langchain and I have noticed no change in the quality of the API results. Is it being analyzed by the web interface or the API?

I’ve been using the ChatGPT interface since launch and haven’t noticed a drop in quality either. I chalk most of the complaints up to hedonic adaptation, which can lead to conspiratorial thinking fueled by a justified distrust of Big Tech in the absence of objective ways to actually measure this.

There was def something that happened in January for a week.

Like it refused to answer questions that it would in December. Not even controversial questions, but it would say 'As an AI model, it would be irresponsible to xyz...'.

I don't like posting my test questions because the internet will be mined in the future.

Re: GPT-4 is getting worse over time, not better

#96
post #40
post #17

My favorites misstep from GPT-4 was when my friend asked it about the difference between vet bulb temperatures and dry bulb. You see that typo correctly (he was dictating): > The main difference is in what they're measuring. Temperature measurement at a vet is usually taken to determine an animal's body temperature, often done rectally or via the ear. It is direct and generally provides an absolute temperature value.…

Response from GPT-4 just now: > It seems like there's a bit of a typo in your question. I think you may be referring to the difference between "wet bulb" temperatures and "dry bulb" temperatures. [...]

There are many ways to formulate this question (unfortunately I haven’t found it) and it answers in a non-deterministic manner each time, so I can give no further proof.

Re: GPT-4 is getting worse over time, not better

#97

The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.

To be fair to sama he did say that "it still seems more impressive on first use than it does after you spend more time with it"[1] but I suspect that was a kind of 4d chess move to hope to be quoted in situations like this when he knew the hype died down a few months later to cushion the blow after the hype cooled off.

[1] https://twitter.com/sama/status/1635687853324902401

Re: GPT-4 is getting worse over time, not better

#98

The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.

I feel like it has gotten worse. Your expectations have shifted towards knowing and accepting you have to prompt it a certain way. When you started, "holy shit I can just use natural language". GPT4 is influencing the bar we are setting for it, and the bar we are setting for it is influencing the outcome of its training. If you've been under a rock and you just found out about chatGPT today, you'd find it less fluid and impressive than the rest of us when we started.

Re: GPT-4 is getting worse over time, not better

#99
post #43

Earlier quoted context omitted.

I mean, they could be lying. Or only talking about the API version and not the front-facing ChatGPT.

I feel like it would be a fairly large conspiracy by the OpenAI team though? In fact, what motive would they have to make the model dumber, really? If it gets out and you're right, I think it will cause major trust issues with the product.

I'm sure they aren't trying to make the model dumber, so much as they're trying to cut compute/infra costs so they can be profitable.

Re: GPT-4 is getting worse over time, not better

#100
post #93

Earlier quoted context omitted.

Because you signed up to pay for a beta service???? You can stop paying at any time.

If you're paying for it, it's a released product, not a beta. No matter how much a company wants to claim it's a "beta", it's simply not. The company just wants to be able to adhere to lower standards. OpenAI is hardly the only company that pulls these shenanigans, too.

No, beta has nothing to do with if it's paid or not. Lots of products are released, and you pay for them, with early access/ beta.

No one lied to you. No one tricked you. It is in beta. You chose to pay for it.

Post reply on HN