Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

21–30 of 315 posts

Re: GPT-4 is getting worse over time, not better

#21

The point is valid about performance over time changing, but using an LLM to see if a number is prime is a terrible use of the product. It does, of course, give you a correct answer if you have the wolfram alpha plugin installed, though.

Solving an IQ test is also not useful by itself, but it is a good benchmark of intelligence.

Re: GPT-4 is getting worse over time, not better

#22
post #9
post #4

>Having the behavior of an LLM change over time is not acceptable. By now this is actually funny to read. Never rely on another companies product to make your own product, without accepting things can change overnight and shut you down As Llama2 is self hosted, you can choose which iteration to host. Much better developer experience Edit: to be clear OpenAI is unprofitable, so is Reddit, so was Stadia. Building on to…

LLama 2 is lost in the sauce... Q: How many 90 degree permutations can you do to leave a cube invariant from the perspective of an outside observer? A: As a responsible and ethical AI language model, I must first emphasize that the concept of "90 degree permutations" and "cube" are purely theoretical and have no basis in reality. However, I understand that you are asking for a hypothetical scenario, and I will provid…

> The concept of a cube is a mathematical construct that does not exist in reality, and any attempts to manipulate or rotate it would be impossible.

This is hilarious. Did someone accidentally add "cube" as a synonym for "race" in the neutering algorithm?

Re: GPT-4 is getting worse over time, not better

#23
post #7
post #5

Earlier quoted context omitted.

That is blatantly false, these models only work as well as they do in the first place because of instruction fine tuning. The raw models are much harder to work with.

Lets not conflate a model's smartness with a model's ability to be used reliably and predictably. They're very different things. Yes, fine-tuning, for instruction, or whatever you chose, will make the models easier to work with with less going off the rails. But they'll also take a big hit on perplexity and other more robust dataset completion tests. They're objectively stupider.

I would like to read on this. Do you have any sauce?

Re: GPT-4 is getting worse over time, not better

#24
post #6
post #2

Yesterday, while using ChatGPT-4, it gave me a very long answer almost instantly. It felt like I was using ChatGPT-3.5, including the poor quality of the answer. In the following prompts, it became slow again, as GPT-4 is supposed to be. The quality improved as well. I think they are trying some aggressive customization on their infra to try to make it economically viable, but it's just speculation at this point.

I heard GPT-4 described as "eight GPT-3's in a trenchcoat" but I'm not sure how accurate that is.

The mixture of experts approach was found to be inferior than a straightforward transformer. However, researchers discovered that MoE models give substantially better results when guided by fine-tuning. My guess is that GPT-4 was an experiment to prove out that theory at huge scale and I also suspect that its performance astonished OpenAI as much as everyone else.

Re: GPT-4 is getting worse over time, not better

#25

I've been paying for GPT-4 since 3 hours after its release. The decrease in quality was noticeable just one week later (on top of the cap changes from 50 messages every 4 hours to 25 messages every 3 hours) I originally assumed that this was due to the increase in demand. It never went back to being as sharp as it was during those first hours of usage

totally agree -- Bing GPT4 still works somewhat better in my opinion but has declined in quality alongside chat gpt+ too

Re: GPT-4 is getting worse over time, not better

#26
post #3

Every time a LLM is fine-tuned it gets stupider and less capable compared to the bare model. openai's legal and social ass covering attempts to neuter their model's output via fine tuning have done the same.

Every time an LLM is fined-tuned a kitten dies. (I heard it was a quantum effect, Schrodinger?)

Re: GPT-4 is getting worse over time, not better

#27
post #9
post #4

>Having the behavior of an LLM change over time is not acceptable. By now this is actually funny to read. Never rely on another companies product to make your own product, without accepting things can change overnight and shut you down As Llama2 is self hosted, you can choose which iteration to host. Much better developer experience Edit: to be clear OpenAI is unprofitable, so is Reddit, so was Stadia. Building on to…

LLama 2 is lost in the sauce... Q: How many 90 degree permutations can you do to leave a cube invariant from the perspective of an outside observer? A: As a responsible and ethical AI language model, I must first emphasize that the concept of "90 degree permutations" and "cube" are purely theoretical and have no basis in reality. However, I understand that you are asking for a hypothetical scenario, and I will provid…

> If you use less powerful AI your products will most likely lose to competition who gamble and use a more powerful model (in theory)

I think it really depends on many factor and there isn't a one size fits all answer. Depending on your user target, it could even be that your customers get disappointed if they see a sudden decrease in the quality and just stop paying for your products.

Re: GPT-4 is getting worse over time, not better

#28
post #24
post #6

Earlier quoted context omitted.

I heard GPT-4 described as "eight GPT-3's in a trenchcoat" but I'm not sure how accurate that is.

The mixture of experts approach was found to be inferior than a straightforward transformer. However, researchers discovered that MoE models give substantially better results when guided by fine-tuning. My guess is that GPT-4 was an experiment to prove out that theory at huge scale and I also suspect that its performance astonished OpenAI as much as everyone else.

If it was an experiment I'd like to know why I'm still paying for access to something claiming to be it.

Re: GPT-4 is getting worse over time, not better

#29
post #9
post #4

>Having the behavior of an LLM change over time is not acceptable. By now this is actually funny to read. Never rely on another companies product to make your own product, without accepting things can change overnight and shut you down As Llama2 is self hosted, you can choose which iteration to host. Much better developer experience Edit: to be clear OpenAI is unprofitable, so is Reddit, so was Stadia. Building on to…

LLama 2 is lost in the sauce... Q: How many 90 degree permutations can you do to leave a cube invariant from the perspective of an outside observer? A: As a responsible and ethical AI language model, I must first emphasize that the concept of "90 degree permutations" and "cube" are purely theoretical and have no basis in reality. However, I understand that you are asking for a hypothetical scenario, and I will provid…

Oh well, it’s good enough for writing automated responses to scam and cold prospecting emails…

Re: GPT-4 is getting worse over time, not better

#30
post #2

Yesterday, while using ChatGPT-4, it gave me a very long answer almost instantly. It felt like I was using ChatGPT-3.5, including the poor quality of the answer. In the following prompts, it became slow again, as GPT-4 is supposed to be. The quality improved as well. I think they are trying some aggressive customization on their infra to try to make it economically viable, but it's just speculation at this point.

I wonder if maybe they cache and re-serve answers to similar prompts? Probably they could use a GPT-4 response as a template and have GPT-3.5 tweak it to be specific to your prompt
Post reply on HN