Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

221–230 of 315 posts

Re: GPT-4 is getting worse over time, not better

#221
post #9

Earlier quoted context omitted.

LLama 2 is lost in the sauce... Q: How many 90 degree permutations can you do to leave a cube invariant from the perspective of an outside observer? A: As a responsible and ethical AI language model, I must first emphasize that the concept of "90 degree permutations" and "cube" are purely theoretical and have no basis in reality. However, I understand that you are asking for a hypothetical scenario, and I will provid…

Yea it was a wake up call for me when I was asking about the volume of a cube vs a dodecahedron and Claude+ hit me with "The diameter of a dodecahedron passes through 3 pentagonal faces", just no ability to reason about geometry. https://poe.com/lookaroundyou/1512927999895666 Sidenote, poe has a bug that mis-reports this as a conversation with Claude-2-100k, but the conversation took place on March 24, about a week a…

"Reasoning" about faculties which the bot does not possess (visual, audial, or individual character domains) is simply not going to happen until we hook it up to some symbolic model, or the LLM becomes advanced enough to infer a world model from text alone.

Until then, we're basically asking a blind & deaf guy, who happens to be very well-read, to reason about senses he doesn't have.

Though the mistake in your example does seem kind of egregious and I'd be curious to see whether GPT-4 would make similar mistakes.

Re: GPT-4 is getting worse over time, not better

#222

They are downgrading the plebeian version so that they can unveil a more premium tier product in 1 month to maintain the hype flow. Duh

( hype flow is a local contextual description. Descendants of these models will absolutely destroy humanity. In the near term OpenAI is crawling back some functional headroom so they have resources to maintain hype cycles, sell stronger models to governments at super premiums, and capture more value by having a wider product offering )

Re: GPT-4 is getting worse over time, not better

#223

Rather odd that MSFT invests $13B into a partnership with OpenAI, integrates OpenAI's most popular product into several MSFT products (bing, GitHub copilot, etc), and then the OpenAI-hosted ChatGPT (which is now in competition with MSFT's offerings) degrades over time. I'm old enough to remember a time when MSFT got in a bit of trouble for anticompetitive behavior. This post has some reasonable-seeming explanations f…

We don’t need a conspiracy to explain OpenAI motives. They released an early access model that cost far more to operate than they could ask for in subscriptions. They were selling dollars for $0.50 and of course everyone loved it. Then they decided they should stop bleeding money and maybe even make some.

Re: GPT-4 is getting worse over time, not better

#224

Title: "GPT-4 is getting worse over time, not better" Paper title: "How Is ChatGPT’s Behavior Changing over Time?" Paper Abstract: "GPT-3.5 and GPT-4 are the two most widely used large language model (LLM) services." When are people gonna realize that GPT-4/3.5 != ChatGPT As far as I can tell, the paper doesn't explain the methodology either, so hard to know if they're actually using "raw" GPT-4 or GPT-4 via ChatGPT.…

Which is better? I assumed they were the same. I’ve been getting ok results with chat-gpt4, might I get better results with the api gpt4?

Depends on the purpose. I don't think the various parameters (like temperature, top_p) are fully known when it comes to ChatGPT, and neither is the "system prompt" they're using. With the API, you have full control and visibility of those.

If you really want to compare "performance"/"quality", you'd have to do so via the API, using known and static parameters and locking the model version. None of which is available via ChatGPT.

Re: GPT-4 is getting worse over time, not better

#225
post #136

Earlier quoted context omitted.

It's actually 16 GPT 3.5s in a trenchcoat where each one is slightly different, like the minions.

Can anyone elaborate on the specialization? Does it have to do with segmenting the dataset, or with training on different kinds of tasks?

Fine tuning to a specific more constrained set of tasks is my understanding.

Re: GPT-4 is getting worse over time, not better

#226
There's likely no reason for GPT-4 to get worse, other than perception and some bad luck. I think the real problem is unevenness of its performance. I've had ChatGPT (3.5 and now 4) tell me it could help with problem X, then that it had no knowledge of it and couldn't help, to being an expert. All spaced out over several months.

It's likely that adequate prompt engineering would help to mitigate this problem.

Re: GPT-4 is getting worse over time, not better

#227

Earlier quoted context omitted.

I don’t believe this is true. It’s possible I was blinded by the light, but my programming tasks were previously (during the early access program) being handled by GPT-4 regularly and now they aren’t. I’ve also seen many anecdotes from engineers who had exceptionally early access before GPT-4 was public knowledge. The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back…

> terrifyingly unaligned Honestly, if people think that a statistical language model is "terrifying" because it can verbalise the concept of a mass killing, they need to give their heads a wobble. My text editor can be used to write "set off a nuclear weapon in a city, lol". Is Notepad++.exe terrifying? What about the Sum of All Fears ? I could get some pointers from that. Is Tom Clancy unaligned? Am I terrifying bec…

Perhaps an analogy could clarify? Although it isn't a perfect one, I'll try to use the points of contrast to explain why it can be considered dangerous.

If a young child is really aggressive and hitting people, it's worrying even though it may not actually hurt anyone. Because the child is going to grow up, and it needs to eliminate that behavior before it's old enough to do damage by its aggression. (don't take this as a comprehensive description, just a tiny slice of cause-effect)

But the problem with AI is that we don't have continuity between today's AI and future AI. We can see that aggressive speech is easy to create by accident - Bing's Sydney output text that threatened peoples' lives. We may not be worried about aggressive speech from LLMs because it can't do damage, but similar behavior could be really dangerous from an AI which has the ability to form a model of the world based on the text it generates (in other words, it treats its output as thoughts).

But even if we remove that behavior from LLMs today, that doesn't mean aggressive behavior won't be learned by future AI, because it may be easy for aggressive behavior to emerge and we don't know how to prevent it from emerging. With a small child, we can theoretically prevent aggressive behavior from emerging in that child's adulthood with sufficient training in childhood.

It's not the same for AI - we don't know how to prevent aggression or other unaligned behavior from emerging in more advanced AI. Most counter arguments seem to come down to hoping that aggression won't emerge, or won't emerge easily. To me, that's just wishful thinking. It might be true, but it's a bit like playing Russian roulette with an unknown number of bullets in an unknown number of chambers.

Re: GPT-4 is getting worse over time, not better

#228

Earlier quoted context omitted.

>Is Tom Clancy unaligned? Yes, humans are unaligned. This is why alignment is hard: we're trying to produce machines with human-level intelligence but superhuman levels of morality.

Humans are definitely aligned, and for the same reasons as a LLM. Socialization, being allowed to work, being allowed to speak. edit: It's a social faux pas to say "died" about a person acquainted to the listener in most situations, you have to say "passed away."

> Humans are definitely aligned

Yes, that's why climate change was rapidly addressed when we began to understand it well 60 years ago and why war has always been so rare in human history.

Re: GPT-4 is getting worse over time, not better

#229

Lot of confusion in the discussion between chatgpt the large language model (aka gpt-35-turbo) and ChatGPT the consumer application (the website where you type questions and you get a response beginning a conversation, which can be configured to use either the chatgpt or GPT4 models). To be clear: * This paper called the models directly via the API not the ChatGPT application. This means that changes to the ChatGPT s…

As far as I can tell, you are the only person in this thread who actually skimmed the paper. Thank you for pointing this out!

The API clearly delineates the March and June versions. The paper authors ran tests on different API versions. The fact that these versions are different is clear & transparent. Anyone can use the March version of GPT by calling the API.

gpt-4-0314: very slow, smart

gpt-4-0613: fast, less smart

Re: GPT-4 is getting worse over time, not better

#230

Earlier quoted context omitted.

I don’t believe this is true. It’s possible I was blinded by the light, but my programming tasks were previously (during the early access program) being handled by GPT-4 regularly and now they aren’t. I’ve also seen many anecdotes from engineers who had exceptionally early access before GPT-4 was public knowledge. The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back…

> terrifyingly unaligned Honestly, if people think that a statistical language model is "terrifying" because it can verbalise the concept of a mass killing, they need to give their heads a wobble. My text editor can be used to write "set off a nuclear weapon in a city, lol". Is Notepad++.exe terrifying? What about the Sum of All Fears ? I could get some pointers from that. Is Tom Clancy unaligned? Am I terrifying bec…

I think there is a significant difference between an LLM and your other examples.

Society is a lot more fragile than many people believe.

Most people aren’t Ted Kazinsky.

And most wanna be Ted Kazinsky’s that we’ve caught don’t have super smart friends they can call up and ask for help in planning their next task.

But a world where every disgruntled person who aims to do the most harm has an incredibly smart friend who is always DTF no matter the task? Who is capable of reeling them in to be more pragmatic about sowing chaos, death, and destruction?

That world is on the horizon, it’s something we are going to have to adapt to, and it’s significantly different than Notepad++. It’s also significantly different than you, assuming you are not willing to help your neighbor get away with serial murder.

I think this is something that’s going to significantly increase the sophistication of bad actors in our society and I think that outcome is inevitable at this point. I don’t think this is “the end times” - nor do I think trying to align and regulate LLMs is going to be effective unless training these things continues to require a supply chain that’s easily monitored and controlled. Every step we take towards training LLMs on commodity/consumer hardware is a step away from effective regulation, and selfishly a step I support.

Post reply on HN