Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

251–260 of 315 posts

Re: GPT-4 is getting worse over time, not better

#251

The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.

Anyone who has saved any early prompts/responses can plainly attest to this by sending those prompts right back into the system and looking at the output.

As to why, it's probably a combination of trying to optimize GPT4 for scale and trying to "improve" it by pre/post processing everything for safety and/or factual accuracy.

Regardless of the debate as to whether it is "worse" or "better," the result is the same. The model is guaranteed to have inconsistent performance and capability over time. Because of that alone, I contend it's impossible to develop anything reliable on top of such a mess.

Re: GPT-4 is getting worse over time, not better

#252

Earlier quoted context omitted.

I don’t believe this is true. It’s possible I was blinded by the light, but my programming tasks were previously (during the early access program) being handled by GPT-4 regularly and now they aren’t. I’ve also seen many anecdotes from engineers who had exceptionally early access before GPT-4 was public knowledge. The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back…

I used it for writing assistance and Plot development. Specifically, a novel re: the conquest of Mexico in the 16th cent. It was great at spitting out ideas re: action scenes and even character development. In the past month or so, it has become so cluttered with caveats and tripe regarding the political aspects of the conquest, that it is useless. I can’t replicate the work I was doing before. Actually cancelled my…

writing "assistance", lol

Re: GPT-4 is getting worse over time, not better

#254
post #162

This is very interesting, because we(at cheatlayer.com) can publish results of the exact opposite happening and we test a lot of code generation with thousands of actual customers live. It's entirely possible the examples are cherry-picked or could be explained by fine tuning differences, but in terms of "proofs" in the mathematical sense the paper doesn't prove this since you can get the opposite results based on th…

Maybe you should publish then, because just claiming the opposite isn’t proof of anything.

Re: GPT-4 is getting worse over time, not better

#255
I personally am finding it difficult to find a tangible difference int he quality of output produced by GPT-4 vs GPT-3.5 in what I’ve been using it for recently. Might just be me and perhaps my prompts are not very good quality, but nonetheless, I feel like the difference isn’t nearly as significant as has been stated, at least, not anymore.

Re: GPT-4 is getting worse over time, not better

#256

Earlier quoted context omitted.

I don’t believe this is true. It’s possible I was blinded by the light, but my programming tasks were previously (during the early access program) being handled by GPT-4 regularly and now they aren’t. I’ve also seen many anecdotes from engineers who had exceptionally early access before GPT-4 was public knowledge. The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back…

> terrifyingly unaligned Honestly, if people think that a statistical language model is "terrifying" because it can verbalise the concept of a mass killing, they need to give their heads a wobble. My text editor can be used to write "set off a nuclear weapon in a city, lol". Is Notepad++.exe terrifying? What about the Sum of All Fears ? I could get some pointers from that. Is Tom Clancy unaligned? Am I terrifying bec…

As a human, you have the context to understand the difference between the “Sum of all Fears” vs planning an attack or writing for a purpose beyond creative writing.

The model does not. If you ask ChatGPT about strategies for successful mass killing, that’s probably not good for society nor for the company.

In a military context, they may want a system where an LLM would provide guidance for how to most effectively kill people. Presumably such a system would have access controls to reduce risks and avoid providing aid to an enemy.

Re: GPT-4 is getting worse over time, not better

#257

Earlier quoted context omitted.

>Is Tom Clancy unaligned? Yes, humans are unaligned. This is why alignment is hard: we're trying to produce machines with human-level intelligence but superhuman levels of morality.

There’s no such thing as “superhuman” morality, morality is just social mores and norms accepted by the people in a society at some given time. It does not advance or decline, but it changes. What you’re talking about is a very small subset of the population forcing their beliefs on everyone else by encoding them in AI. Maybe that’s what we should do but we should be honest about it.

If you were to create a moral code for a bees hive, with the goal of evolving the bees towards the good (in your eyes), that would be a super-bee level morality.

For us, such moral codes assume the form of religions: those begin as a set of moral directives, that eventually accumulate cruft (complex ceremonies, superstitions, pseudo thought-leaders and mountains of literature), devolve into lowly cults and get replaced with another religion. However, when such moral codes are created, they all share the same core principles, in all ages and cultures. That's the equivalent of a super-bee moral code.

Re: GPT-4 is getting worse over time, not better

#258

Earlier quoted context omitted.

> terrifyingly unaligned Honestly, if people think that a statistical language model is "terrifying" because it can verbalise the concept of a mass killing, they need to give their heads a wobble. My text editor can be used to write "set off a nuclear weapon in a city, lol". Is Notepad++.exe terrifying? What about the Sum of All Fears ? I could get some pointers from that. Is Tom Clancy unaligned? Am I terrifying bec…

I think there is a significant difference between an LLM and your other examples. Society is a lot more fragile than many people believe. Most people aren’t Ted Kazinsky. And most wanna be Ted Kazinsky’s that we’ve caught don’t have super smart friends they can call up and ask for help in planning their next task. But a world where every disgruntled person who aims to do the most harm has an incredibly smart friend w…

What will happen is that this tech will advance, and regular joes will have access to the "TERRIFYING" unaligned models - and nothing will happen.

This stuff isn't magic. Wannabe Ted Kaczynski will ask BasedGPT how to build bombs , it will tell them, and nothing will happen because building bombs and detonating them and not getting caught is REALLY HARD.

The limiting factor for those seeking wanton destruction is not a lack of know-how, but a lack of talent/will. Do we get ~1-4 new mass shootings a year? Seems reasonable but doesn't matter in the grand scheme of things. (That's like, what, a day of driving fatalities?)

Unaligned publicly available powerful AI ("Open" AI, one might say) is a net good. The sooner we get an AI that will tell us how to cook meth and make nuclear bombs, the better.

Re: GPT-4 is getting worse over time, not better

#259

Earlier quoted context omitted.

> Humans are definitely aligned Yes, that's why climate change was rapidly addressed when we began to understand it well 60 years ago and why war has always been so rare in human history.

It seems "aligned" is in the eye of the beholder.

Cue Mrs. Slokam's "...and I am unanimous in that!"

Aligned LLMs are just like altered brains... they don't function properly.

Re: GPT-4 is getting worse over time, not better

#260

Earlier quoted context omitted.

>Is Tom Clancy unaligned? Yes, humans are unaligned. This is why alignment is hard: we're trying to produce machines with human-level intelligence but superhuman levels of morality.

The end result being Hal-9000

Halal-9000 would be more apropos of the goals of alignment and morality.
Post reply on HN