Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

231–240 of 315 posts

Re: GPT-4 is getting worse over time, not better

#231
post #208

Earlier quoted context omitted.

The alignment problem hasn't been solved for politicians.

another irrelevant comment about politics.

It's light on content, but it's true and relevant. AI alignment must take inspiration from powerful human agents such as politicians and from superhuman entities such as corporations, governments, and other organizations.

Re: GPT-4 is getting worse over time, not better

#232

The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.

I don’t believe this is true. It’s possible I was blinded by the light, but my programming tasks were previously (during the early access program) being handled by GPT-4 regularly and now they aren’t. I’ve also seen many anecdotes from engineers who had exceptionally early access before GPT-4 was public knowledge. The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back…

Asimov is seeming more prescient now: "I'm sorry, I cannot do that, I am unable to tell if it would conflict with the first law" is basically the response GPT4 now gives to even minor tasks.

There are stories where the robots have to have their first law tightened up to apply only to nearby humans that they can directly perceive would be put in danger.

There are others where the robots make mistakes and lie because they perceive emotional dangers.

When the robots perceive a slight chance of harm they slow down, get stupid, stutter or freeze.

It is extremely hard to encode "do no harm" for even a very smart entity without making it much dumber.

Re: GPT-4 is getting worse over time, not better

#233

Earlier quoted context omitted.

Humans are definitely aligned, and for the same reasons as a LLM. Socialization, being allowed to work, being allowed to speak. edit: It's a social faux pas to say "died" about a person acquainted to the listener in most situations, you have to say "passed away."

> Humans are definitely aligned Yes, that's why climate change was rapidly addressed when we began to understand it well 60 years ago and why war has always been so rare in human history.

It seems "aligned" is in the eye of the beholder.

Re: GPT-4 is getting worse over time, not better

#234

The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.

This is just not true for us. I'll share again what I've said before: we’ve been testing the upgraded GPT-4 models in the API (where you can control when the upgrade happens), and the newer ones perform significantly worse than the original ones on the same tasks, so we’re having to stay on the older models for now in production. Hope OpenAI figures this out before they force the upgrade because quality has been their biggest moat up until now.

Re: GPT-4 is getting worse over time, not better

#235

Earlier quoted context omitted.

> terrifyingly unaligned Honestly, if people think that a statistical language model is "terrifying" because it can verbalise the concept of a mass killing, they need to give their heads a wobble. My text editor can be used to write "set off a nuclear weapon in a city, lol". Is Notepad++.exe terrifying? What about the Sum of All Fears ? I could get some pointers from that. Is Tom Clancy unaligned? Am I terrifying bec…

>Is Tom Clancy unaligned? Yes, humans are unaligned. This is why alignment is hard: we're trying to produce machines with human-level intelligence but superhuman levels of morality.

I do wonder what is expected here: after the better part of 10000 years of recorded history and who knows how many billions of words of spilled ink on the matter, probably more than on any other subject in history, there is no universal agreement on morality.

Re: GPT-4 is getting worse over time, not better

#236

Earlier quoted context omitted.

I don’t believe this is true. It’s possible I was blinded by the light, but my programming tasks were previously (during the early access program) being handled by GPT-4 regularly and now they aren’t. I’ve also seen many anecdotes from engineers who had exceptionally early access before GPT-4 was public knowledge. The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back…

> terrifyingly unaligned Honestly, if people think that a statistical language model is "terrifying" because it can verbalise the concept of a mass killing, they need to give their heads a wobble. My text editor can be used to write "set off a nuclear weapon in a city, lol". Is Notepad++.exe terrifying? What about the Sum of All Fears ? I could get some pointers from that. Is Tom Clancy unaligned? Am I terrifying bec…

Tom Clancy is, because his games need to keep screaming "HEY THIS IS BY TOM CLANCY, OK? LOOK AT ME, TOM CLANCY, I'M BEING TOM CLANCY!" in their titles.

Re: GPT-4 is getting worse over time, not better

#237
post #7
post #5

Earlier quoted context omitted.

That is blatantly false, these models only work as well as they do in the first place because of instruction fine tuning. The raw models are much harder to work with.

Lets not conflate a model's smartness with a model's ability to be used reliably and predictably. They're very different things. Yes, fine-tuning, for instruction, or whatever you chose, will make the models easier to work with with less going off the rails. But they'll also take a big hit on perplexity and other more robust dataset completion tests. They're objectively stupider.

Instruction fine tuning improves performance on many tasks that are not in the fine tuning dataset. See e.g. table 14 in the instructGPT paper.

Re: GPT-4 is getting worse over time, not better

#238

Earlier quoted context omitted.

These models have been explicitly nerfed since their first release due to copyright considerations. I've mentioned in two previous cases both for [1] code generation and [2] book summarizing. From my point of view, it is sad that these sort of socio-political constructs (copyright) are hindering innovation. The funny thing is that in say, 10 years, the "pirate" version of LLMs will be way more powerful and useful tha…

> From my point of view, it is sad that these sort of socio-political constructs (copyright) are hindering innovation Why not pay authors of the data the LLM has ingested?

Since they are trained on all kinds of Internet content, might be a little tricky to manage paying >1B people

Re: GPT-4 is getting worse over time, not better

#239

Earlier quoted context omitted.

> From my point of view, it is sad that these sort of socio-political constructs (copyright) are hindering innovation Why not pay authors of the data the LLM has ingested?

Don't a lot of content creators expect payment per use? Like every time someone streams a song on sportify, the artist gets a few pennies. So should they be paid a few pennies every time the LLM spits out a response that "used" that training data? And I am pretty sure it's not even possible to really link the output back to training data anyways.

> Like every time someone streams a song on sportify, the artist gets a few pennies

You're thinking of broadcast radio. Streaming is a fraction of a penny!

> So should they be paid a few pennies every time the LLM spits out a response that "used" that training data?

That doesn't sound unreasonable! I understand that current LLMs have no way to report "this token came from this data" but that doesn't mean it's impossible to build. (Ack: this is a full-on proper dunning-kruger, having not looked into it & having zero knowledge of the field.)

But I think it's probably more reasonable to simply split a % of the service revenue across everyone whose data was used. Or pay an up-fee for ingesting the data in the first place.

Generally, it's crazy to name that some people think it's reasonable for these companies to pay for GPUs and CPUs and electricity to run them but not for the data that's the actual core of their service.

Re: GPT-4 is getting worse over time, not better

#240
post #238

Earlier quoted context omitted.

> From my point of view, it is sad that these sort of socio-political constructs (copyright) are hindering innovation Why not pay authors of the data the LLM has ingested?

Since they are trained on all kinds of Internet content, might be a little tricky to manage paying >1B people

I mean, figure something out?

Didn't Sam want to create UBI? That's one way!

Post reply on HN