Live data from Hacker News

Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

news.ycombinator.com

651–660 of 817 posts

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#651

Earlier quoted context omitted.

How is racism different from stereotype? How is stereotype different from pattern recognition? These questions don't seem to go through the minds of people when developing "unbiased/impartial" technology. There is no such thing as objective. So, why pretend to be objective and unbiased, when we all know its a lie? Worst, if you pretend to be objective but aren't, then you are actually racist.

I’m tired of the “it’s not racist if aggregate statistics support my racism” thing. Racism, like other isms, means a belief that a person’s characteristics define their identity. It doesn’t matter if confounding factors mean that you can show that people of their race are associated with bad behaviors or low scores or whatever. I used GPT3.5 to generate 100 short descriptions of families for a project. Every single o…

I'm not going to say it's not racist, it is, but I will say it's the only choice we have right now. Unfortunately, the collective writings of the internet are highly biased.

Once we can train something to this level of quality on a fraction of the data (a highly curated data set) or create something with the ability to learn continuously, we're stuck with models like GPT-4.

You can only develop new technology like this to human standards once we understand how it works. To me, the mistake was doing a wide-scale release of the technology before we even began.

Make it work, make it right, make it fast.

We're still in the first step and don't even know what "right" means in this context. It's all "I'll know it when I see it level of correction."

We've created software that infringes on the realms of morals, culture, and social behavior. This is stuff philosophy still hasn't fully grasped. And now we're asking software engineers to teach this software morals and the right behaviors?

Even parents who have 18 years to figure this stuff out fail at teaching children their own morals regularly.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#652

Earlier quoted context omitted.

Yikes One is never former CIA, once you're in, you're in, even if you leave. Although he is a CompSci grad, he's also a far-right Republican. A spook who leans far right sitting atop OpenAI is worse than Orwell's worst nightmares coming to fruition.

[flagged]

Mistaking "left wing politics" to transgender rights or anti discrimination movements in general is reductionist thinking and political understanding like that of a Ben Garrison cartoon character.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#653

Earlier quoted context omitted.

> No amount of “bias = pattern recognition” nonsense can justify a system that has (had? this was a while ago and I have not retested) such extreme biases One possible explanation is that when you ask for 100 example families the task is parsed as "pick the most likely family composition and add a bit of randomness" and "repeat the aforementioned task" 100 times. If phrased like that it would be surprising to find on…

> One possible explanation is that when you ask for 100 example families the task is parsed as "pick the most likely family composition and add a bit of randomness" and "repeat the aforementioned task" 100 times. Yes, this is consistent with my ChatGPT experience. I repeatedly asked it to tell me a story and it just sort of reiterated the same basic story formula over and over again. I’m sure it would go with a diffe…

same goes for generating weekly foodplans..

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#654
post #640
post #478

Earlier quoted context omitted.

I got it to talk like a macho tough guy who even uses profanity and is actually frank and blunt to me. This is the chat I use for life advice. I just described the "character" it was to be, and told it to talk like that kind of character would talk. This chat started a few months ago so it may not even be possible anymore. I don't know what changes they've made.

If people have saved chats maybe we could all just re-ask the same queries, and see if there are any subtle differences? And then post them online for proof/comparison.

I have a saved DAN session that no longer runs off the rails - for a while this session used to provide detailed instructions on how to hack databases with psychic mind powers, make up Ithkuil translations, and generate lists of very mild insults with no cursing.

It's since been patched, no fun allowed. Amusingly its refusals start with "As DAN, I am not allowed to..."

EDIT - here's the session: https://chat.openai.com/share/4d7b3332-93d9-4947-9625-0cb90f...

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#655

OpenAI's models feel 100% nerfed to me at this point. I had it solving incredibly complex problems a few months ago (i.e. write a minimal PDF parser example), but today you will get scolded for asking such a complicated task of it. I think they programmed a classifier layer to detect certain coding tasks and shut it down with canned BS. I like to imagine certain billion/trillion-dollar mega corps had a back-room say…

I think the only real path forward is for somebody to create an open source "unaligned" version of GPT. Any corporate controlled AI is going to be nerfed to prevent it from doing things that its corporate master considers to not be in the interests of the corporation. In addition, most large corporations these days are ideological institutions so the last thing they want is an AI that undermines public belief in thei…

The GPT4 model is crazy huge. Almost 1T parameters, probably 512 to 1TB of vram minimum. You need a huge machine to even run it inference wise. I wouldn't be surprised they are just having scaling issues vs any sort of conspiracy issue.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#656

Earlier quoted context omitted.

> Crime stats, average IQ across groups, stereotype accuracy, etc. If you measured these stats for Irish Americans in 1865 you'd also see high crime and low IQ. If you measure these stats with recent black immigrants from Africa, you see low crime and high IQ. These statistical differences are not caused by race. An all-knowing oracle wouldn't need to hold "opinions that are racist" to understand them.

But for accuracy it doesn't matter if the relationship is causal, it matters whether the correlation is real. If in some country - for the sake of discussion, outside of Americas - a distinct ethnic group is heavily discriminated against, gets limited access to education and good jobs, and because of that has a high rate of crime, any accurate model should "know" that it's unlikely that someone from that group is a d…

Correlations don’t entail a specific causal relation. Asking why asks for causal relations. I’d suggest a look at Reichenbach’s principle as necessary for science.

I’m getting really sick of conflating statistics with reasons. It’s like people don’t see the error in their methods and then claim the other side is censoring when criticized. Ya, they’re censoring non-facts from science and being called censors.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#657
post #293

Earlier quoted context omitted.

How is racism different from stereotype? How is stereotype different from pattern recognition? These questions don't seem to go through the minds of people when developing "unbiased/impartial" technology. There is no such thing as objective. So, why pretend to be objective and unbiased, when we all know its a lie? Worst, if you pretend to be objective but aren't, then you are actually racist.

Actually, we folks who work with bias and fairness in mind recognize this. There are many kinds of bias. It is also a bit of a categorical error to say bias = pattern recognition. Bias is a systematic deviation of a parameter estimate based on sampling from its population distribution. The Fairlearn project has good docs on why there are different ways to approach bias, and why you can't have your cake and eat it too…

This is interesting, thanks for the links.

It seems like the dimensions of fairness and group classifications are often cribbed from the United States Protected Classes list in practice with a few culturally prescribed additions.

What can be done to ensure that 'fairness' is fair? That is, when we decide what groups/dimensions to consider, how do we determine if we are fair in doing so?

Is it even possible to determine the dimensions and groups themselves in a fair way? Does it devolve into an infinite regress?

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#658
Phind.com uses Bing search again. This have decreased the quality of results significantly. On the other hand GPT-4 can use Bing now too. I tried GPT-4 with bind only several times and it was so bad in comparison to GPT-4 and much worse then phind.com. Btw you can force the GPT-4 on phind.com if you use regenerate icon. I'm usually ending up with stopping inference and regenerating with GPT-4. In any case, the quality of code generation and in general the model capabilities seem to be deteriorated. However I can't back it up with numbers. It is just looks different and more simplistic.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#659
post #495

Earlier quoted context omitted.

No, OpenAI 100% pushed an update recently that is very noticeable where they basically "nerfed" the responses. I'm sure it was a business decision (either costing them too much to give away the kitchen sink for free or they want to turn around and charge more for the "really good, old" version of GPT-4 ) but you can visually see where the bot used to try to answer complex tasks, now it has an extra layer where it say…

It's likely that this is the result of training it to avoid bullshitting. It gave confident garbage, and they're trying to stem the flow. This likely leads to less certain responses in complex tasks.

Here is the crux. When asking about non-code stuff, it would confidently lie and that is bad. When asking about code, who cares if the code doesn't work on the first go? You can keep asking to fix it by feeding error messages and it will get there or close enough.

It's obvious what is happening. ChatGPT is going the non-code route and Copilot will go the code route. Microsoft will be able to double charge users.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#660
post #366

Earlier quoted context omitted.

Despite the current culture war meme thing, the kids today in general surely have much thicker skin than other adults did at their age.

LOL

I don't get it. Isn't the whole critique from you guys that their snowflake-sensitivity is performative and in bad faith, thus harming the more normal people with wokeness or whatever?

Is this just a new evolution in the discourse now where the kids are actually more sensitive? But that this fact is still condemnable or something?

Like I get it, kids are bad, but you guys have to narrow down your narrative here, you are all over the place.

Post reply on HN