Live data from Hacker News

Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

news.ycombinator.com

571–580 of 817 posts

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#571
post #257

Earlier quoted context omitted.

They're up against a pretty difficult barrier - if we had a perfect all-knowing oracle it might easily have opinions that are racist. Statistics alone suggest there will be racist truths. We're dealing with groups of people who are observably different from each other in correlated ways. GPT would need to reach a convincing balance of lying and honesty if it is supposed to navigate that challenge. It'd have to be dee…

[flagged]

> the actual reason (Black men tend to be larger and faster, which are useful)

If that's the case, why aren't NHL players mostly Black? Being larger and faster helps there too. I actually agree that small differences in means of normal distributions lead to large differences at the tail end, which amplifies the effect of any genetic differences, racial included. But clearly that's only one reason, not the reason -- and it's not even the most important, or the NHL would look similar.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#572

Earlier quoted context omitted.

Who has the necessary resources to run, let alone train the model?

folding@home has been doing cool stuff for ages now. There's nothing to say that distributed computing couldn't also be used for this kind of stuff, albeit a bit slower and fragmented than running on a huge clusters of H100 with NVLink. In terms of training feedback I suppose there's a few different ways of doing it. Gamification, mech turk, etc. Hell free filesharing sites could get on the action and have you comple…

[deleted]

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#573

OpenAI's models feel 100% nerfed to me at this point. I had it solving incredibly complex problems a few months ago (i.e. write a minimal PDF parser example), but today you will get scolded for asking such a complicated task of it. I think they programmed a classifier layer to detect certain coding tasks and shut it down with canned BS. I like to imagine certain billion/trillion-dollar mega corps had a back-room say…

`Microsoft is a big stakeholder and they might not want to get sued...`

Wait I thought Microsoft has bought the right to get profits from OpenAI, not actually shares in the company? Can someone correct me if I'm wrong>

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#574

Earlier quoted context omitted.

Currently, not at all. You need low latency, high bandwidth links between the GPUs to be able to shard the model usefully. There is no way you can fit an 1T (or whatever) parameter model on a MacBook, or any current device, so sharding is a requirement. Even if it that problem disappeared, propagating the model weight updates between training steps poses an issue in itself. It's a lot of data, at this size.

What if you split up the training down to the literal vector math, and treated every macbook like a thread in a gpu, with just one big computer acting as the orchestrator?

You would need each MacBook to have an internet connection capable of multiple terabytes per second, with sub millisecond latency to every other MacBook.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#575
post #410

Earlier quoted context omitted.

I’m tired of the “it’s not racist if aggregate statistics support my racism” thing. Racism, like other isms, means a belief that a person’s characteristics define their identity. It doesn’t matter if confounding factors mean that you can show that people of their race are associated with bad behaviors or low scores or whatever. I used GPT3.5 to generate 100 short descriptions of families for a project. Every single o…

What was your prompt? LLMs take previous output into account when generating the next token. If it had already output 20 families of a similar shape, number 21 is more likely to match that shape.

Multiple one-shot prompts with no history. I don't have the exact prompt handy but it was something like "Create a short biography of a family, summarizing each person's age and personality".

I just ran that prompt 3 times (no history, new sessions, that prompt for first query) and got:

1. Hard-working father, stay at home mother, artistic daughter, adventurous son, empathic ballet-loving daughter

2. Busy architect father, children's book author mother, environment- and animal-loving daughter, technology-loving son, dance-loving daughter

3. Hard-working engineer father, English-teaching mother, piano- and book-loving daughter, basketball- and technology-loving son, comedic dog (!)

I'm summarizing because the responses were ~500 words each. But you can see the patterns: fathers work hard (and come first!), mothers largely nurture, daughters love art and dance, sons love technology.

It's not the end of the world, and as AI goes this is relatively harmless. But it is a pretty deep bias and a reminder that AI reflects implicit bias in training materials and feedback. You could make as many families as you want with that prompt and it will not approximate any real society.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#576
post #536

Earlier quoted context omitted.

No, OpenAI 100% pushed an update recently that is very noticeable where they basically "nerfed" the responses. I'm sure it was a business decision (either costing them too much to give away the kitchen sink for free or they want to turn around and charge more for the "really good, old" version of GPT-4 ) but you can visually see where the bot used to try to answer complex tasks, now it has an extra layer where it say…

>OpenAI 100% pushed an update recently that is very noticeable where they basically "nerfed" the responses >they want to turn around and charge more for the "really good, old" version of GPT- If this was the case it is provable by submitting queries to `gpt-4-0314` and comparing them to `gpt-4`.

To me, it doesn't feel like the gpt-4 in the playground (and in the API) is the same as the model used in chatgpt. It's hard to prove though since both are nondeterministic.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#577
post #200

Earlier quoted context omitted.

Same here. If I have a choice between honesty and political correctness, I always pick honesty.

What makes you think the "unaligned" version necessarily has more honesty? Rather than just being generally easier to prompt to say whatever the user wants it to say, true or not, horrible or not? Or even easier to unintentionally make it confabulate/hallucinate stuff? Does not seem to follow, and does not seem to be a true dichotomy. Edginess does not equal honesty.

I always prefer a model that can be prompted to say anything I want over a model that can only say things that a centralized corporation with political ties wants it to say.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#578
I think it is a combination of issues (from most unlikely to most likely):

- They might've not been prepared for the growth;

- They were prepared, but decided to focus on the 95% of "regular people with easy questions" that treat it as a curiosity instead of the fewer people with difficult questions. Most regular people have no idea who is OpenAI, what a LLM is or how a GPT works, but they know "ChatGPT, the brand". Since it became a household name so quickly, it would be far better that the AI is just a little underwhelming sometimes than for it to not be able to serve so many people.

- The corpus used to generate it was trained on a staggering amount of content. This includes fringe and unmoderated content. Imagine you asked a question about WW2, and being trained, lets say on 4chan, the model responds with a very charitable bias about the solid reasoning behind the reich's actions at the time... It does not look good for investors, for the company, and attracts scrutiny. Even more innocuous themes are enough to invite all kinds of bad faith debate, radical criticism and whatnot... and the darling of the "IA revolution" certainly does not need (or want) coverage outside of their roadmap wishes.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#579

I believe they introduced a sort of rate limiting where the expectation will be that the user is doing more thinking and due diligence and asking more precise questions, so that they're not just attempting to get a lot of hand holding with very broad prompts when they can otherwise think about what code they're being given back from more specific and structured questioning. This is useful because it will preserve the…

> it will preserve the value proposition of AI while avoiding a machine just being milked

What's the point of a dairy cow that can't be milked? The whole point of AI is to milk it.

> the user is doing more thinking and due diligence and asking more precise questions

If I need to do "due diligence" and learn how to ask exactly precise questions to get useful answers, why am I using the AI at all? The actual "value proposition" of AI is that I don't need to learn how to use it. I'm not going to learn Structured Prompting Language just to being the process of having an AI show me how to learn some new programming language; it's easier to just learn that programming language instead.

> Someone who is non-technical should not be given a firecracker before they've ever even just turned on a microwave or a stove.

"We found a next-level genius Einstein physicist, but we have to make sure the public is not allowed to ask him any questions! He must be kept secret at all times! It is too dangerous to allow any fool with a high-school physics textbook to ask just a powerful genius a question!"

See the problem?

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#580

Earlier quoted context omitted.

I’m tired of the “it’s not racist if aggregate statistics support my racism” thing. Racism, like other isms, means a belief that a person’s characteristics define their identity. It doesn’t matter if confounding factors mean that you can show that people of their race are associated with bad behaviors or low scores or whatever. I used GPT3.5 to generate 100 short descriptions of families for a project. Every single o…

> a person’s characteristics define their identity They do though. Your personality, culture and appearance are the main components of how people perceive you, your identity. The main thing you can associate with bad behaviour is domestic culture. It's not racist to say that African Americans have below-average educational attainment and above-average criminality. This is as contrasted to African immigrants to Americ…

> Your personality, culture and appearance are the main components of how people perceive you, your identity

I'm not sure if this is bad rhetoric (defining identity as how you are perceived rather than who you are) or if you really think of your own identity as the judgements that random people make about you based on who knows what. Either way, please rethink.

> Your personality, culture and appearance are the main components of how people perceive you, your identity

Ah, so if you asked for 100 numbers between 1-100, there's no reason not to expect 100 numbers very close to 50?

> Why do you think that every sample has to be implicitly representative of the US population?

That is a straw man that I am not suggesting. I am suggesting that there should be some variation. It doesn't have to represent the US population, but can you really think of ANY context where a sample of 100 families turns up every single one having one male and one female parent, who are still married and alive?

You're bringing a culture war mindset to a discussion about implicit bias in AI. It's not super constructive.

Post reply on HN