Live data from Hacker News

Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

news.ycombinator.com

201–210 of 817 posts

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#201
post #176

Earlier quoted context omitted.

They can only do this in the USA. The battle bottleneck is going to be GPUs and countries that have access to it will leap ahead of the US if arbitrary regulations are put in place to neuter LLMs in the US.

You can lobby in every country, what do you mean

I think the healthcare industry case in the US is a great example of how regulatory capture works against the consumer.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#202
This was my final attempt when I asked for ASCII art of a person standing on a tall bed.

I apologize for the confusion. Here's another attempt at creating an ASCII art representation of a person standing on a tall bed:

        __
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
       |  |
  _____|__|_____
  Bed         Person

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#203
You can still run the original gpt-4-0314 model (March 14th) on the API playground:

https://platform.openai.com/playground?mode=chat&model=gpt-4...

Costs $0.12 per thousand tokens (~words), and I find even fairly heavy use rarely exceeds a dollar a day.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#204

This was my final attempt when I asked for ASCII art of a person standing on a tall bed. I apologize for the confusion. Here's another attempt at creating an ASCII art representation of a person standing on a tall bed: __ | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | _____|__|_____ Bed Person

Its trolling you ... right?

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#205
post #201

Earlier quoted context omitted.

You can lobby in every country, what do you mean

I think the healthcare industry case in the US is a great example of how regulatory capture works against the consumer.

Okay. you can do regulatory capture in many countries too

You’re just saying random things about US industries as if it’s insightful because you saw a documentary once, well that’s what its reading like

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#206
post #124

The reason it's worse is basically because it's more 'safe' (not racist, etc). That of course sounds insane, and doesn't mean that safety shouldn't be strived for, etc - but there's an explanation as to how this occurs. It occurs because the system essentially does a latent classification of problems into 'acceptable' or 'not acceptable' to respond to. When this is done, a decent amount of information is lost regardi…

If I run a model like LLaMA locally would it be subject to the same restrictions? In other words is the safety baked into the model or a separate step separate from the main model?

Both approaches are valid, but I would hope they are using a separate model to validate responses, rather than crippling the base model(s). In OpenAI's case, we don't know for sure, but it seems like a combination of both, resulting in lower quality responses overall.

I imagine LLaMA was fed highly-vetted training data, as opposed to being "fixed" afterwards.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#207
post #200

Earlier quoted context omitted.

I would immediately jump to an AI not being "aligned" by SF techies (or anyone else).

Same here. If I have a choice between honesty and political correctness, I always pick honesty.

It's not about honesty vs. political correctness, it is about safety. There's real concern that the model can cause harm to humans, in a variety of ways, which is and should be unethical. If we have to argue about that in 2023, that's concerning.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#208

To me, it feels like it's started giving superficial responses and encouraging follow-up elsewhere -- I wouldn't be surprized if its prompt has changed to something to that effect. Before, if I had an issue with a library or debugging issue, it would try to be helpful and walk me through potential issues, and ask me to 'let it know' if it worked or not. Now it will try to superficially diagnose the problem and then a…

>To me, it feels like it's started giving superficial responses and encouraging follow-up elsewhere -- I wouldn't be surprized if its prompt has changed to something to that effect. That's the vibe I've been getting. The responses feel a little cagier at times than they used to. I assume it's trying to limit hallucinations in order to increase public trust in the technology, and as a consequence it has been nerfed a…

Seems like the metric they're optimising for is reducing the number of bad answers, not the proportion of bad answers, and giving non-answers to a larger fraction of questions will achieve that.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#209
I had a very strange experience yesterday where I asked about git authentication, explicitly telling Bing that I was asking about git itself rather than GitHub, but I was not using GitHub and not to include results referencing GitHub. Bing did not understand.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#210

Is it consistently worse or just sometimes/often worse than before? Any extreme power users or GPT-whisperers here? If it’s only noticeably worse X% of the time my bet would be experimentation. One of my least favorite patterns that tech companies do is use “Experimentation” overzealously or prematurely. Mainly, my problem is they’re not transparent about it, and it creates an inconsistent product experience that jus…

It is not worse for me. I do notice the novelty has worn off. Asking chatGPT4 about why people would think this on here I think it nails it with the novelty effect lol:

"Indeed, the performance of an AI model like ChatGPT doesn't deteriorate over time. However, human perception of its performance can change due to a variety of psychological factors:

Expectation Bias: As users become more familiar with AI capabilities, their expectations may increase over time. When AI doesn't meet these heightened expectations, users might perceive this as a decline in performance.

Novelty Effect: At the beginning, the novelty of interacting with an AI could lead to positive experiences. However, as the novelty wears off, users may start to focus more on the limitations, creating a perception of decreased performance."

Without this thread I would have said it got stronger with the May 12th update. I don't think that is really true though. There is this random aspect of streaks in asking questions it is good at answering vs streaks of asking questions it is less good at answering.

Post reply on HN