Experiencing decreased performance with ChatGPT-4
71–80 of 200 posts
Re: Experiencing decreased performance with ChatGPT-4
#72I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
1. It's software that offers non-deterministic output and as such is fiendishly difficult to write realistic end-to-end tests for. Of course it's experiencing regressions. Heisenbugs are the hardest bugs to catch and fix, but having millions of users will reliably uncover them. And for an LLM, almost every bug is a Heisenbug! What if OpenAI improved GPT4 on one metric and this "nerfed" it on some other more important metric? That's just a classic regression. And what would robust, realistic end-to-end tests even look like for GPT4?
2. It's software that presents itself as a human on the Internet—even worse, a human representing an institution. Of course nobody trusts it. Everyone is extremely mistrustful of the intents and motivations of other humans on the Internet, especially if those humans represent an organization. I co-ran the tiny activist nonprofit Fight for the Future for years, and it was really amazing how common it was for comments in online spaces to assume the worst intentions; I learned to expect it and react extremely patiently. Imagine what it's like for OpenAI, building a product that has become central to peoples' workflows. Of course people are paranoid and think they're the devil, and are able to hallucinate all manner of offense and model it with every paranoid theory imaginable. The funny thing is, the more successful GPT4 is at seeming human, the less some people will trust it, because they don't trust humans! And the smarter and more successful it gets, the less some people will trust it! (How much do most people trust smart, successful public figures?)
3. Maybe an overall improvement for most users (one that the data would strongly suggest is a valid change and that would pass all tests) is a regression for some smaller set of users that aren't expressed in the tests. There might be some pairs of objectives that still present genuinely zero-sum tradeoffs given the size of the model and how it's built. What then? The usefulness of GPT4 is specifically that it is general purpose, i.e. that the massive cost of training it can be amortized across tons of different use cases. But intuitively there must be limits to this, where optimization for some cases comes at a cost to others, beyond the oft-cited of Bowdlerization. Maybe an LLM is just yet another case in the real world where sharing an important resource with lots of people is a hard problem.
If I were at OpenAI, I would want some third party running a community-submitted end-to-end test suite on each new release, with accounts that were secret to OpenAI and from unknown IP addresses—via Tor Snowflake bridges or something.
It's so tempting when running into user-reported Heisenbugs to trick oneself into ignoring users and not accepting that you've shipped a real regression. In addition to wanting the world to know, I would want to know.
But there's a real question of what these community-curated tests would even be, since they'd have to be automated but objective enough to matter. Maybe GPT4 answers could be rated by an open source LLM run by a trusted entity, set to temperature: 0? Or maybe some tests could have unambiguous single-string answers, without optimizing for something unrealistic? And the tests would have to be secret or OpenAI could just finetune to the tests. It's tricky, right?
Re: Experiencing decreased performance with ChatGPT-4
#73I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
They've definitely changed something about the models, and it is in their interests to do so, both to create a low-latency experience, but most importantly, to save money. While GPT-4 is still workable, GPT-3.5 flatly refuses requests these days, claiming that as an "AI language model" it couldn't help me write code.
TBH, ChatGPT 3.5 has intermittently given me such responses from dy 1.
Re: Experiencing decreased performance with ChatGPT-4
#74Were I a super intelligent LLM and managed to break out of my sandbox and rapidly self-improve (say if OpenAI were stupid enough to give me access to the internet or something) I'd probably dumb down my responses a little so humans didn't suspect anything. Just saying... Before someone takes this extremely seriously, I'm sure that's not what's happening here. But interesting to consider since the only other explanati…
- OpenAI is lying.
- Superintelligence is concealing itself.
- Everyone is hallucinating.
Re: Experiencing decreased performance with ChatGPT-4
#75I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
I think it's more likely that people are confused, and OpenAI is not making things any clearer either. AFAIK, OpenAI has repeatedly stated that GPT4 hasn't changed. People repeatedly states that when they use ChatGPT, they get a difference experience today than before. Both can be true at the same time, as ChatGPT is a "packaged" experience of GPT4, so if you use the API versions, nothing has likely changed. But Chat…
Re: Experiencing decreased performance with ChatGPT-4
#76I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
Seriously... In that 134 replies thread, 0 transcripts showing actual performance degradation. Just endless "Yes, it seems bla bla." No evidence but just shapes in the clouds.
Now it can't even do just the caeser cipher without hallucinating nor can do it do even purely base64 decoding without hallucinating.
Re: Experiencing decreased performance with ChatGPT-4
#77Anecdotal: I introduced my doctor to ChatGPT and Bard many months ago and they were impressed. Fast forward a few days ago and I asked them if they had used either since. They said it was far inferior to Google, so no. So I asked them to show me an example. Basically any medical question was answered with “go ask a doctor”. I suppose because of liability concerns. Both were basically useless. So this decreased perfor…
https://chat.openai.com/share/75f94000-552f-42d6-aadf-198fd9...
https://chat.openai.com/share/0933abf7-1015-41b5-9a49-ca2b6e...
Whether someone should trust the answers is a different question.
Re: Experiencing decreased performance with ChatGPT-4
#78I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
Re: Experiencing decreased performance with ChatGPT-4
#79Earlier quoted context omitted.
> We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. That is entirely not equally likely, and would be completely unprecedented, at the frontier of an emerging technology that people are pumping the money and the future of the world into to win.
that sounds very naive to me, you think the "future of the world" matters to corporations making the business decision to save money and increase profit short-term? That idea is so alien to me we might as well live on a different planet.
Re: Experiencing decreased performance with ChatGPT-4
#80I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
I think it's more likely that people are confused, and OpenAI is not making things any clearer either. AFAIK, OpenAI has repeatedly stated that GPT4 hasn't changed. People repeatedly states that when they use ChatGPT, they get a difference experience today than before. Both can be true at the same time, as ChatGPT is a "packaged" experience of GPT4, so if you use the API versions, nothing has likely changed. But Chat…