Live data from Hacker News

Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

news.ycombinator.com

61–70 of 817 posts

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#61
It's untenable to simply trust LLM API providers that the models they are serving through an API endpoint is the model they claim it is. They could easily switch the model with a cheaper one whenever they wanted, and since LLM outputs are non-deterministic (presuming a random seed), it would be impossible to prove this.

LLM's integrated into any real product requires a model hash, otherwise the provider of the model has full control over any sort of deception.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#62
post #4

Yes! It didn't even try on my question of Jarvis standings desks, which is a fairly old product that hasn't changed up.. Their typical "My knowledge cutoff..." response doesn't even make sense. It screwed up another question I asked it about server uptime and four-9s, Bard got it right. I've moved back to Bard for the time being...It's way faster as well. And GPT-4's knowledge cutoff thing is getting old fast. Exampl…

Were the numbers accurate or hallucinated?

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#65
post #38

Earlier quoted context omitted.

I’d like to see a model with the effluent of the internet intelligently filtered from the pretraining data by LLM and human curation, and much more effort to include digitised archival sources and the entirety of books and high quality media transcripts. I imagine it would yield far better baseline quality outputs with much less than current “requirements” for (over)correction with ultimately disastrous RLHF masking.

I'd love to play with a version of GPT 4 fine-tuned with every science textbook written in the last few decades, every published science paper (not just preprints from ArXiV), and everything generated by every large research institute. Think NASA, CERN, etc... Or one tuned with every fiction novel ever written, along with every screenplay.

I would gladly pay triple digits a month for exactly that.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#66
post #14

Yes. Before the update, when its avatar was still black, it solved pretty complex coding problems effortlessly and gave very nuanced, thoughtful answers to non-programming questions. Now it struggles with just changing two lines in a 10-line block of CSS and printing this modified 10-line block again. Some lines are missing, others are completely different for no reason. I'm sure scaling the model is hard, but they l…

"The original GPT-4 felt like magic to me"

You never had access to that original. Watch this talk by one of the people that integrated GPT-4 in Bing telling how they noticed GPT-4 releases they got from OpenAI got iteratively and significantly nerfed even during the project.

https://www.youtube.com/watch?v=qbIk7-JPB2c

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#68
post #47
post #40

Earlier quoted context omitted.

“Bard isn’t currently supported in your country. Stay tuned!”

Google's passion for region locking is insane to me

Its a legal thing, not something they want to do

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#69

This is the normal workflow for drug dealers too. The first fix is free. The second one will cost you money. The third one will be laced with fillers and have degraded quality.

I wonder how much people are relying on it already, to what extent and so on. Would be a good study.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#70
post #4

Yes! It didn't even try on my question of Jarvis standings desks, which is a fairly old product that hasn't changed up.. Their typical "My knowledge cutoff..." response doesn't even make sense. It screwed up another question I asked it about server uptime and four-9s, Bard got it right. I've moved back to Bard for the time being...It's way faster as well. And GPT-4's knowledge cutoff thing is getting old fast. Exampl…

  > The fully assembled Jarvis Bamboo Standing Desk weighs 92 pounds. The desktop itself weighs 38 pounds, and the frame weighs 54 pounds. The desk can hold a maximum weight of 350 pounds.
That sounds like a linguistically valid sentences, exactly what I would expect from a novel LLM. Did you check that it is also factually correct? Factually correctness is _not_ the goal of a typical LLM.
Post reply on HN