LLM's integrated into any real product requires a model hash, otherwise the provider of the model has full control over any sort of deception.
Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
61–70 of 817 posts
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#62Yes! It didn't even try on my question of Jarvis standings desks, which is a fairly old product that hasn't changed up.. Their typical "My knowledge cutoff..." response doesn't even make sense. It screwed up another question I asked it about server uptime and four-9s, Bard got it right. I've moved back to Bard for the time being...It's way faster as well. And GPT-4's knowledge cutoff thing is getting old fast. Exampl…
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#63Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#64Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#65Earlier quoted context omitted.
I’d like to see a model with the effluent of the internet intelligently filtered from the pretraining data by LLM and human curation, and much more effort to include digitised archival sources and the entirety of books and high quality media transcripts. I imagine it would yield far better baseline quality outputs with much less than current “requirements” for (over)correction with ultimately disastrous RLHF masking.
I'd love to play with a version of GPT 4 fine-tuned with every science textbook written in the last few decades, every published science paper (not just preprints from ArXiV), and everything generated by every large research institute. Think NASA, CERN, etc... Or one tuned with every fiction novel ever written, along with every screenplay.
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#66Yes. Before the update, when its avatar was still black, it solved pretty complex coding problems effortlessly and gave very nuanced, thoughtful answers to non-programming questions. Now it struggles with just changing two lines in a 10-line block of CSS and printing this modified 10-line block again. Some lines are missing, others are completely different for no reason. I'm sure scaling the model is hard, but they l…
You never had access to that original. Watch this talk by one of the people that integrated GPT-4 in Bing telling how they noticed GPT-4 releases they got from OpenAI got iteratively and significantly nerfed even during the project.
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#67Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#68Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#69This is the normal workflow for drug dealers too. The first fix is free. The second one will cost you money. The third one will be laced with fillers and have degraded quality.
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#70Yes! It didn't even try on my question of Jarvis standings desks, which is a fairly old product that hasn't changed up.. Their typical "My knowledge cutoff..." response doesn't even make sense. It screwed up another question I asked it about server uptime and four-9s, Bard got it right. I've moved back to Bard for the time being...It's way faster as well. And GPT-4's knowledge cutoff thing is getting old fast. Exampl…
> The fully assembled Jarvis Bamboo Standing Desk weighs 92 pounds. The desktop itself weighs 38 pounds, and the frame weighs 54 pounds. The desk can hold a maximum weight of 350 pounds.
That sounds like a linguistically valid sentences, exactly what I would expect from a novel LLM. Did you check that it is also factually correct? Factually correctness is _not_ the goal of a typical LLM.