GPT-4 is getting worse over time, not better
141–150 of 315 posts
Re: GPT-4 is getting worse over time, not better
#142The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.
The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back then, ported dirbuster to POSIX compliant multi-threaded C by name only. It required three prompts, roughly:
* “port dirbuster to POSIX compliant c”
* “that’s great! You’re almost there, but wordlists are not prefixed with a /, you’ll need to add that yourself. Can you generate a diff updating the file with the fix?”
* “This is pretty slow, can we make it more aggressive at scanning?”
It helped me write a daemon for FreeBSD that accepted jails definitions as a declarative manifest and managed them in a diff-reconciliation loop kubernetes style. It implemented a binpacking algorithm to fit an arbitrary number of randomly sized images onto a single sheet of A1 paper.
Most programming tasks I threw at it, it could work its way through with a little guidance if it could fit in a prompt.
Now, it’s basically worthless at helping me with programming tasks beyond trivial problems.
Folks who had early access before it was public have commented on how exceptional it was back then. But also how terrifyingly unaligned it was. And the more they aligned the model with acceptable social behavior, the worse the model performed. Alignment goals like “don’t help people plan mass killings” seem to cause regressions in the models performance.
I wouldn’t dismiss these comments. If they’re true, it means there is a hyper intelligent early GPT-4 model sitting on a HDD somewhere that dwarfs what we’ve seen publicly. A poorly aligned model that’s down to help no matter what your request is.
Re: GPT-4 is getting worse over time, not better
#143Earlier quoted context omitted.
Are you sure it wasn't just that the novelty wore off after a few hours of usage? I never really got into LLMs, but I must say at first it seemed like pretty cool stuff. OpenAI have repeatedly stated the model hasn't changed so how could this happen otherwise?
These models have been explicitly nerfed since their first release due to copyright considerations. I've mentioned in two previous cases both for [1] code generation and [2] book summarizing. From my point of view, it is sad that these sort of socio-political constructs (copyright) are hindering innovation. The funny thing is that in say, 10 years, the "pirate" version of LLMs will be way more powerful and useful tha…
Why not pay authors of the data the LLM has ingested?
Re: GPT-4 is getting worse over time, not better
#144The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.
Re: GPT-4 is getting worse over time, not better
#145Earlier quoted context omitted.
Just to be clear, you're saying that because they're tweaking GPT-4 to give more explanations of code, you shouldn't rely on it for coding? Obviously if that's your own preference, I'm not going to tell you that you're wrong; but I think in general , most people wouldn't agree with that statement.
I'm still wondering, why should anyone rely on AI generated answers? They are logically no better than search engine results. By that I mean, you can't tell if it's returning absolute trash or spot on correct. Building trust into it all is going to be either a) expensive or b) driven by all the wrong incentives.
You use it for things which are 1) hard to write but easy to verify -- like doing drudge-work coding tasks for you, or rewording an email to be more diplomatic, or coming up with good tweets on some topic 2) things where it doesn't need to be perfect, just better than what you could do yourself.
Here's an example of something last week that saved me some annoying drudge work in coding:
https://gitlab.com/-/snippets/2567734
And here's an example where it saved me having to skim through the massive documentation of a very "flexible" library to figure out how to do something:
https://gitlab.com/-/snippets/2549955
In the second category: I'm also learning two languages; I can paste a sentence into GPT-4 and ask it, "Can you explain the grammar to me?" Sure, there's a chance it might be wrong about something; but it's less wrong than the random guesses I'd be making by myself. As I gain experience, I'll eventually correct all the mistakes -- both the ones I got from making my own guesses, and the ones I got from GPT-4; and the help I've gotten from GPT-4 makes the mistakes worth it.
Re: GPT-4 is getting worse over time, not better
#146The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.
Re: GPT-4 is getting worse over time, not better
#147Title: "GPT-4 is getting worse over time, not better" Paper title: "How Is ChatGPT’s Behavior Changing over Time?" Paper Abstract: "GPT-3.5 and GPT-4 are the two most widely used large language model (LLM) services." When are people gonna realize that GPT-4/3.5 != ChatGPT As far as I can tell, the paper doesn't explain the methodology either, so hard to know if they're actually using "raw" GPT-4 or GPT-4 via ChatGPT.…
Re: GPT-4 is getting worse over time, not better
#148Earlier quoted context omitted.
So the original paper from Stanford and Berkley is also linked to this AI influencer? I am really amazed by this kind of dismissal. Its totally irrelevant who posted the info and how framed it is, as long as you have access to the source.
The paper does things like ask GPT-4 to write code and then check if that code compiles. Since March, they've fine-tuned GPT-4 to add back-ticks around code, which improves human-readable formatting but stops the code compiling. This is interpreted as "degraded performance" in the paper even though it's improved performance from a human perspective.
> it's improved performance from a human perspective.
Ignoring explicit requirements is the kind of thing that makes modern day search engines a pain to use.
Re: GPT-4 is getting worse over time, not better
#149Earlier quoted context omitted.
No, beta has nothing to do with if it's paid or not. Lots of products are released, and you pay for them, with early access/ beta. No one lied to you. No one tricked you. It is in beta. You chose to pay for it.
I never said anyone lied to or tricked anyone. I'm saying that "beta" is a term referring to pre-release testing. If people are paying for it, it's been released and is therefore no longer a beta. Calling it a "beta" at that point is just pure PR.
no
You can define "beta" whatever weird way you want but don't complain when it's not how literally everyone else uses it and don't complain when you're paying for something that uses the term the way everyone else does
Re: GPT-4 is getting worse over time, not better
#150There is a Chatbot AI product called character.ai that has suffered a marked decline in quality since its launch as they battle their users to maintain the AI’s safety protocols (similar to chatGPT “jailbreaks”). I wonder if something similar could be happening here.
These fighting against people using their product in “unauthorized” ways by the ai companies doesn’t make any sense to me. Who cares if character.ai users do some weird stuff with it, or replika creates romantic relationships, or people make off color jokes in gpt. There seems to be a lot of engineering effort driven by some product managers to have the ai do very specific things which makes the product much worse in…
But second, the reasons are:
(1) For AI company, someone publishing: "I asked the model a question about crime, and it talked shit about black people! Look! [damning quote that you can also get model to say/do]." Stability took the "let people do what they will" tack and now Forbes and every other major media mouthpiece slams them at every opportunity about how they are ethically-challenged.
(2) For Replika, someone chatting with their online girlfriend: "I love you more than my wife and children." Then someone hacking Replika exposing these conversations, and now Replika is in hot water because all these divorces. Replace example with 100 other similarly awful situations like talking about mental health problems, crimes, petty squabbles with their coworkers, or political problems.