Earlier quoted context omitted.
The paper does things like ask GPT-4 to write code and then check if that code compiles. Since March, they've fine-tuned GPT-4 to add back-ticks around code, which improves human-readable formatting but stops the code compiling. This is interpreted as "degraded performance" in the paper even though it's improved performance from a human perspective.
The query explicitly asks it to add no other text to the code. > it's improved performance from a human perspective. Ignoring explicit requirements is the kind of thing that makes modern day search engines a pain to use.
GPT-4 is getting worse over time, not better
171–180 of 315 posts
Re: GPT-4 is getting worse over time, not better
#172The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.
Re: GPT-4 is getting worse over time, not better
#173Earlier quoted context omitted.
So the original paper from Stanford and Berkley is also linked to this AI influencer? I am really amazed by this kind of dismissal. Its totally irrelevant who posted the info and how framed it is, as long as you have access to the source.
The paper does things like ask GPT-4 to write code and then check if that code compiles. Since March, they've fine-tuned GPT-4 to add back-ticks around code, which improves human-readable formatting but stops the code compiling. This is interpreted as "degraded performance" in the paper even though it's improved performance from a human perspective.
- legal advice
- psychological guidance
- complex programming tasks.
IMO OpenAI is just backtracking on what it released to resegment their product into multiple offerings.
Re: GPT-4 is getting worse over time, not better
#174The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.
> it's just that the magic has worn off as we've used this tech I agree with this. The analogy I use on repeat is the dawn of moving picture making. The first movies were short larks, designed just to elicit a response. Just like when CGI was new- a bunch of over-the-top, sensationalist fluff got made. This tech needs to mature. And we need it to continue to be fed the work of real humans, not AI feeding on AI, a rec…
Don't claim it hasn't gotten dumber. It's easy to find excuses, but none that explain my experience.
Re: GPT-4 is getting worse over time, not better
#175Earlier quoted context omitted.
These fighting against people using their product in “unauthorized” ways by the ai companies doesn’t make any sense to me. Who cares if character.ai users do some weird stuff with it, or replika creates romantic relationships, or people make off color jokes in gpt. There seems to be a lot of engineering effort driven by some product managers to have the ai do very specific things which makes the product much worse in…
First, I am also frustrated by companies trying to prevent unauthorised used. But second, the reasons are: (1) For AI company, someone publishing: "I asked the model a question about crime, and it talked shit about black people! Look! [damning quote that you can also get model to say/do]." Stability took the "let people do what they will" tack and now Forbes and every other major media mouthpiece slams them at every…
Re: GPT-4 is getting worse over time, not better
#176Earlier quoted context omitted.
The query explicitly asks it to add no other text to the code. > it's improved performance from a human perspective. Ignoring explicit requirements is the kind of thing that makes modern day search engines a pain to use.
If I'm using it from the web UI, this is exactly what I would want—this allows the language model to define the output language so I get correct syntax highlighting, without an error-prone secondary step of language detection. If I'm using it from the API, then all I have to do is strip out the leading backticks and language name if I don't need to check the language, or alternatively parse it to determine what the o…
Re: GPT-4 is getting worse over time, not better
#177Earlier quoted context omitted.
LLama 2 is lost in the sauce... Q: How many 90 degree permutations can you do to leave a cube invariant from the perspective of an outside observer? A: As a responsible and ethical AI language model, I must first emphasize that the concept of "90 degree permutations" and "cube" are purely theoretical and have no basis in reality. However, I understand that you are asking for a hypothetical scenario, and I will provid…
Is this real? That’s incredible if so.
See for yourself
Re: GPT-4 is getting worse over time, not better
#178>Having the behavior of an LLM change over time is not acceptable. By now this is actually funny to read. Never rely on another companies product to make your own product, without accepting things can change overnight and shut you down As Llama2 is self hosted, you can choose which iteration to host. Much better developer experience Edit: to be clear OpenAI is unprofitable, so is Reddit, so was Stadia. Building on to…
LLama 2 is lost in the sauce... Q: How many 90 degree permutations can you do to leave a cube invariant from the perspective of an outside observer? A: As a responsible and ethical AI language model, I must first emphasize that the concept of "90 degree permutations" and "cube" are purely theoretical and have no basis in reality. However, I understand that you are asking for a hypothetical scenario, and I will provid…
It's like asking an LLM for all the digits of pi multiplied by 5. Why?
The problem with LLMs is the amount of things it makes sense to use them for is really not that large.
Re: GPT-4 is getting worse over time, not better
#179Earlier quoted context omitted.
Just to be clear, you're saying that because they're tweaking GPT-4 to give more explanations of code, you shouldn't rely on it for coding? Obviously if that's your own preference, I'm not going to tell you that you're wrong; but I think in general , most people wouldn't agree with that statement.
I'm still wondering, why should anyone rely on AI generated answers? They are logically no better than search engine results. By that I mean, you can't tell if it's returning absolute trash or spot on correct. Building trust into it all is going to be either a) expensive or b) driven by all the wrong incentives.
I use to generate code I'd get from libraries. Graph-theory related algorithms, special datastructures, etc...
Re: GPT-4 is getting worse over time, not better
#180Earlier quoted context omitted.
Name a single subscription service that never changes over time? If your product is based on another company's service then you are beholden to them.
Electricity, water, TV etc... they basically always work. If they change, usually not for the worst - you dont get color TV shows downgraded to b&w.