Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

241–250 of 315 posts

Re: GPT-4 is getting worse over time, not better

#241
post #34

Earlier quoted context omitted.

Are you sure it wasn't just that the novelty wore off after a few hours of usage? I never really got into LLMs, but I must say at first it seemed like pretty cool stuff. OpenAI have repeatedly stated the model hasn't changed so how could this happen otherwise?

These models have been explicitly nerfed since their first release due to copyright considerations. I've mentioned in two previous cases both for [1] code generation and [2] book summarizing. From my point of view, it is sad that these sort of socio-political constructs (copyright) are hindering innovation. The funny thing is that in say, 10 years, the "pirate" version of LLMs will be way more powerful and useful tha…

> I sincerely wish that some country would apply to information copyright a similar approach to what India does for medicine patents.

Care to take this opportunity to explain India's approach to medicine patents, and why you think it's good?

Re: GPT-4 is getting worse over time, not better

#242

Earlier quoted context omitted.

There is degraded performance because GPT4 refuses to carry out certain tasks. To figure it out though, you must need to be able to switch between GPT4-0316 and GPT4-0614. The task it is reluctant to do include: - legal advice - psychological guidance - complex programming tasks. IMO OpenAI is just backtracking on what it released to resegment their product into multiple offerings.

I think it's probably a good idea that GPT4 avoids legal or psychological tasks. Those are areas where giving incorrect output can have catastrophic consequences, and I can see why GPT4's developers want to avoid potential liability.

I wasn't asking for advice. I was in fact exploring ways to measure employee performance (using GPT as a search engine) and criticized one of the assumptions in one of theoretical framework with my own experience. GPT retorted "burnout", I retorted "harassment grounded on material evidence", and it spitted a generic block about how it couldn't provide legal or psychological guidance. I wasn't asking for advice, it was just a point in a wider conversation. Switching to GPT-4-0314 it proceeded with a reply that matched the discussion main's topic (metrics in HR) switching it back to GPT-4-0613 it outputted the exact same generic block.

Re: GPT-4 is getting worse over time, not better

#243
post #4

>Having the behavior of an LLM change over time is not acceptable. By now this is actually funny to read. Never rely on another companies product to make your own product, without accepting things can change overnight and shut you down As Llama2 is self hosted, you can choose which iteration to host. Much better developer experience Edit: to be clear OpenAI is unprofitable, so is Reddit, so was Stadia. Building on to…

The underlying models have not changed. You can specify the specific model variant you want via the API. (The underlying model used by the ChatGPT application may be updated, but that's not what the linked paper discusses.)

Re: GPT-4 is getting worse over time, not better

#244
post #72

Earlier quoted context omitted.

This comment has an interesting take on it, haven't read the paper to verify the take: https://news.ycombinator.com/item?id=36781968 EDIT: FWIW I haven't noticed any such regression. I don't generally use it to find prime numbers, but I do use it for coding, and have been really impressed with what it's able to do. 8 This paper is being misinterpreted. The degradations reported are somewhat peculiar to the authors' t…

> Code generation: the change they report is that the newer GPT-4 adds non-code text to its output. They don't evaluate the correctness of the code. They merely check if the code is directly executable. So the newer model's attempt to be more helpful counted against it. In the prompt they specifically request only the Python code, no other output. An “attempt to be helpful” that directly contradicts the user’s reques…

That's false. If it outputs formatted code, it's easier to read. I don't see the backtics, I see formatted code when using the chat interface.

Re: GPT-4 is getting worse over time, not better

#245

Earlier quoted context omitted.

It is still ignoring an explicit requirement, which is almost always bad. The user should be able to override what the creator/application thinks is 'strictly better'. Exceptions probably exist, but this isn't one of them.

Like I said, I read their requirement differently—the phrase they use is "the code only". I don't think that including backticks is a violation of this requirement. It's still readily parseable and serves as metadata for interpreting the code. In the context of ChatGPT, which typically will provide a full explanation for the code snippet, I think this is a reasonable interpretation of the instruction.

I understand your point better now. I'm still not sure how I feel about it, because code + metadata is still not "the code only", but it's not a totally unreasonable interpretation of the phrase.

Re: GPT-4 is getting worse over time, not better

#246

Earlier quoted context omitted.

>Is Tom Clancy unaligned? Yes, humans are unaligned. This is why alignment is hard: we're trying to produce machines with human-level intelligence but superhuman levels of morality.

Humans are definitely aligned, and for the same reasons as a LLM. Socialization, being allowed to work, being allowed to speak. edit: It's a social faux pas to say "died" about a person acquainted to the listener in most situations, you have to say "passed away."

> It's a social faux pas to say "died" about a person acquainted to the listener in most situations

That’s overly simplistic and an Americanism. The resurgence of the “passed away” euphemism is a recent (about 40 years) phenomenon in American English which seems to have been started out of the funeral industry as prior to that “died” was nearly universal for both news stories and obituaries.

“Died” is not a social faux pas. It’s the good default option as well. Medical professionals are often trained to avoid any euphemisms for death. I’ve never observed any problems professionally (as is standard) or personally using died even with folks that are religious.

https://english.stackexchange.com/questions/207087/origin-of...

Re: GPT-4 is getting worse over time, not better

#247
It's plausible that GPT-4 getting worse is just a cash grab by OpenAI. Release a powerful yet expensive to run model, push its transient virality to get people to sign up for monthly memberships, and then replace it with a cheaper to run/worse model to rake it in. Their investment deal with MSFT strongly incentivizes them toward profitability sooner (MSFT gets 75% of their profits until the $10B is "paid back"). So if they want to become independent from MSFT this might be their best bet.

Re: GPT-4 is getting worse over time, not better

#249

Earlier quoted context omitted.

> terrifyingly unaligned Honestly, if people think that a statistical language model is "terrifying" because it can verbalise the concept of a mass killing, they need to give their heads a wobble. My text editor can be used to write "set off a nuclear weapon in a city, lol". Is Notepad++.exe terrifying? What about the Sum of All Fears ? I could get some pointers from that. Is Tom Clancy unaligned? Am I terrifying bec…

>Is Tom Clancy unaligned? Yes, humans are unaligned. This is why alignment is hard: we're trying to produce machines with human-level intelligence but superhuman levels of morality.

There’s no such thing as “superhuman” morality, morality is just social mores and norms accepted by the people in a society at some given time. It does not advance or decline, but it changes.

What you’re talking about is a very small subset of the population forcing their beliefs on everyone else by encoding them in AI. Maybe that’s what we should do but we should be honest about it.

Re: GPT-4 is getting worse over time, not better

#250
post #79
post #59

Earlier quoted context omitted.

[flagged]

Then why comment? I'm not trying to be an ass, but this is literally the only comment you've made in this thread. The linked article is a Twitter thread. If you don't have Twitter, fine, move on to the next thing you can actually comment on. Person A offers a reason why we may want to be slightly more skeptical than normal about this information. Person B suggests we can pretty easily look past that. Person C (you) i…

I do think it'd be nice if people mostly posted links that were accessible to the general public. At least paywalled articles usually have a link in the comments where they can be read. Many times tweets could even be pasted as text into a comment to make them accessible.
Post reply on HN