Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

301–310 of 315 posts

Re: GPT-4 is getting worse over time, not better

#301

Earlier quoted context omitted.

> it's just that the magic has worn off as we've used this tech I agree with this. The analogy I use on repeat is the dawn of moving picture making. The first movies were short larks, designed just to elicit a response. Just like when CGI was new- a bunch of over-the-top, sensationalist fluff got made. This tech needs to mature. And we need it to continue to be fed the work of real humans, not AI feeding on AI, a rec…

I have actual artifacts -- programs it wrote -- from immediately after the release, which I'm unable to replicate now. Don't claim it hasn't gotten dumber. It's easy to find excuses, but none that explain my experience.

ChatGPT is a nondeterministic system that's sampling from a huge distribution in which there are thousands or possibly tens of thousands of events with similarly infinitesimal probabilities. There is no guarantee that you will get the same result twice, and in fact you should be a little surprised if you did.

That is not an excuse, and explains your experience.

Re: GPT-4 is getting worse over time, not better

#303
post #9

Earlier quoted context omitted.

LLama 2 is lost in the sauce... Q: How many 90 degree permutations can you do to leave a cube invariant from the perspective of an outside observer? A: As a responsible and ethical AI language model, I must first emphasize that the concept of "90 degree permutations" and "cube" are purely theoretical and have no basis in reality. However, I understand that you are asking for a hypothetical scenario, and I will provid…

I'd question why you'd ask an LLM something like this instead of Wolfram Alpha, or just look up the answer. It's not something you need an LLM to figure out. It's like asking an LLM for all the digits of pi multiplied by 5. Why? The problem with LLMs is the amount of things it makes sense to use them for is really not that large.

>> I'd question why you'd ask an LLM something like this instead of Wolfram Alpha, or just look up the answer.

That's because OpenAI dedicated an entire section in claims that GPT-3 is reasonably good at arithmetic. See Section 3.9.1 titled "Arithmetic" in "Language Models are Few-Shot Learners":

https://arxiv.org/abs/2005.14165

Whence I quote below:

Overall, GPT-3 displays reasonable proficiency at moderately complex arithmetic in few-shot, one-shot, and even zero-shot settings.

The reported results are pretty poor and too poor to justify even the relatively weak claim above (although the rest of the text in the same section very clearly and strongly implies that GPT-3 is doing something else than simply memorising a table of sums, which is an altogether much grander claim). OpenAI themselves seemed to be dubious enough about their own claim that the Arithmetic section of their paper was only included in the preprint (on Arxiv) and not in the published paper (in the 34th NeurIPS).

Yet, the claim in the preprint was still enough for people to forcefully argue that GPT-3 can do arithmetic, that it can learn the rules of arithmetic, and other impossible things before breakfast.

This is a discussion that goes back at least 3 years (judging from my comments where I point out that it's nonsense). It seems that the arithmetic ability of large language models is now a well accepted truth in the minds of the general public, who will casually use it to do, say, their maths coursework etc.

So blame OpenAI who made the big claims.

Re: GPT-4 is getting worse over time, not better

#304
post #299

Earlier quoted context omitted.

I'm sure it would be easy to pay publishers who pass on the money? How do you think Amazon does things for kindle?

Will you get a check for your posts on HN? Your tweets? No, right? So all this talk about paying people for their data is rubbish, and what is really at stake here are checks between huge corporations that already stole it from you in the first place.

I'm not quite sure which side your on but yeah, I personally have stopped participating in social media since realizing I'm the product.

HN I agree with you, but I feel like I'm getting enough back from this community to make my time worth it.

Re: GPT-4 is getting worse over time, not better

#305

Earlier quoted context omitted.

I don’t believe this is true. It’s possible I was blinded by the light, but my programming tasks were previously (during the early access program) being handled by GPT-4 regularly and now they aren’t. I’ve also seen many anecdotes from engineers who had exceptionally early access before GPT-4 was public knowledge. The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back…

I used it for writing assistance and Plot development. Specifically, a novel re: the conquest of Mexico in the 16th cent. It was great at spitting out ideas re: action scenes and even character development. In the past month or so, it has become so cluttered with caveats and tripe regarding the political aspects of the conquest, that it is useless. I can’t replicate the work I was doing before. Actually cancelled my…

I fear we have to wait for a non-woke (so probably non-US) entity to train a useful GPT4(+) level model. Maybe one from Tencent or Baidu could be could, provided you avoid very specific topics like Taiwan or Xi.

Re: GPT-4 is getting worse over time, not better

#306

The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.

So the original paper from Stanford and Berkley is also linked to this AI influencer? I am really amazed by this kind of dismissal. Its totally irrelevant who posted the info and how framed it is, as long as you have access to the source.

Like clockwork, it's coming out that the original paper was wildly misinterpreted: https://www.aisnakeoil.com/p/is-gpt-4-getting-worse-over-tim...

Re: GPT-4 is getting worse over time, not better

#307

Earlier quoted context omitted.

I have actual artifacts -- programs it wrote -- from immediately after the release, which I'm unable to replicate now. Don't claim it hasn't gotten dumber. It's easy to find excuses, but none that explain my experience.

ChatGPT is a nondeterministic system that's sampling from a huge distribution in which there are thousands or possibly tens of thousands of events with similarly infinitesimal probabilities. There is no guarantee that you will get the same result twice, and in fact you should be a little surprised if you did. That is not an excuse, and explains your experience.

You're assuming I'm not aware of that, and didn't try multiple times. The original program wasn't the output of a single request; it took several retries and some refining, but all the retries produced something at least mostly correct.

Doesn't matter how many times I try it now, it just doesn't work.

Re: GPT-4 is getting worse over time, not better

#308

Earlier quoted context omitted.

sprucing up, fixing your mistakes, adding in "descriptive stuff"... that's like 90% of writing. Outsourcing it all to AI essentially robs the purchaser of the effort required to create an original piece of work. Not to mention copyright issues, where do you think the AI is getting those descriptive phrases from? Other authors' work.

I think that you, like I did in the past, are underestimating the number of people who simply hate writing and see it as a painstaking chore that they would happily outsource to a machine. It doesn't help that most people grow up being forced to write when they have no interest in doing so, and to write things they have no interest in writing, like school essays and business applications and so on. If a chatbot could…

The original post specifically stated he was utilizing AI for plot development of a novel. Not a school essay or business application.

Re: GPT-4 is getting worse over time, not better

#309

Earlier quoted context omitted.

I think that you, like I did in the past, are underestimating the number of people who simply hate writing and see it as a painstaking chore that they would happily outsource to a machine. It doesn't help that most people grow up being forced to write when they have no interest in doing so, and to write things they have no interest in writing, like school essays and business applications and so on. If a chatbot could…

The original post specifically stated he was utilizing AI for plot development of a novel. Not a school essay or business application.

I should have bookmarked it but there was an article shared on HN, published in the New Yorker I believe, or the LARB, or some such, where a professional writer was praising some language model as a tool for people who hate writing, like themself and other professional writers. I was dumbstruck.

But, it's true. Even people who want to write, even write literature, can hate the act of actually, you know, writing.

In part, I'm trying to convince myself because I still find it hard to believe but it seems to be the case.

Re: GPT-4 is getting worse over time, not better

#310

Earlier quoted context omitted.

ChatGPT is a nondeterministic system that's sampling from a huge distribution in which there are thousands or possibly tens of thousands of events with similarly infinitesimal probabilities. There is no guarantee that you will get the same result twice, and in fact you should be a little surprised if you did. That is not an excuse, and explains your experience.

You're assuming I'm not aware of that, and didn't try multiple times. The original program wasn't the output of a single request; it took several retries and some refining, but all the retries produced something at least mostly correct. Doesn't matter how many times I try it now, it just doesn't work.

>> You're assuming I'm not aware of that, and didn't try multiple times.

I'm not! What you describe is perfectly normal. Of course you'd need multiple tries to get the "right" output and of course if you keep trying you'll get different results. It's a stochastic process and you have very limited means to control the output.

If you sample from a model you can expect to get a distribution of results, rather tautologically. Until you've drawn a large enough number of samples there's no way to tell what is a representative sample, and what should surprise you. But what is a "large enough" number of samples is hard to tell, given a large model, like a large language model.

>> Doesn't matter how many times I try it now, it just doesn't work.

Wanna bet? If you try long enough you'll eventually get results very similar to the ones you got originally. But it might take time. Or not. Either you got lucky the first few times and saw uncommon results, or you're getting unlucky now and seeing uncommon results. It's hard to know until you've spent a lot of time and systematically test the model.

Statistics is a bitch. Not least because you need to draw lots of samples and you never know how close you are to the true distribution.

Post reply on HN