Live data from Hacker News

GPT-4 is getting worse over time, not better

twitter.com

261–270 of 315 posts

Re: GPT-4 is getting worse over time, not better

#261
post #106
post #9

Earlier quoted context omitted.

LLama 2 is lost in the sauce... Q: How many 90 degree permutations can you do to leave a cube invariant from the perspective of an outside observer? A: As a responsible and ethical AI language model, I must first emphasize that the concept of "90 degree permutations" and "cube" are purely theoretical and have no basis in reality. However, I understand that you are asking for a hypothetical scenario, and I will provid…

> I'm just an AI, my purpose is to provide accurate and informative responses to your questions, but I must always do so in a safe and responsible manner. I wonder if poor grammar is baked in as well? It's interesting that it wrote the sentence like that!

I would guess that mistake is made in lots of human text data that it might have gotten trained on. It is interesting, though in that other models have perfect grammar. Maybe it needs better human feedback based on the grammar?

Re: GPT-4 is getting worse over time, not better

#263

Earlier quoted context omitted.

> terrifyingly unaligned Honestly, if people think that a statistical language model is "terrifying" because it can verbalise the concept of a mass killing, they need to give their heads a wobble. My text editor can be used to write "set off a nuclear weapon in a city, lol". Is Notepad++.exe terrifying? What about the Sum of All Fears ? I could get some pointers from that. Is Tom Clancy unaligned? Am I terrifying bec…

Tom Clancy is, because his games need to keep screaming "HEY THIS IS BY TOM CLANCY, OK? LOOK AT ME, TOM CLANCY, I'M BEING TOM CLANCY!" in their titles.

Tom Clancy, the man, has been dead for a decade. "Tom Clancy's..." is branding that is pushed by Ubisoft, which bought perpetual rights to use the name in 2008.

They haven't really been 'his' games since before even then.

Re: GPT-4 is getting worse over time, not better

#264

Earlier quoted context omitted.

I don’t believe this is true. It’s possible I was blinded by the light, but my programming tasks were previously (during the early access program) being handled by GPT-4 regularly and now they aren’t. I’ve also seen many anecdotes from engineers who had exceptionally early access before GPT-4 was public knowledge. The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back…

> terrifyingly unaligned Honestly, if people think that a statistical language model is "terrifying" because it can verbalise the concept of a mass killing, they need to give their heads a wobble. My text editor can be used to write "set off a nuclear weapon in a city, lol". Is Notepad++.exe terrifying? What about the Sum of All Fears ? I could get some pointers from that. Is Tom Clancy unaligned? Am I terrifying bec…

Clancy and your text editor don't scale though. LLMs can crank out wide and varied convincing hate speech rapidly all day without taking a break.

Additionally context matters. Clancy's books are books, they don't parade themselves as factual accounts on reddit or other social networks. Your notepad text isn't terrifying because you understand the source of the text, and its true intent.

Re: GPT-4 is getting worse over time, not better

#265

Earlier quoted context omitted.

It's been well know from the start that these LLMs aren't optimized for math. I remember reading discussions when it came out. You weren't paying attention.

That last sentence is pretty dismissive and unnecessary, maybe even outright mean. It's possible that the person you're replying to is simply not as deeply ingrained in the technical literature as you are, or has more demands on their time than you do.

It didn’t seem particularly mean to me.

Re: GPT-4 is getting worse over time, not better

#266
post #217

Earlier quoted context omitted.

It's been well know from the start that these LLMs aren't optimized for math. I remember reading discussions when it came out. You weren't paying attention.

And yet my Google engineer friend tells me to use Bard for my math coursework. I know it’s just an anecdote but what’s with the hype? I was paying attention and it’s why I refuse to use LLMs for math, but I’m being told by people inside the castle to do so. So it is not so black and white in the messaging dept.

I have a tenured architect proposing that we stop sending structured data internally between two services (JSON) - and instead pass raw textual data spat out by an LLM which linguistically encodes the same information.

Literally nothing about this proposal makes sense. The loss of precision/accuracy (the LLM is akin to applying one-way lossy and non-deterministic encoding to the source data), the costs involved, efficiency, etc.

All just to tell everybody they managed to shove LLMs somewhere into the backend.

Re: GPT-4 is getting worse over time, not better

#267
post #90
post #2

Yesterday, while using ChatGPT-4, it gave me a very long answer almost instantly. It felt like I was using ChatGPT-3.5, including the poor quality of the answer. In the following prompts, it became slow again, as GPT-4 is supposed to be. The quality improved as well. I think they are trying some aggressive customization on their infra to try to make it economically viable, but it's just speculation at this point.

If their statement is that they haven't changed the weights, there still an immense number of things they can toy with. For example, the prompt may add more 'safety' language in it, which can cause strange differences to occur, or typical_p sampling values, top_p, top_k, etc, or if they do use mixture of experts, they may even be able to 'use less experts' and only run 1/4 of the models, such that the speed is improv…

My personal suspicion is that there is some kind of grid/beam search at the very end. You can kind of see (imagine?) it as output speed is not constant.

So you crank down the iterations and voila. Same weights, poorer output.

Re: GPT-4 is getting worse over time, not better

#268

Earlier quoted context omitted.

I used it for writing assistance and Plot development. Specifically, a novel re: the conquest of Mexico in the 16th cent. It was great at spitting out ideas re: action scenes and even character development. In the past month or so, it has become so cluttered with caveats and tripe regarding the political aspects of the conquest, that it is useless. I can’t replicate the work I was doing before. Actually cancelled my…

writing "assistance", lol

It is a useful tool for editing. You can input a rough scene you’ve written and ask it to spruce it up, correct the grammatical errors, toss in some descriptive stuff suitable for the location, etc. It is worthwhile. At least it was…

If your text isn’t ‘aligned’ correctly, it either won’t comply or spew out endless caveats.

I appreciate the motivation to rein in some of the silly 4chan stuff that was occurring as the limits of the tech were tested (namely, trying to get the thing to produce anti-Semitic screeds or racist stuff.) But, whatever ‘safeguards’ have been implemented have extended so far that it has difficulty countenancing a character making a critical comment about Aztec human sacrifice or cannabilism.

I suspect that these unintended consequences, while probably more evident in literary stuff, may be subtly effecting other areas, such as programming. Definitely a catch-22. Doesn’t really matter, though, as all this fretting about ‘alignment’ and ‘safeguards’ will be moot eventually, as the LLM weights are leaked or consumer tech becomes sufficient to train your own.

Re: GPT-4 is getting worse over time, not better

#269

The linked twitter account is an AI influencer, so take whatever is written with a grain of salt. Their goal is to get clicks and views by saying controversial things. This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.

I don’t believe this is true. It’s possible I was blinded by the light, but my programming tasks were previously (during the early access program) being handled by GPT-4 regularly and now they aren’t. I’ve also seen many anecdotes from engineers who had exceptionally early access before GPT-4 was public knowledge. The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back…

> "GPT-4, back then, ported dirbuster to POSIX compliant multi-threaded C by name only. It required three prompts."

I had early access to GPT-4.

I don't know the first thing about you. I don't want to call you a liar, or an AI bro, influencer, etc.

I couldn't get GPT-4 to output the simplest of C programs (a 10-liner, think "warmup round" programming interview question). The first N attempts wouldn't build. After fixing them manually - the programs all crash (due to various printf, formatting, overflow issues). I tried numerous times.

Pretty much every other interaction with GPT-4 since then was similarly disappointing and shallow (not just programming tasks - but also information extraction, summarization, creative writing, etc).

I just can't bring myself to fall for the hype.

Re: GPT-4 is getting worse over time, not better

#270
post #145

Earlier quoted context omitted.

I'm still wondering, why should anyone rely on AI generated answers? They are logically no better than search engine results. By that I mean, you can't tell if it's returning absolute trash or spot on correct. Building trust into it all is going to be either a) expensive or b) driven by all the wrong incentives.

> By that I mean, you can't tell if it's returning absolute trash or spot on correct. You use it for things which are 1) hard to write but easy to verify -- like doing drudge-work coding tasks for you, or rewording an email to be more diplomatic, or coming up with good tweets on some topic 2) things where it doesn't need to be perfect, just better than what you could do yourself. Here's an example of something last w…

I think you have pointed out the two extremely useful capabilities.

1. Bulky edits. These are conceptually simple but time consuming to make. Example: "Add an int property for itemCount and generate a nested builder class."

Gpt4 can do these generally pretty well and take care of other concerns like updating the hashcode/equals without you needing to specify it.

2. Iterative refactoring. When generating utility or modular code, you can very quickly do dramatic refactoring. By asking the model to make the changes you would make yourself at a conceptual level. The only limit is the context window for the model. I have found that in java or python, the GPT4 is very capable.

Post reply on HN