What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.
Check out prompt 2, "Write a limerick about a dog". The models undeniably get better at writing limericks, but I think the answers are progressively less interesting. GPT-1 and GPT-2 are the most interesting to read, despite not following the prompt (not being limericks.) They get boring as soon as it can write limericks, with GPT-4 being more boring than text-davinci-001 and GPT-5 being more boring still.
OpenAI Progress
271–280 of 372 posts
Re: OpenAI Progress
#272My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…
To me 4 to 5 got much faster, but also worse. It is much more often ignoring explicit instructions like: "generate 10 song-titles with varying length" and it generates 10 song titles that are nearly identical length. This worked somewhat well with version 3 already..
Re: OpenAI Progress
#273Earlier quoted context omitted.
> I'm not brave enough to draw a public conclusion about what this could mean. I'm brave enough to be honest: it means nothing. LLMs execute a very sophisticated algorithm that pattern matches against a vast amount of data drawn from human utterances. LLMs have no mental states, minds, thoughts, feelings, concerns, desires, goals, etc. If the training data were instead drawn from a billion monkeys banging on typewrit…
LLMs are not people, but they are still minds, and to deny even that seems willfully luddite. While they are generating tokens they have a state, and that state is recursively fed back through the network, and what is being fed back operates not just at the level of snippets of text but also of semantic concepts. So while it occurs in brief flashes I would argue they have mental state and they have thoughts. If we bu…
Re: OpenAI Progress
#274Earlier quoted context omitted.
I did initial tests so that I don't have to do it anymore.
Everyone else has done tests that indicate that you do.
Comment sections are never good at being accountable for how vibes-driven they are when selecting which anecdotes to prefer.
Re: OpenAI Progress
#275Earlier quoted context omitted.
It got it right with thinking which was the challenge I posed. https://chatgpt.com/share/68a0b897-f8dc-800b-8799-9be2a8ad54...
The point you're missing is it's not always right. Cherry-picking examples doesn't really bolster your point. Obviously it works for you (or at least you think it does), but I can confidently say it's fucking god-awful for me.
That was never their argument. And it's not cherry picking to make an argument that there's a definable of examples where it returns broadly consistent and accurate information that they invite anyone to test.
They're making a legitimate point and you're strawmanning it and randomly pointing to your own personal anecdotes, and I don't think you're paying attention to the qualifications they're making about what it's useful for.
Re: OpenAI Progress
#276Earlier quoted context omitted.
The poster didn't use "thinking" model. That was my original challenge!! Why don't you try the original prompt using thinking model and see if I'm cherry picking?
Every time I use ChatGPT I become incredibly frustrated with how fucking awful it is. I've used it more than enough, time and time again (just try the new model, bro!), to know that I fucking hate it. If it works for you, cool. I think it's dogshit.
I'm sorry but this is a lazy and unresponsive string of comments that's degrading the discussion.
Re: OpenAI Progress
#277Earlier quoted context omitted.
[flagged]
Sorry but no. It's still early fooled and confused. Here's a trivial example: https://chatgpt.com/share/688b00ea-9824-8007-b8d1-ca41d59c18...
Re: OpenAI Progress
#278Re: OpenAI Progress
#279Cynical TLDR; We have plateaued and it has become obvious that fancy autocomplete is not and can never be close to reasoning, regardless of how many hacks and tweaks we are making.
Re: OpenAI Progress
#280Earlier quoted context omitted.
> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.
They outperform asking humans, unless you are asking an expert. On average