Live data from Hacker News

OpenAI Progress

progress.openai.com

271–280 of 372 posts

Re: OpenAI Progress

#271
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

Check out prompt 2, "Write a limerick about a dog". The models undeniably get better at writing limericks, but I think the answers are progressively less interesting. GPT-1 and GPT-2 are the most interesting to read, despite not following the prompt (not being limericks.) They get boring as soon as it can write limericks, with GPT-4 being more boring than text-davinci-001 and GPT-5 being more boring still.

I don't know if that is bad. The most intelligent person on a party is usually also the most boring one.

Re: OpenAI Progress

#272
post #251

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

To me 4 to 5 got much faster, but also worse. It is much more often ignoring explicit instructions like: "generate 10 song-titles with varying length" and it generates 10 song titles that are nearly identical length. This worked somewhat well with version 3 already..

Shows that they can't solve the fundamental problems as the technology, while amusing and with some utility, is also a dead end if we are going after cognition.

Re: OpenAI Progress

#273
post #250
post #209

Earlier quoted context omitted.

> I'm not brave enough to draw a public conclusion about what this could mean. I'm brave enough to be honest: it means nothing. LLMs execute a very sophisticated algorithm that pattern matches against a vast amount of data drawn from human utterances. LLMs have no mental states, minds, thoughts, feelings, concerns, desires, goals, etc. If the training data were instead drawn from a billion monkeys banging on typewrit…

LLMs are not people, but they are still minds, and to deny even that seems willfully luddite. While they are generating tokens they have a state, and that state is recursively fed back through the network, and what is being fed back operates not just at the level of snippets of text but also of semantic concepts. So while it occurs in brief flashes I would argue they have mental state and they have thoughts. If we bu…

There is no hidden state in a recurrent nets sense. Each new token just has all the previous tokens and that’s it.

Re: OpenAI Progress

#274
post #185

Earlier quoted context omitted.

I did initial tests so that I don't have to do it anymore.

Everyone else has done tests that indicate that you do.

And this is why you can't use personal anecdotes to settle questions of software performance.

Comment sections are never good at being accountable for how vibes-driven they are when selecting which anecdotes to prefer.

Re: OpenAI Progress

#275

Earlier quoted context omitted.

It got it right with thinking which was the challenge I posed. https://chatgpt.com/share/68a0b897-f8dc-800b-8799-9be2a8ad54...

The point you're missing is it's not always right. Cherry-picking examples doesn't really bolster your point. Obviously it works for you (or at least you think it does), but I can confidently say it's fucking god-awful for me.

>The point you're missing is it's not always right.

That was never their argument. And it's not cherry picking to make an argument that there's a definable of examples where it returns broadly consistent and accurate information that they invite anyone to test.

They're making a legitimate point and you're strawmanning it and randomly pointing to your own personal anecdotes, and I don't think you're paying attention to the qualifications they're making about what it's useful for.

Re: OpenAI Progress

#276

Earlier quoted context omitted.

The poster didn't use "thinking" model. That was my original challenge!! Why don't you try the original prompt using thinking model and see if I'm cherry picking?

Every time I use ChatGPT I become incredibly frustrated with how fucking awful it is. I've used it more than enough, time and time again (just try the new model, bro!), to know that I fucking hate it. If it works for you, cool. I think it's dogshit.

They just spent like six comments imploring you to understand that they were making a specific point: generally reliable on non-niche topics using thinking mode. And that nuance bounced off of you every single time as you keep repeating it's not perfect, dismiss those qualifications as cherry picking and repeat personal anecdotes.

I'm sorry but this is a lazy and unresponsive string of comments that's degrading the discussion.

Re: OpenAI Progress

#277
post #74
post #70

Earlier quoted context omitted.

[flagged]

Sorry but no. It's still early fooled and confused. Here's a trivial example: https://chatgpt.com/share/688b00ea-9824-8007-b8d1-ca41d59c18...

That worked great though? The question is confusing and unclear and it found an interpretation that made some sense and ran with it.

Re: OpenAI Progress

#278
it would be interesting to get GPT-OSS 120B and 20B responses to all of these questions to see how they compare.

Re: OpenAI Progress

#279
post #270

Cynical TLDR; We have plateaued and it has become obvious that fancy autocomplete is not and can never be close to reasoning, regardless of how many hacks and tweaks we are making.

Do you think the OP supports this claim? Don't you think the answers shown from GPT-5 are better than those from 4?

Re: OpenAI Progress

#280
post #134

Earlier quoted context omitted.

> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.

They outperform asking humans, unless you are asking an expert. On average

When I have a question, I don't usually "ask" that question and expect an answer. I figure out the answer. I certainly don't ask the question to a random human.
Post reply on HN