Live data from Hacker News

OpenAI Progress

progress.openai.com

111–120 of 372 posts

Re: OpenAI Progress

#111

I’m baffled by claims that AI has “hit a wall.” By every quantitative measure, today’s models are making dramatic leaps compared to those from just a year ago. It’s easy to forget that reasoning models didn’t even exist a year back! IMO Gold, Vibe coding with potential implications across sciences and engineering? Those are completely new and transformative capabilities gained in the last 1 year alone. Critics argue…

thanks OpenAI, very cool!

Re: OpenAI Progress

#112

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

I must be crazy, because I clearly remember chatgpt 4 being downgraded before they released 4o, and I felt it was a worse model with a different label, I even choose the old chatgpt 4 when they would give me the option. I canceled my subscription around that time.

Not crazy. 4o was a hallucination machine. 4o had better “vibes” and was really good at synthesizing information in useful ways, but GPT-4 Turbo was a bigger model with better world knowledge.

Re: OpenAI Progress

#113

Earlier quoted context omitted.

> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.

Modern ChatGPT will (typically on its own; always if you instruct it to) provide inline links to back up its answers. You can click on those if it seems dubious or if it's important, or trust it if it seems reasonably true and/or doesn't matter much. The fact that it provides those relevant links is what allows it to replace Google for a lot of purposes.

In my experience, 80% of the links it provides are either 404, or go to a thread on a forum that is completely unrelated to the subject.

Im also someone who refuses to pay for it, so maybe the paid versions do better. who knows.

Re: OpenAI Progress

#114
post #46
post #6

GPT-5 IS an incredible breakthrough! They just don't understand! Quick, vibe-code a website with some examples, that'll show them!11!!1

5 is a breakthrough at reducing OpenAI's electric bills.

As someone who likes this planet, I'm grateful for that.

Re: OpenAI Progress

#115

I’m baffled by claims that AI has “hit a wall.” By every quantitative measure, today’s models are making dramatic leaps compared to those from just a year ago. It’s easy to forget that reasoning models didn’t even exist a year back! IMO Gold, Vibe coding with potential implications across sciences and engineering? Those are completely new and transformative capabilities gained in the last 1 year alone. Critics argue…

thanks OpenAI, very cool!

[dead]

Re: OpenAI Progress

#116

Earlier quoted context omitted.

I can't reproduce it. Or similar ones. Why do yout think that is?

Because it’s embarrassing and they manually patch it out every time like a game of Whack-a-Mole?

Except people use the same examples like blueberry and strawberry, which were used months ago, as if they're current.

These models can also call Counter from python's collections library or whatever other algorithm. Or are we claiming it should be a pure LLM as if that's what we use in the real world.

I don't get it, and I'm not one to hype up LLMs since they're absolutely faulty, but the fixation over this example screams of lack of use.

Re: OpenAI Progress

#117
post #38

Earlier quoted context omitted.

Maybe you should fact check your AI outputs more if you think it only hallucinates in niche topics

The accuracy is high enough that I don't have to fact check too often.

I totally get that you meant this in a nuanced way, but at face value it sort of reads like...

Joe Rogan has high enough accuracy that I don't have to fact check too often. Newsmax has high enough accuracy that I don't have to fact check too often, etc.

If you accept the output as accurate, why would fact checking even cross your mind?

Re: OpenAI Progress

#119
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

Honestly my quick take on the prompt was some sort of horror theme and GPT-1’s response fits nicely.
Post reply on HN