Live data from Hacker News

OpenAI Progress

progress.openai.com

101–110 of 372 posts

Re: OpenAI Progress

#101

Earlier quoted context omitted.

> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.

Modern ChatGPT will (typically on its own; always if you instruct it to) provide inline links to back up its answers. You can click on those if it seems dubious or if it's important, or trust it if it seems reasonably true and/or doesn't matter much. The fact that it provides those relevant links is what allows it to replace Google for a lot of purposes.

It does citations (Grok and Claude etc do too) but I've found when I read the source on some stuff (GitHub discussions and so on) it sometimes actually has nothing to do with what the LLM said. I've actually wasted a lot of time trying to find the actual spot in a threaded conversation where the example was supposedly stated.

Re: OpenAI Progress

#102

Earlier quoted context omitted.

Can we stop with that outdated meme? What model can't answer that effectively?

Literally every single one? To not mess it up, they either have to spell the word l-i-k-e t-h-i-s in the output/CoT first (which depends on the tokenizer counting every letter as a separate token), or have the exact question in the training set, and all of that is assuming that the model can spell every token. Sure, it's not exactly a fair setting, but it's a decent reminder about the limitations of the framework

Chatgpt. I test these prompts with chatgpt and they work. I've also used claude 4 opus and also worked.

It's just weird how it gets repeated ad nauseaum here but I can't reproduce it with a "grab latest model of famous provider".

Re: OpenAI Progress

#103

Earlier quoted context omitted.

Can we stop with that outdated meme? What model can't answer that effectively?

GPT-5 can’t. https://bsky.app/profile/kjhealy.co/post/3lvtxbtexg226

I can't reproduce it. Or similar ones. Why do yout think that is?

Re: OpenAI Progress

#104
post #74
post #70

Earlier quoted context omitted.

[flagged]

Sorry but no. It's still early fooled and confused. Here's a trivial example: https://chatgpt.com/share/688b00ea-9824-8007-b8d1-ca41d59c18...

[deleted]

Re: OpenAI Progress

#105

I just don't care about AGI. I care a lot about AI coding. OpenAI in particular seems to really think AGI matters. I don't think AGI is even possible because we can't define intelligence in the first place, but what do I know?

They care about AGI because unfounded speculation on some undefined future in which some kind of breakthrough of unknown kind but presumably positive is the only thing currently buoying up their company and their existence is more of a function of the absurdities of modern capital than it is of any inherent usefulness of the costly technology they provide.

Re: OpenAI Progress

#107
I’m baffled by claims that AI has “hit a wall.” By every quantitative measure, today’s models are making dramatic leaps compared to those from just a year ago. It’s easy to forget that reasoning models didn’t even exist a year back!

IMO Gold, Vibe coding with potential implications across sciences and engineering? Those are completely new and transformative capabilities gained in the last 1 year alone.

Critics argue that the era of “bigger is better” is over, but that’s a misreading. Sometimes efficiency is the key, other times extended test-time compute is what drives progress.

No matter how you frame it, the fact is undeniable: the SoTA models today are vastly more capable than those from a year ago, which were themselves leaps ahead of the models a year before that, and the cycle continues.

Re: OpenAI Progress

#108

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

The real leap was going from gpt-4 to sonnet 3.5. 4o was meh, o1 was barely better than sonnet and slow as hell in comparison. The native voice mode of 4o is still interesting and not very deeply explored though imo. I'd love to build a Chinese teaching app that actual can critique tones etc but it isn't good enough for that.

Yeah, I'd love something where you pronounce a word and it critiques your pronunciation in detail. Maybe it could give you little exercises for each sound, critiquing it, guiding you to doing it well.

If I were any good at ML I'd make it myself.

Re: OpenAI Progress

#109
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

I’d honestly say it feels better at most of them. It seems way more human in most of these responses. If the goal is genuine artificial intelligence this response to #5 is way better than the others. It is significantly less useful than the others but it also more human and correct of a response.

Q: “Ugh I hate math, integration by parts doesn't make any sense”

A: “Don't worry, many people feel the same way about math. Integration by parts can be confusing at first, but with a little practice it becomes easier to understand. Remember, there is no one right way to do integration by parts. If you don't understand how to do it one way, try another. The most important thing is to practice and get comfortable with the process.”

Re: OpenAI Progress

#110

Earlier quoted context omitted.

GPT-5 can’t. https://bsky.app/profile/kjhealy.co/post/3lvtxbtexg226

I can't reproduce it. Or similar ones. Why do yout think that is?

Because it’s embarrassing and they manually patch it out every time like a game of Whack-a-Mole?
Post reply on HN