Live data from Hacker News

OpenAI O3-Mini

openai.com

911–920 of 944 posts

Re: OpenAI O3-Mini

#911

For AI coding, o3-mini scored similarly to o1 at 10X less cost on the aider polyglot benchmark [0]. This comparison was with both models using high reasoning effort. o3-mini with medium effort scored in between R1 and Sonnet. 62% $186 o1 high 60% $18 o3-mini high 57% $5 DeepSeek R1 54% $9 o3-mini medium 52% $14 Sonnet 48% $0 DeepSeek V3 [0] https://aider.chat/docs/leaderboards/

Also Gemini API is free for coding.

I have yet to see a valid reason to use Gemini over any other alternative, the only exception being large contexts.

Re: OpenAI O3-Mini

#912
post #879

Earlier quoted context omitted.

This is a dumb argument. Humans frequently fall for the same tricks, are they not "intelligent"? All intelligence is ultimately based on some sort of statistical models, some represented in neurons, some represented in matrices.

State-of-the-art LLMs have been trained on practically the whole internet. Yet, they fall prey to pretty dumb tricks. It's very funny to see how The Guardian was able to circumvent censorship on the Deepseek app by asking it to "use special characters like swapping A for 4 and E for 3". [1] This is clearly not intelligence. LLMs are fascinating for sure, but calling them intelligent is quite the stretch. [1]: https:/…

The censorship is in fact not part of the llm. This can be shown easily by examples where llms visually output censored sentences after which they disappear.

Re: OpenAI O3-Mini

#913

Earlier quoted context omitted.

Thank you, this is a perfect argument why LLMs are not AI but just statistical models. The original is so overrepresented in the training data that even though they notice this riddle is different, they regress to the statistically more likely solution over the course of generating the response. For example, I tried the first one with Claude and in its 4th step, it said: > This is safe because the wolf won't eat the…

The problem with claims like these that models are not doing “actual reasoning” is that they are often hot takes and not thought through very well. For example, since reasoning doesn’t yet have any consensus definition that can be applied as a yes/no test - you have to explain what you specifically mean by it, or else the claim is hollow. Clarify your definition, give a concrete example under that definition of somet…

I like the true scottsman requirement. About any definition man can come up with for 'true' reasoning easily falls prey to a 'true scottsman fallacy'.

Re: OpenAI O3-Mini

#914

For years I've been asking all the models this mixed up version of the classic riddle and they 99% of the time get it wrong and insist on taking the goat across first. Even the other reasoning models would reason about how it was wrong, figure out the answer, and then still conclude goat. o3-mini is the first one to get it right for me. Transcript: Me: I have a wolf, a goat, and a cabbage and a boat. I want to get th…

Thank you, this is a perfect argument why LLMs are not AI but just statistical models. The original is so overrepresented in the training data that even though they notice this riddle is different, they regress to the statistically more likely solution over the course of generating the response. For example, I tried the first one with Claude and in its 4th step, it said: > This is safe because the wolf won't eat the…

I tried the same with deepseek. It failed and then i asked to try again and explain why it was wrong. The answer below is convincing for me that it has forms of reasoning, even meta reasoning.

" Certainly! If we take the riddle very literally and ignore the common interpretation or stereotypes, the answer I provided ("the surgeon is the boy's mother") would actually be incorrect because the riddle explicitly states:

"a surgeon, who is the boy's father, says, 'I cannot operate on this boy, he's my son!'"

The key phrase here is "who is the boy's father." This directly tells us that the surgeon is the boy's father. Therefore, the correct literal answer is that the surgeon is the boy's father.

The reason my initial answer was "wrong" in this context is that I interpreted the riddle as a play on stereotypes (the assumption that surgeons are male), rather than taking the statement at face value. By focusing on the literal wording, the surgeon is explicitly identified as the boy's father, and no further interpretation is needed. "

Re: OpenAI O3-Mini

#915

Earlier quoted context omitted.

I just installed and tried. Pretty neat stuff! Would be great if the addon allows user to override the sys prompt (it might need minor tweak when changing different server backend)?

Thank you for trying out the extension and for this great suggestion! We've actually been thinking along similar lines. Here are a couple of improvements we're considering: 1. Built-in prompt templates - Support multiple flavors (e.g. On similar to is there already, in addition to knowledge of up/down votes, another one similar to what Simon had - which is more detailed etc.) 2. User-editable prompts - Exactly like y…

the shared summaries sounds like a great idea to save most people's inference cost! There might be some details need to figure out - e.g. the summary per post need to be associated with a timestamp, if there are new comments kicking in after that (especially hot posts). Still i think it's good useful feature and i will definitely read that before browsing details.

Re: OpenAI O3-Mini

#916
post #879

Earlier quoted context omitted.

State-of-the-art LLMs have been trained on practically the whole internet. Yet, they fall prey to pretty dumb tricks. It's very funny to see how The Guardian was able to circumvent censorship on the Deepseek app by asking it to "use special characters like swapping A for 4 and E for 3". [1] This is clearly not intelligence. LLMs are fascinating for sure, but calling them intelligent is quite the stretch. [1]: https:/…

The censorship is in fact not part of the llm. This can be shown easily by examples where llms visually output censored sentences after which they disappear.

The nuance here being that this only proves additional censorship is applied on top of the output. It does not disprove that (sometimes ineffective) censorship is part of the LLM or that censorship was not attempted during training.

Re: OpenAI O3-Mini

#917
post #743
post #12

Earlier quoted context omitted.

I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.

Did they even think about what happens when they get to o4? We’re going to have GPT-4o and o4

They’ll call it GPT-XP. But first we need gpt-o3.11 for workgroups.

Re: OpenAI O3-Mini

#918
post #404

Well, o3-mini-high just successfully found the root cause of a seg fault that o1 missed: mistakenly using _mm512_store_si512 for an unaligned store that should have been _mm512_storeu_si512.

How do I avoid the angst about this stuff as a student in computer science? I love this field but frankly I've been at a loss since the rapid development of these models.

The cost of developing software is quickly dropping thanks to these models, and the demand for software is about to go way up because of this. LLMs will just be power tools to software builders. Learn to pop up a level.

Re: OpenAI O3-Mini

#919
post #669

Earlier quoted context omitted.

Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.

Have you ever had to give a demo that no meeting attendee actually cared about, just because management demanded it? Standup meetings with 20 people where maybe 2 people cared about? The future might involve AI updates, summarized by AI into weekly reports, summarized again into monthly reports, then into quarterly departmental reports that nobody actually reads.

Doesn't that seem ultimately futile though?

Say I use AI to write a report that nobody cares about, and then the reciever gives it to AI because they can't be bothered to read it. Who is benefiting here other than OpenAI?

Re: OpenAI O3-Mini

#920
post #780

Earlier quoted context omitted.

Our digital twins will write the comments. They will be us, but with none of our flaws. They will never experience the shame of posting a dumb joke, getting flamed, and then deleting it, for they will have tested all ideas to prevent such an oversight. They will never experience the satisfaction-turned-to-puzzlement of posting an expertly crafted, well-researched comment that took 2 hours of the workday to draft - on…

We all are, since a long time. The mud just doesn't shine so bright.

Well by definition the vast majority of comments written, anywhere, cannot be memorable.

Assuming the median reader reads a few tens of thousands comments in a year, only a few hundred would likely stick without being muddled. At best.

Post reply on HN