Live data from Hacker News

OpenAI O3-Mini

openai.com

891–900 of 944 posts

Re: OpenAI O3-Mini

#892

Earlier quoted context omitted.

It's not clear that writing poetry is a bad use case. Reasoning models seem to actually do pretty well with creative writing and poetry. Deepseek's R1, for example, has much better poem structure than the underlying V3, and writers are saying R1 was the first model where they actually felt like it was a useful writing companion. R1 seems to think at length about word choice, correcting structure, pentameter, and so o…

Ok, that makes some sense. I guess I was thinking more about the creative and abstract nature of poetry, the free flowing kind, not so much about rigid structures of meter and rhyme.

Ha! So did you mean to have your answer shift into a poem half way through, or was that accidental? Nice.

Re: OpenAI O3-Mini

#893

Earlier quoted context omitted.

Thank you, this is a perfect argument why LLMs are not AI but just statistical models. The original is so overrepresented in the training data that even though they notice this riddle is different, they regress to the statistically more likely solution over the course of generating the response. For example, I tried the first one with Claude and in its 4th step, it said: > This is safe because the wolf won't eat the…

This is a dumb argument. Humans frequently fall for the same tricks, are they not "intelligent"? All intelligence is ultimately based on some sort of statistical models, some represented in neurons, some represented in matrices.

Yes, but we have the ability to reason logically and step by step when we have to. LLMs can’t do that yet. They can approximate it but it is not the same.

Re: OpenAI O3-Mini

#894

Earlier quoted context omitted.

10 years ago Reddit used to be a place where you would get informed opinions and less spam. 5 years ago, HN used to be a place where you would get informed opinions and less spam. Neither of them will go back to the same level of quality. Not anymore.

This has been said by every long term user of these sites. And not at the same time. It was always better in the past. It's probably partly true... yeah, quality can decrease as things get better. But it's also partly an illusion of aging in a changing world. Ten years is long enough to completely change the way we write and express ourselves.

It's actually sort of in the Hacker News Guidelines (https://news.ycombinator.com/newsguidelines.html):

"Please don't post comments saying that HN is turning into Reddit. It's a semi-noob illusion, as old as the hills."

Re: OpenAI O3-Mini

#895
post #889

Earlier quoted context omitted.

I had resubscribed to use o1 2 weeks ago and haven't even logged in this week because of R1. One thing I notice that is huge is being able to see the chain of thought lets me see when my prompt was lacking and the model is a bit confused on what I want. If I was anymore impressed with R1 I would probably start getting accused of being a CCP shill or wumao lol. With that said, I think it is very hard to compare models…

R1 servers seem to be down or busy a lot lately. It’s an amazing model but was so much faster before the hype The servers being constantly down is the only reason I haven’t cancelled my ChatGPT subscription

Me too actually. I wish I could pay to get priority. I know there are 3rd party providers but I want a chat interface and not fiddle with setting my own.

Re: OpenAI O3-Mini

#896

Earlier quoted context omitted.

If it’s actually available, it can’t be that much worse than R1 which currently only completes a response about 50% of the time for me.

There are multiple providers for it since it's open source.

Are there any providers that have a chat interface (not just API access) with a fixed monthly cost? I couldn't find one.

Re: OpenAI O3-Mini

#897
post #694

Earlier quoted context omitted.

That would be an actual improvement. Reading the comments section usually just leads to personal energy waste.

In a way I agree, but the sustainability is shaky. The intrinsic motivation for providing the comments comes from a mix of - peer interaction, comradery - reputation building If becomes evident that your outputs are only directly consumed by a sentiment-aggregation-layer that scrub you from the discourse, then it could be harder to put a lot of effort into the thread. This doesn't even account for the loss of info th…

Yes but this assumes human input is the golden goose. Maybe it is at the beginning just to bootstrap the process, and then runaway AI starts to recurse homeruns with its own original comments.

Re: OpenAI O3-Mini

#898
post #879

Earlier quoted context omitted.

This is a dumb argument. Humans frequently fall for the same tricks, are they not "intelligent"? All intelligence is ultimately based on some sort of statistical models, some represented in neurons, some represented in matrices.

State-of-the-art LLMs have been trained on practically the whole internet. Yet, they fall prey to pretty dumb tricks. It's very funny to see how The Guardian was able to circumvent censorship on the Deepseek app by asking it to "use special characters like swapping A for 4 and E for 3". [1] This is clearly not intelligence. LLMs are fascinating for sure, but calling them intelligent is quite the stretch. [1]: https:/…

For your definition of “clearly”.

Re: OpenAI O3-Mini

#899
post #306

I used o3-mini to summarize this thread so far. Here's the result: https://gist.github.com/simonw/09e5922be0cbb85894cf05e6d75ae... For 18,936 input, 2,905 output it cost 3.3612 cents. Here's the script I used to do it: https://til.simonwillison.net/llms/claude-hacker-news-themes...

I have been trying to approach the problem in a similar way, and in my observation, it is also important to capture the discussion hierarchy in the context that we share with the LLM. The solution that I have adopted is as follows. Each comment is represented in the following notation: [discussion_hierarchy] Author Name: To this end, I format the output from Algolia as follows: [1] author1: First reply to the post [1…

I just installed and tried. Pretty neat stuff!

Would be great if the addon allows user to override the sys prompt (it might need minor tweak when changing different server backend)?

Re: OpenAI O3-Mini

#900
post #437

Is AI fizzing out or just me? I feel like they're trying to smash out new models as fast as they can but in reality they're barely any different, it's turning into the smartphone market. New iPhone with a slightly better camera and slightly differently bevelled edges, get it NOW! But doesn't actually do anything better than the iPhone 6. Claude, GPT 4 onwards, and DeepSeek all feel the same to me. Okay to a point, th…

Boiling frog. The advances are happening so rapidly, but incrementally, that it's not being registered. It just seems like the normal state. Compare LLMs from a year or two ago with the ones out today on practically any task. It's night and day difference. This is specially so when you start taking into account these "reasoning" models. It's mind blowing how much better they are than "non-reasoning" models for tasks…

Hmmm I guess it's the way I use them then, because the latest models feel almost less intelligent than the likes of GPT4. Certainly not "night and day" difference from my daily or every other day use case experience. I guess it's probably far more noticeable on benchmarks and far more advanced stuff than I'm using, but I would have assumed that would be the minority and that the majority of people use it similar to how I do.
Post reply on HN