Live data from Hacker News

Open models by OpenAI

openai.com

291–300 of 909 posts

Re: Open models by OpenAI

#291
post #172

Earlier quoted context omitted.

I'll accept Meta's frontier AI demise if they're in their current position a year from now. People killed Google prematurely too (remember Bard?), because we severely underestimate the catch-up power bought with ungodly piles of cash.

And boy, with the $250m offers to people, Meta is definitely throwing ungodly piles of cash at the problem. But Apple is waking up too. So is Google. It's absolutely insane, the amount of money being thrown around.

It's insane numbers like that that give me some concern for a bubble. Not because AI hits some dead end, but due to a plateau that shifts from aggressive investment to passive-but-steady improvement.

Re: Open models by OpenAI

#293

Earlier quoted context omitted.

It's funny because I was thinking the opposite, the pricing seems way too high for a 5B parameter activation model.

Sure you're right, but if I can squeeze out o4-mini level utility out of it, but its less than quarter the price, does it really matter?

Yes

Re: Open models by OpenAI

#294
I just tried it on open router but i was served by cerebras. Holy... 40,000 tokens per second. That was SURREAL.

I got a 1.7k token reply delivered too fast for the human eye to perceive the streaming.

n=1 for this 120b model but id rank the reply #1 just ahead of claude sonnet 4 for a boring JIRA ticket shuffling type challenge.

EDIT: The same prompt on gpt-oss, despite being served 1000x slower, wasn't as good but was in a similar vein. It wanted to clarify more and as a result only half responded.

Re: Open models by OpenAI

#296
post #60

Earlier quoted context omitted.

It only seems like that if you haven't been following other open source efforts. Models like Qwen perform ridiculously well and do so on very restricted hardware. I'm looking forward to seeing benchmarks to see how these new open source models compare.

Agreed, these models seem relatively mediocre to Qwen3 / GLM 4.5

Yes, but they are suuuuper safe. /s

So far I have mixed impressions, but they do indeed seem noticeably weaker than comparably-sized Qwen3 / GLM4.5 models. Part of the reason may be that the oai models do appear to be much more lobotomized than their Chinese counterparts (which are surprisingly uncensored). There's research showing that "aligning" a model makes it dumber.

Re: Open models by OpenAI

#297

Open models are going to win long-term. Anthropics' own research has to use OSS models [0]. China is demonstrating how quickly companies can iterate on open models, allowing smaller teams access and augmentation to the abilities of a model without paying the training cost. My personal prediction is that the US foundational model makers will OSS something close to N-1 for the next 1-3 iterations. The CAPEX for the fou…

> Open models are going to win long-term.

[1 of 3] For the sake of argument here, I'll grant the premise. If this turns out to be true, it glosses over other key questions, including:

For a frontier lab, what is a rational period of time (according to your organizational mission / charter / shareholder motivations*) to wait before:

1. releasing a new version of an open-weight model; and

2. how much secret sauce do you hold back?

* Take your pick. These don't align perfectly with each other, much less the interests of a nation or world.

Re: Open models by OpenAI

#298

I wonder if this is a PR thing, to save face after flipping the non-profit. "Look it's more open now". Or if it's more of a recruiting pipeline thing, like Google allowing k8s and bazel to be open sourced so everyone in the industry has an idea of how they work.

I think it’s both of them, as well as an attempt to compete with other makers of open-weight models. OpenAI certainly isn’t happy about the success of Google, Facebook, Alibaba, DeepSeek…

Re: Open models by OpenAI

#300
> Training: The gpt-oss models trained on NVIDIA H100 GPUs using the PyTorch framework [17] with expert-optimized Triton [18] kernels2. The training run for gpt-oss-120b required 2.1 million H100-hours to complete, with gpt-oss-20b needing almost 10x fewer.

This makes DeepSeek's very cheap claim on compute cost for r1 seem reasonable. Assuming $2/hr for h100, it's really not that much money compared to the $60-100M estimates for GPT 4, which people speculate as a MoE 1.8T model, something in the range of 200B active last I heard.

Post reply on HN