Live data from Hacker News

Open source AI must win

opensourceaimustwin.com

531–538 of 538 posts

Re: Open source AI must win

#531

Earlier quoted context omitted.

It's the most naive opinion that keeps getting shoveled around. You have a product that is viewed as essential by businesses, with revenue growing by 10x a year and geopolitical ramifications that have continued to rear their heads and your opinion is "this is all an unprofitable shill". It is extraordinary to me that people really believe this. Whether or not labs run at a loss today is absolutely irrelevant. There…

That businesses view it as essential...is not a profitability argument. Businesses also bought dot com infrastructure, telecom fiber, crypto platforms, metaverse tools, and overbuilt SaaS. The question is whether the AI application layer can charge more than its full cost and the costs are inference, infrastructure, depreciation, R&D, customer acquisition, support, compliance, security, and error remediation. The num…

> That businesses view it as essential...is not a profitability argument.

What do you mean, as in it doesn't imply profitability (and profitability of what?) today?

> Businesses also bought dot com infrastructure, telecom fiber, crypto platforms, metaverse tools, and overbuilt SaaS.

Yes, and like I said, there is absolutely a stable state where token economics are profitable.

> The question is whether the AI application layer can charge more than its full cost and the costs are inference, infrastructure, depreciation, R&D, customer acquisition, support, compliance, security, and error remediation.

Of course they can. If they couldn't, they wouldn't, and then someone would come in and charge the correct price for it because there is a tremendous amount of demand that will spur supply. I'm not sure what the mystery is here.

> The numbers so far do not inspire confidence. OpenAI reportedly did $4.3B in revenue in the first half of 2025 while burning $2.5B, and Microsoft said OpenAI related losses reduced its own quarterly net income by $3.1B. An MIT 2025 enterprise AI study found $30 to 40B spent on GenAI with 95% of organizations seeing zero return.

I do not find this particular argument inspires any confidence either -- everyone is trying to capture market share. How on earth, on a blog with ycombinator literally in the URL, do people not understand that lack of profitability now says absolutely nothing about profitability in the end state, after this growth phase?

You seem to imply by your MIT study citation that there is "evidence" that AI is not even profitable. Reading between the lines here, your theory is: the entire market is wrong, AI is a net negative ROI (and in your mind, for how long? forever?), and is inherently unprofitable. That is a crazy thing to think in light of all of the datapoints we have.

> One of the core technical reason is that hallucination destroy enterprise economics. If SAP hallucinated 2% of invoices, or Oracle returned fake rows 2% of the time, nobody would call that early stage friction. They would call it unusable for core operations.

Yet we see absolutely incredible adoption? In your view we'd see a spike and then all of those dumb CEOs will say "oh oops we were dumb" and adoption will go down. Hallucination does not destroy enterprise economics. Hallucination rates, by the way, have continued to decline after every model generation. Believe it or not there are incredibly valuable, smart, sustainable ways to incorporate AI. Even if it's just coding agent licenses, that alone is a powerful enough revenue driver.

> In legal AI, even specialized tools have been measured hallucinating 30% of the time. The problem is that as AI gets better it is confidently, plausibly wrong. That forces humans to verify it.

You seem to be conflating two things: one is that there is an inevitable eventuality where everyone will finally shut up about AI being unprofitable because we exit this growth phase and reach a more steady state market economics. The other thing is that we are not yet there. I'm not arguing about us not yet being there, though I do think the evidence even there is a bit cherry picked.

> So the cost does not disappear. It moves from doing the work to checking the work. AI coding has the same issue. If an autopilot got you there faster but one flight in ten became unstable unless the pilot constantly supervised it, that is not productivity.

That's just not going to be true for long. The reason you see coding agent adoption and other step changes in adoption + capability is that error ("hallucination" and other error modes) get to a low enough point where they are acceptable for a given application. This will eventually eat most applications. Yes we have to check things today, but this will be less and less important once we reach a point where humans are even less reliable than AI at _checking_ output from either themselves or other AI systems.

> For the bull case to work, the usage must explode, the quality must improve, prices must fall, reliability must rise, legal risk must shrink, and margins must expand and all this at once.

What??! this is all happening _right now_ right in front of you. Usage _is_ exploding in ways we didn't even dream about 8 months ago. Quality _is_ improving in very predictable ways (scaling laws anyone? benchmark trends?). Reliability _is_ rising _and thats why you see the adoption level you see for coding agents_.

Re: Open source AI must win

#532
post #519

Earlier quoted context omitted.

> The analogy falls apart very quickly. Without the training data, your modifications amount to virtually nothing compared to what these "versions" are, and the idea that you can maintain and improve on these models without the continual support of the company that owns the training data AND harnesses AND in general build instructions is not very credible. This is completely wrong, and sort of shows why what you are…

> You can post-train any LLM very easily without access to the original training data. Are you claiming this is e.g. what Alibaba spends their time doing? My point is that the usefulness of this is limited _in comparison to the one provided by having their training data AND mechanisms_.

> what Alibaba spends their time doing?

Not most of the time (pre-training takes a long time), but post-training is where most of the value is, yes.

Famously it is all that OpenAI did between GPT 4o and GPT 5.3 (or 5.2?) - they didn't manage to complete a pre-training run[1], and all their progress was done with post-training (!)

Post training what Cursor spends their time doing, and that has built a model that is competitive with the best coding models out there.

It isn't limited at all.

If you want to complain about something not being open source, complain about the lack of good open source RL environments (Prime Intellect excepted).

[1] https://newsletter.semianalysis.com/p/tpuv7-google-takes-a-s...

Re: Open source AI must win

#535
post #532

Earlier quoted context omitted.

> You can post-train any LLM very easily without access to the original training data. Are you claiming this is e.g. what Alibaba spends their time doing? My point is that the usefulness of this is limited _in comparison to the one provided by having their training data AND mechanisms_.

> what Alibaba spends their time doing? Not most of the time (pre-training takes a long time), but post-training is where most of the value is, yes. Famously it is all that OpenAI did between GPT 4o and GPT 5.3 (or 5.2?) - they didn't manage to complete a pre-training run[1], and all their progress was done with post-training (!) Post training what Cursor spends their time doing, and that has built a model that is co…

> It isn't limited at all.

Your very message is already showing that indeed it is limited, so dunno where you get that "that's where most of the value is". It is definitely not , and your very own link is showing that the the limit is there. This is not to say that there is _no_ value whatsoever, but that the value is negligible compared to what someone with _the real source_ could do. See the Rio model for another example.

> If you want to complain about something not being open source, complain about the lack of good open source RL environments (Prime Intellect excepted).

This is implicitly included when I was emphasizing "AND the software used to build the model", which I did for a reason.

Re: Open source AI must win

#536

Earlier quoted context omitted.

If humanity is over-reliant on frontier labs' models to perform work, the result is a dependence on the actual intelligence of these models -- not on human intelligence. This could be a small reason, on top of many others, why investors are throwing hundreds of billions of dollars a bit "carelessly" to these labs. It's fascinating seeing the models do the "hard work" (the deep, challenging thinking) for you. The conu…

I am going to try to cheer you up. Hear me out. One day, not long from now, I am going to buy a humanoid bot for 40k. This human android will 1) get my groceries, 2) make my elderly parents meals, 3) go to the backyard and plant 1 acre of corn, 4) paint my neighbors house. 5) get the kids from school 6) change my oil. What will happen? Massive. Deflation. What will you pay for an oil change? Corn? Meals? Everything i…

> am going to try to cheer you up. Hear me out. One day, not long from now, I am going to buy a humanoid bot for 40k. This human android will 1) get my groceries, 2) make my elderly parents meals, 3) go to the backyard and plant 1 acre of corn, 4) paint my neighbors house. 5) get the kids from school 6) change my oil.

Hahahahahahahha. You're delusional if you think this is true.

>Everything is about to be free.

Hahahahahahaha.

Re: Open source AI must win

#537
post #532

Earlier quoted context omitted.

> what Alibaba spends their time doing? Not most of the time (pre-training takes a long time), but post-training is where most of the value is, yes. Famously it is all that OpenAI did between GPT 4o and GPT 5.3 (or 5.2?) - they didn't manage to complete a pre-training run[1], and all their progress was done with post-training (!) Post training what Cursor spends their time doing, and that has built a model that is co…

> It isn't limited at all. Your very message is already showing that indeed it is limited, so dunno where you get that "that's where most of the value is". It is definitely not , and your very own link is showing that the the limit is there. This is not to say that there is _no_ value whatsoever, but that the value is negligible compared to what someone with _the real source_ could do. See the Rio model for another e…

> See the Rio model for another example

The Rio model (assuming you mean https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B) is a merge of Nex-N2_pro and Qwen: https://github.com/nex-agi/Nex-N2/issues/4

Neither of these ship training data.

> This is implicitly included when I was emphasizing "AND the software used to build the model", which I did for a reason.

If that's what you meant then sure, I agree.

Most providers do this though (at least to some extent) - the problem is that they can't eg ship Excel in their RL environment, and AFAIK there aren't any with alternatives.

Post reply on HN