Live data from Hacker News

Open source AI must win

opensourceaimustwin.com

311–320 of 538 posts

Re: Open source AI must win

#311

Isn't training material the biggest problem for truly open source LLMs (such that could compete with top tier models)? The computation part can be solved with money, but compiling a comprehensive training set that could be freely shared and free of copyright issues is pretty much impossible.

You don't need to have fully copyright-unencumbered datasets to build Open Source AI, as that (as you say) would be impossible. https://opensource.org/ai

Re: Open source AI must win

#312
It will win - in the sense that AI too will become a freely available resource. You can't stop progress.

My bet is that once cost-efficiency becomes a priority, we will figure out ways to get away from the expensive GPU infrastructure on figure out how to architect models for CPUs. I still remember that Microsoft paper about ternary weights.

Re: Open source AI must win

#315
post #84

Earlier quoted context omitted.

Ever calculate the cost of a computer in the 1960s, adjusted for inflation? Training is unfathomably expensive right now . What if a bunch of universities pooled their money? Or a bunch of nations pooled their money? Breakthroughs will eventually happen, optimization will occur, etc. People questioned whether there could ever be a viable open source operating system, yet Linux has been a viable option for a desktop e…

Yes, but have you seen what's happened to hardware improvements over the past 20 years? From the 1960s to the mid-2000s, every 10 years you'd have a big enough improvement in computing power that you could basically throw out the old computers and replace them with two new ones that were each massive improvements for the same cost (this varied, of course, from hyperbole to massive understatement). We achieved this by…

Moore's law isn't as relevant with parallel workloads. If you can keep building more lanes you don't have to worry about making faster cars.

Re: Open source AI must win

#316
post #84

Earlier quoted context omitted.

Ever calculate the cost of a computer in the 1960s, adjusted for inflation? Training is unfathomably expensive right now . What if a bunch of universities pooled their money? Or a bunch of nations pooled their money? Breakthroughs will eventually happen, optimization will occur, etc. People questioned whether there could ever be a viable open source operating system, yet Linux has been a viable option for a desktop e…

Yes, but have you seen what's happened to hardware improvements over the past 20 years? From the 1960s to the mid-2000s, every 10 years you'd have a big enough improvement in computing power that you could basically throw out the old computers and replace them with two new ones that were each massive improvements for the same cost (this varied, of course, from hyperbole to massive understatement). We achieved this by…

The bottleneck right now isn't making hardware more powerful, it's manufacturing it fast enough. Hardware right now is expensive because of scarcity, and those with a monopoly on it have no incentive to change that.

The Chinese would love to produce AI hardware much cheaper, but are blocked from doing so because US sanctions stop a Dutch company from selling them the machines capable of doing so. Coincidentally the companies with a monopoly happen to be in the US.

Re: Open source AI must win

#317
post #198

Who is going to fund it? Training is unfathomably expensive. You have either VC funded models looking for a return on investment, or CCP funded models looking to solidify authoritarian "model Chinese society". Maybe there are some university 4B models, but I doubt those will carry far.

Tbh, there really needs to be some legal precedent set that makes model distillation a legal activity. If the model makers can rip everyone else's work and launder information as if it's their own without giving credit back to the original creators, I don't see why it should be illegal to distill the models. It's the same thing the frontier model makers are doing to IP everywhere else.

I agree. But this won't happen in the US because Anthropic / OpenAI is a big ol economic recession risk because we levered ourselves to the tits and put our chips on them.

Re: Open source AI must win

#318
post #197

Earlier quoted context omitted.

As I replied to a child comment - this is a nice idea that just isn't tenable in reality. AI hardware isn't just hilariously faster than consumer GPUs, it's also hilariously more power-efficient and has hilariously better connectivity. Every one of these dimensions kills the idea. The far, FAR superior power efficiency means that even if you did harness every public GPU or GPU-like device on earth, you'd end up consu…

AI hardware is for inference, not training. Training uses normal HPC crap. Superpods aren't really power efficient, it's kind of a meme, and it stems from limiting the power draw of other components by having less of them. It's more of a rounding error. > you'd end up consuming so much excess electricity it would be cheaper on net to simply take the money that would have gone to the power bill and spend it on your ow…

Are you sure most of frontier cost isn't inference in RL environments?

Re: Open source AI must win

#319
There is not much open source AI .. there is open weight .. but anyways. Deepseek v4 is pretty much at the same level as the agents we had last year around November and it is an open weight model so I am hopeful.

Re: Open source AI must win

#320
I am really curious how long will it take for the open source models to hit current fable/mythos capabilities, KIMI 2.7 was launched recently and its quiet good for open source models its as good as Opus 4.6 maybe in practical applications not benchmarks so like 6 months to an year behind, after which the next step will be to wait for the day when we will be able to run mythos level intelligence on local hardware, Remember when 5MB storage was the size of a table?

A loooooot of work to be done for the above to happen

Post reply on HN