Live data from Hacker News

The state of open source AI

stateofopensource.ai

71–80 of 379 posts

Re: The state of open source AI

#71
post #58
post #33

Earlier quoted context omitted.

Open models are probably also comparatively astronomically expensive to train - just less so than the frontier models because they’re somewhat smaller, +/- the creators are more incentivised to focus on getting more from less compute because they’re have to, +/- they rely on distillation of the frontier models and this is more efficient. But efficiencies aside; creation of open models still requires a lot of money an…

How does it work if people flock to open models but they're too expensive to train? What is the financial incentive to do so? I seem to understand open models are mostly coming from China, and the benefit of training and releasing them for 'free' is a powerful geopolitical weapon against the Western/US economy that at this point depends on OpenAI & co. to succeed. Will the West make open models illegal?

> Will the West make open models illegal?

We better not.

> What is the financial incentive to do so?

If we'd been sharing all along (as we should have been), we probably would have gotten even further along in the development of the tech.

Think of everything we could do if every researcher on the planet had first class access to the frontier. No academic fallback models. No crude API access. No limits, but direct access to the weights and the ability to lobotomize, splice, and dice.

We could pour intelligence from one container to the next without paying a tax or wearing a blindfold. All without spilling a drop.

*Open* *Must* *Win*

Re: The state of open source AI

#72
post #35

> Mozilla exists because one company tried to own the front door to the web, and an open community rose up to make sure it never could. I'd say that the front door to the web is pretty much owned by Google and Apple at this point given Firefox current marketshare. And maybe that's enough, maybe a future where a low percentage of open models keep the rest of the system honest but that doesn't seem the argument of this…

Mozilla exists because Google gives them billions to keep Google as default search engine.

Re: The state of open source AI

#73
post #70

Speculation: open models is what will kill Anthropic and OpenAI. Hyperscalers can run the models without a licensing fee. Apple can make them smaller and put them on the device. The frontier models are an edge and a liability. They're astronomically expensive to train. Without them, their models will fade into obscurity. Their marketing depends on people believing the models are meaningfully different, as people have…

Completely agree. Once I can reliably get open models doing what I am on Fable ultra I imagine I will switch for good. I am fortunate to have access to a decent bit of local RAM, 192GB of DDR5 at an OK speed. It is not enough and costs are well past absurd. In a few years time I envisage a setup that is sub $10k which can accomplish such tasks. The pace so far has been breakneck. That is all I personally need. That m…

This is easier to say as Fable is good (even SOTA). But people have been were saying this continuously for the current model and for now the improvement are still coming.

A better question is would you settle for o3 now or pay 20$ or 200$/month for fable ? Because o3 quality is available OSS.

It is like the new IPhone, in some sort. At some point come a feature many would like to have, despite diminishing returns.

We will see how long labs can keep up and what the scaling curve look like, but I would be more worried into losing sota status to Chinese companies than letting them take the open non-sota approach.

Re: The state of open source AI

#74
post #62
post #61

There isn't any open-source AI. There is Open AI (not to be confused with the closed company called OpenAI, which was unable to trademark its name). There's no open source AI both because the open source community doesn't have the resources to train a useful AI and because AI doesn't have source code.

Is training code and dataset not source?

Are they open?

Re: The state of open source AI

#75

Speculation: open models is what will kill Anthropic and OpenAI. Hyperscalers can run the models without a licensing fee. Apple can make them smaller and put them on the device. The frontier models are an edge and a liability. They're astronomically expensive to train. Without them, their models will fade into obscurity. Their marketing depends on people believing the models are meaningfully different, as people have…

Eventually they will kill the hyperscalers too because of privacy issues. It's better for a company to pay an uprfont cost and then run everything on premise that uploading their entire codebase to a third party service.

Re: The state of open source AI

#76

Speculation: open models is what will kill Anthropic and OpenAI. Hyperscalers can run the models without a licensing fee. Apple can make them smaller and put them on the device. The frontier models are an edge and a liability. They're astronomically expensive to train. Without them, their models will fade into obscurity. Their marketing depends on people believing the models are meaningfully different, as people have…

I still strongly believe Google Gemini has the best position for one simple reason: model maintenance. Accurate information is a moving target.

Open models are indeed very capable, but they will eventually become more specialized to the application to keep an edge. It makes perfect sense that the future shape of AI conforms to the landscape it was born out of.

Re: The state of open source AI

#77
post #36
post #33

Earlier quoted context omitted.

Open models are probably also comparatively astronomically expensive to train - just less so than the frontier models because they’re somewhat smaller, +/- the creators are more incentivised to focus on getting more from less compute because they’re have to, +/- they rely on distillation of the frontier models and this is more efficient. But efficiencies aside; creation of open models still requires a lot of money an…

I’m not exactly sure on the “how” but it only makes logical sense for (non-AI) companies to band together to fund the training of a shared model. Apple is a great example, AI is not their core business but they still require it. The only thing that took us down a different path is the vast sums of VC funding pumped into the AI companies.

If not for VC-funded LLMs there wouldn't be any LLMs.

Re: The state of open source AI

#78
post #29

Earlier quoted context omitted.

Just like opensource search engines killed google oh wait

I don’t even know the names of any open source search engines, but the open source models perform decently on various benchmarks and in personal experience. Was it ever even a claim that open source search engines were trying to outperform google, let alone kill it?

Yacy tried in the 2000s. I'm sure some magazines made headlines posing the question whether yacy is a google-killer

Re: The state of open source AI

#79
post #77
post #36

Earlier quoted context omitted.

I’m not exactly sure on the “how” but it only makes logical sense for (non-AI) companies to band together to fund the training of a shared model. Apple is a great example, AI is not their core business but they still require it. The only thing that took us down a different path is the vast sums of VC funding pumped into the AI companies.

If not for VC-funded LLMs there wouldn't be any LLMs.

[citation needed]

Historically speaking a lot of inventions have come about without things like VC investment. Either way, there’s probably little point in debating it, just because VC funded companies control the market now doesn’t mean they should indefinitely.

Re: The state of open source AI

#80
post #58
post #33

Earlier quoted context omitted.

Open models are probably also comparatively astronomically expensive to train - just less so than the frontier models because they’re somewhat smaller, +/- the creators are more incentivised to focus on getting more from less compute because they’re have to, +/- they rely on distillation of the frontier models and this is more efficient. But efficiencies aside; creation of open models still requires a lot of money an…

How does it work if people flock to open models but they're too expensive to train? What is the financial incentive to do so? I seem to understand open models are mostly coming from China, and the benefit of training and releasing them for 'free' is a powerful geopolitical weapon against the Western/US economy that at this point depends on OpenAI & co. to succeed. Will the West make open models illegal?

"You wouldn't download an LLM"
Post reply on HN