Live data from Hacker News

Open source AI must win

opensourceaimustwin.com

251–260 of 538 posts

Re: Open source AI must win

#251

I've been contemplating a decentralized model training system for some time using volunteer machines that we all contribute. But, it is astronomically difficult. The communication speeds are untenable. And, there is the issue of data poisoning from untrusted nodes. I've almost cracked that last issue with a self-healing checkpointed rollback system that doesn't have to throw out anything that follows the corrupt datu…

This could be of interest to you: https://thealliance.ai/projects/tapestry

Man, that project is such bait for my particular sensibilities but just looking at the copy about not sharing your data and only sharing weights has me feeling very disappointed in the project already. I would want a project like this to not elide fact that sharing your weight updates probably effectively means sharing your data too.

Re: Open source AI must win

#252
post #198

Earlier quoted context omitted.

Tbh, there really needs to be some legal precedent set that makes model distillation a legal activity. If the model makers can rip everyone else's work and launder information as if it's their own without giving credit back to the original creators, I don't see why it should be illegal to distill the models. It's the same thing the frontier model makers are doing to IP everywhere else.

And which leading country is going to go for allowing other countries to distill their models?

It doesn’t have to be the leading countries, if the EU allows it, it is good enough to create a market for distilled models

Re: Open source AI must win

#253

Earlier quoted context omitted.

And which leading country is going to go for allowing other countries to distill their models?

It doesn’t have to be the leading countries, if the EU allows it, it is good enough to create a market for distilled models

But EU is way behind right?

Re: Open source AI must win

#254
post #197

I've been contemplating a decentralized model training system for some time using volunteer machines that we all contribute. But, it is astronomically difficult. The communication speeds are untenable. And, there is the issue of data poisoning from untrusted nodes. I've almost cracked that last issue with a self-healing checkpointed rollback system that doesn't have to throw out anything that follows the corrupt datu…

As I replied to a child comment - this is a nice idea that just isn't tenable in reality. AI hardware isn't just hilariously faster than consumer GPUs, it's also hilariously more power-efficient and has hilariously better connectivity. Every one of these dimensions kills the idea. The far, FAR superior power efficiency means that even if you did harness every public GPU or GPU-like device on earth, you'd end up consu…

What makes you think Deepseek or GLM won't catch up to Fable level? Why would there be a break in the trend now?

Re: Open source AI must win

#255

Earlier quoted context omitted.

Open source models don't need to be anywhere near as good as Claude Mythos or even Claude Sonnet to 'win'. Open source 'winning' just means that there exists at least one open source alternative to closed models which is as good as, say, GPT 4... I mean, we're essentially there already with Google Gemma models. As a software engineer, I didn't notice any difference in my productivity since Sonnet. Of course Opus is b…

> Open source 'winning' just means that there exists at least one open source alternative to closed models which is as good as, say, GPT 4... I mean, we're essentially there already with Google Gemma models. Is this really true? We just don't know what the maximum capability of AI is. If it turns out AI can be as intelligent and capable as something like Data from Star Trek, no one is going to be thinking GPT 4 is go…

>>We just don't know what the maximum capability of AI is

For all theory purposes there is no limit. Thats what the latest loop engineering trend is about, you are asking AI to find solutions to a problem going by listing steps, and if solution not found in those steps, to treat each step as a separate problem and repeat the process until the master solution to the master problem is found.

Once a solution is found, or new data/insights are generated through this process, the LLM can be trained on this. So in theory you can just keep going like this forever.

Secondly. This is as close to agency you can build inside a machine.

Practically speaking, hardware is a limit. But that can scale up with time.

So we are already looking at some kind of runaway intelligence even if not sentient.

Re: Open source AI must win

#256

I've been contemplating a decentralized model training system for some time using volunteer machines that we all contribute. But, it is astronomically difficult. The communication speeds are untenable. And, there is the issue of data poisoning from untrusted nodes. I've almost cracked that last issue with a self-healing checkpointed rollback system that doesn't have to throw out anything that follows the corrupt datu…

The biggest problem is accuracy and integrity of the actors in the project.

Re: Open source AI must win

#257

I've been contemplating a decentralized model training system for some time using volunteer machines that we all contribute. But, it is astronomically difficult. The communication speeds are untenable. And, there is the issue of data poisoning from untrusted nodes. I've almost cracked that last issue with a self-healing checkpointed rollback system that doesn't have to throw out anything that follows the corrupt datu…

[deleted]

Re: Open source AI must win

#258

Earlier quoted context omitted.

Monopoly capitalism and finance capitalism took reigns of markets more than a century ago. The state serves these huge interests. Everybody knows AI firms pirated to train, nothing will come of it. A plain example of classist application of law. The reason for the willy nilly application of their own laws will always be 'national security', of course, since they own infrastructure their interests are a national secur…

No state, anywhere, has the right to rule or even exist. All states are terroristic parasite gangs, all states [no exceptions] . Your state exists because there is no one else capable of challenging it [no outsider or internal armed militia]. Your state is merely the gang which reigns supreme in your territory - constitutions, democracy, and other grievance pressure relief systems be damned. You don't get to vote or…

>No state, anywhere, has the right to rule or even exist.

No person has an inherent right to exist either. Rights, just like states, or property, or gender, are social constructs. They exist because enough people believe they exist and behave accordingly.

Re: Open source AI must win

#259
post #198

Earlier quoted context omitted.

Tbh, there really needs to be some legal precedent set that makes model distillation a legal activity. If the model makers can rip everyone else's work and launder information as if it's their own without giving credit back to the original creators, I don't see why it should be illegal to distill the models. It's the same thing the frontier model makers are doing to IP everywhere else.

And which leading country is going to go for allowing other countries to distill their models?

If your country doesn't have any leading models, why not legalize distillation, either explicitly or implicitly?

(Chinese labs famously distilled American models, and that seems to be going well for them. They now have a competitive industry, home-grown talent choosing not to leave, and they now can truly compete without distillation).

Re: Open source AI must win

#260

Who is going to fund it? Training is unfathomably expensive. You have either VC funded models looking for a return on investment, or CCP funded models looking to solidify authoritarian "model Chinese society". Maybe there are some university 4B models, but I doubt those will carry far.

It’s expensive, but not unfathomably, esp in an open source setting where capable people might contribute high quality data for post training (worked problems, code reviews, feedback, …) gratis instead of at immense cost.
Post reply on HN