Live data from Hacker News

Open source AI is the path forward

about.fb.com

581–590 of 936 posts

Re: Open source AI is the path forward

#581

"Eventually though, open source Linux gained popularity – initially because it allowed developers to modify its code however they wanted ..." I find the language around "open source AI" to be confusing. With "open source" there's usually "source" to open, right? As in, there is human legible code that can be read and modified by the user? If so, then how can current ML models be open source? They're very large matric…

Coming up with the words and concepts to describe the models is a challenge.

Does the training data require permission from the copyright holder to use? Are the weights really open source or more like compiled assembly?

Re: Open source AI is the path forward

#584
Ok one notable difference: did the linux researchers of yore warn about adversarial giants getting this tech? Or is this unique to the current moment? That for me is the largest question when considering the logical progression on "linux open is better therefore ai open is better".

Re: Open source AI is the path forward

#585

Ok one notable difference: did the linux researchers of yore warn about adversarial giants getting this tech? Or is this unique to the current moment? That for me is the largest question when considering the logical progression on "linux open is better therefore ai open is better".

We can't open source Linux because bad people might run servers?

Can you imagine the disinformation they could spread with those? With enough of them you could have a massively global site made entirely for spreading it. God what if such a thing got into the hands of an egocentric billionaire?

Re: Open source AI is the path forward

#586
post #43

Earlier quoted context omitted.

It would be better that the government removes IP on such technology for public use, like drugs got generics. This way the government pays 2'500 USD per card, not 40'000 USD or whatever absurd.

> better that the government removes IP on such technology for public use, like drugs got generics You want to punish NVIDIA for calling its shots correctly? You don't see the many ways that backfires?

They said remove legally-enforced monopolies on what they produce. Many of these big firms made their tech with millions to billions of taxpayer dollars at various points in time. If we’ve given them millions, shouldn’t we at least get to make independent implementations of the tech we already paid for?

Re: Open source AI is the path forward

#587
post #91

Earlier quoted context omitted.

How about using some of that money to develop CUDA alternatives so everyone is not paying the Nvidia tax?

Either you port Tensorflow (Apple)[1] or PyTorch to your platform or you allow CUDA to run on your hardware (AMD) [2]. Companies are incentives to not have NVIDIA having a monopoly but the thing is that CUDA is a huge moat due to compatibility of all frameworks and everyone knows it. Also, all of the cloud or on premises providers use NVIDIA regardless. [1] https://developer.apple.com/metal/tensorflow-plugin/ [2] htt…

>> Either you port Tensorflow (Apple)[1] or PyTorch to your platform or you allow CUDA to run on your hardware (AMD) [2]. Companies are incentives to not have NVIDIA having a monopoly but the thing is that CUDA is a huge moat due to compatibility of all frameworks and everyone knows it. Also, all of the cloud or on premises providers use NVIDIA regardless.

This never made sense to me -- Apple could easily hire top talent to write Apple Silicon bindings for these popular libraries. I work at a creative ad agency, we have tons of high end apple devices yet the neural cores sit unused most of the time.

Re: Open source AI is the path forward

#588

Earlier quoted context omitted.

How in the heck is an open source model that is free and open today going to lock me down, down the line? This is nonsense. You can literally run this model forever if you use NixOS (or never touch your windows, macos or linux install again). Zuck can't come back and molest it. Ever. The best I can tell is that their self-interest here is more about gathering mindshare. That's not a terrible motive; in fact, that's a…

Yeah because history isn't absolutely littered with examples of shiny things being dangled in front of people with the intent to entrap them /s. Can you really say this model will still be useful in 2 years, 5 years for you ? And that FB's stance on these models will still be open source at that time once they incrementally make improvements? Maybe, maybe not. But FB doesn't give anything away for free, and the fact…

> But FB doesn't give anything away for free, and the fact that you think so is your blindness, not mine

Is it, though? They are literally giving this away "for free". https://dev.to/llm_explorer/llama3-license-explained-2915 Unless you build a service with it that has over 700 million monthly users (read: "problem anyone would love to have"), you do not have to re-negotiate a license agreement with them. Beyond that, it can't "phone home" or do any other sorts of nefarious shite. The other limitations there, which you can plainly read, seem not very restrictive.

Is there a magic secret clause conspiracy buried within the license agreement that you believe will be magically pulled out at the worst possible moment? >..Sometimes, good things happen. Sorry you're "too blinded" by past hurt experience to see that, I guess

Re: Open source AI is the path forward

#589
post #352

Earlier quoted context omitted.

As you've rightly pointed out, we have the mechanism, now let's fund it properly! I'm in Canada, and our science funding has likewise fallen year after year as a proportion of our GDP. I'm still benefiting from A100 clusters funded by tax payer dollars, but think of the advantage we'd have over industry if we didn't have to fight over resources.

Where do you get access to those as a member of the general public?

I'm going to guess it's Compute Canada, which I don't think we non-academics have access to.

Re: Open source AI is the path forward

#590

Earlier quoted context omitted.

You can find the entire Llama 3.0 pretraining set here: https://huggingface.co/datasets/HuggingFaceFW/fineweb 15T tokens, 45 terrabytes. Seems fairly open source to me.

Where has Facebook linked that? I can't find anywhere that they actually published that.

Many companies stopped publishing their data sets after people published evidence they were mass, copyright infringement. They dropped the specifics of pretraining data from the model cards.

Aside from licensing content, that content creators don’t like redistribution means a lawful model would probably only use Gutenberg’s collection and permissive code. Anything else, including Wikipedia, usually has licensing requirements they might violate.

Post reply on HN