Live data from Hacker News

Open source AI is the path forward

about.fb.com

211–220 of 936 posts

Re: Open source AI is the path forward

#211

Only if it is truly open source (open data sets, transparent curation/moderation/censorship of data sets, open training source code, open evaluation suites, and an OSI approved open source license). Open weights (and open inference code) is NOT open source, but just some weak open washing marketing. The model that comes closest to being TRULY open is AI2’s OLMo. See their blog post on their approach: https://blog.all…

Yeah, though I do wonder for a big model like 405B if the original training recipe, really matters for where models are heading, practically speaking which is smaller and more specific?

I imagine its main use would be to train other models by distilling them down with LoRA/Quantization etc(assuming we have a tokenizer). Or use them to generate training data for smaller models directly.

But, I do think there is always a way to share without disclosing too many specifics, like this[1] lecture from this year's spring course at Stanford. You can always say, for example:

- The most common technique for filtering was using voting LLMs (without disclosing said llms or quantity of data).

- We built on top of a filtering technique for removing poor code using ____ by ____ authors (without disclosing or handwaving how you exactly filtered, but saying that you had to filter).

- We mixed certain proportion of this data with that data to make it better (without saying what proportion)

[1] https://www.youtube.com/watch?v=jm2hyJLFfN8&list=PLoROMvodv4...

Re: Open source AI is the path forward

#212

Earlier quoted context omitted.

I will steelman the idea that a tokenizer and weights are all you need for the "source" of an LLM. They are components that can be modified, redistributed and when put together, reproduce the full experience intended. If we insist upon the release of training data with Open models, you might as well kiss the idea of usable Open LLMs out the door. Most of the content in training datasets like The Pile are not licensed…

> Most of the content in training datasets like The Pile are not licensed for redistribution in any way shape or form. But distributing the weights is a "form" of distribution. You can recover many items of the dataset (most easily, the outliers) by using the weights. Just because they are codified in a non-readily accessible way, does not mean that you are not distributing them. It's scary to think that "training" i…

The weights are a transformed, lossy and non-complete permutation of the training material. You cannot recover most of the dataset reliably, which is what stops it from being an outright replacement for the work it's trained on.

> does not mean that you are not distributing them.

Except you literally aren't distributing them. It's like accusing me of pirating a movie because I sent a screenshot or a scene description to my friend.

> It's scary to think that "training" is becoming a thinly veiled way to strip copyright of works.

This is the way it's been for years. Google is given Fair Use for redistributing incomplete parts of copywritten text materials verbatim, since their application is transformative: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

Or Corellium, who won their case to use copywritten Apple code in novel and transformative ways: https://www.forbes.com/sites/thomasbrewster/2023/12/14/apple...

Copyright has always been a limited power.

Re: Open source AI is the path forward

#215

It's alarming that he refers to llama as if it was open source. The definition of free software (and open source, for that mater), is well-established. The same definition applies to all programs, whether they are "AI" or not. In any case, if a program was built by training against a dataset, the whole dataset is part of the source code. Llama is distributed in binary form, and it was built based on a secret dataset.…

> In any case, if a program was built by training against a dataset, the whole dataset is part of the source code. I'm not sure why I keep seeing this. What is the equivalent of the training data for something like the Linux kernel?

> What is the equivalent of the training data for something like the Linux kernel?

It's the source code.

For the linux kernel:

          compile(sourcecode) = binary
For llama:

          train(data) = weights

Re: Open source AI is the path forward

#217

Earlier quoted context omitted.

> Why do people keep mislabeling this as Open Source? The whole point of calling something Open Source is that the "magic sauce" of how to build something is publicly available, so I could built it myself if I have the means. But without the training data publicly available, could I train Llama 3.1 if I had the means? I don't think not releasing the commit history of a project makes it not Open Source, this seems lik…

For the freedom to change to be effective, a user must be given the software in a form they can modify. Can you tweak an LLM once it's built? (I genuinely don't know the answer)

Yes, you can finetune Llama: https://llama.meta.com/docs/how-to-guides/fine-tuning/

Re: Open source AI is the path forward

#218
post #104

> Today we’re taking the next steps towards open source AI becoming the industry standard. We’re releasing Llama 3.1 405B, the first frontier-level open source AI model, Why do people keep mislabeling this as Open Source? The whole point of calling something Open Source is that the "magic sauce" of how to build something is publicly available, so I could built it myself if I have the means. But without the training d…

> Why do people keep mislabeling this as Open Source? The whole point of calling something Open Source is that the "magic sauce" of how to build something is publicly available, so I could built it myself if I have the means. But without the training data publicly available, could I train Llama 3.1 if I had the means? I don't think not releasing the commit history of a project makes it not Open Source, this seems lik…

> I don't think not releasing the commit history of a project makes it not Open Source,

Right, I'm not talking about the commit history, but rather that anyone (with means) should be able to produce the final artifact themselves, if they want. For weights like this, that requires at least the training script + the training data. Without that, it's very misleading to call the project Open Source, when only the result of the training is released.

> What's important is you can download it, run it, modify it, and re-release it

But I literally cannot download the project, build it and run it myself? I can only use the binaries (weights) provided by Meta. No one can modify how the artifact is produced, only modify the already produced artifact.

That's like saying that Slack is Open Source because if I want to, I could patch the binary with a hex editor and add/remove things as I see fit? No one believes Slack should be called Open Source for that.

Re: Open source AI is the path forward

#219
post #111

Earlier quoted context omitted.

> reducing the duration of the government granted monopolies on chip technology that is obsolete well before the default duration of 20 years is over Why do you think these private entities are willing to invest the massive capital it takes to keep the frontier advancing at that rate? > I do want to limit the amount we reward NVIDIA for calling the shots correctly to maximize the benefit to society Why wouldn't NVIDI…

> Why do you think these private entities are willing to invest the massive capital it takes to keep the frontier advancing at that rate? Because whether they make 100x or 200x they make a shitload of money. > Why wouldn't NVIDIA be a solid steward of that capital given their track record? The problem isn't who is the steward of the capital. The problem is that economically efficient thing to do for a single company…

> Because whether they make 100x or 200x they make a shitload of money.

It's not a certainty that they 'make a shitload of money'. Reducing the right tail payoffs absolutely reduces the capital allocated to solve problems - many of which are risky bets.

Your solution absolutely decreases capital investment at the margin, this is indisputable and basic economics. Even worse when the taking is not due to some pre-existing law, so companies have to deal with the additional uncertainty of whether & when future people will decide in retrospect that they got too large a payoff and arbitrarily decide to take it from them.

Re: Open source AI is the path forward

#220
post #91

“The Heavy Press Program was a Cold War-era program of the United States Air Force to build the largest forging presses and extrusion presses in the world.” This ”program began in 1944 and concluded in 1957 after construction of four forging presses and six extruders, at an overall cost of $279 million. Six of them are still in operation today, manufacturing structural parts for military and commercial aircraft” [1].…

How about using some of that money to develop CUDA alternatives so everyone is not paying the Nvidia tax?

Please start with the Windows Tax first for Linux users buying hardware...and the Apple Tax for Android users...
Post reply on HN