Live data from Hacker News

Meta to release open-source commercial AI model

zdnet.com

101–110 of 168 posts

Re: Meta to release open-source commercial AI model

#101
post #73

Earlier quoted context omitted.

I read it in all such discussions. What does it mean? I just have a very high level understanding of AI models. No idea how things work under the hood or what knobs can be tweaked.

The source code is all the supporting code needed to run inference on the weights. This is usually python and in the case of llama it's already open source. Usually the source code is referred to as the "model". You can kind of think of the weights as a settings file in a normal desktop application. The desktop app has its own source code and loads in the settings file at runtime. It can load different settings files…

I thought model is the output of training. It's a binary file black box. That's what I had read somewhere.

Re: Meta to release open-source commercial AI model

#102
post #71

Earlier quoted context omitted.

Sometimes, I wonder what if someone in XYZ country downloads whole of Z-Library/Libgen, all the books ever printed, and all the papers ever published, all the newspapers and so on. and releases the model open source. There are jurisdictions with Lax rules. And they will have much better knowledge, answers, etc than the western, Lawyer approved models. Sometimes knowledge needs to be set free I guess.

The production of knowledge needs to be funded as it isn’t “free”. Copyright and licensing is one model that has worked for a long time. It has flaws, but it has produced good things. At this point with the quality of current web content and the collapse of journalism as an industry I think we can say online ads have utterly failed as a replacement income stream. Unless you want all LLM to say “I’m sorry the data I w…

But again, "funding" is merely common and/or one step in the process. It's not always necessary and is definitely never sufficient, and I think when you bring it up, the mental model that people have is of the incorrect scale?

Put differently, we consider -- but don't think a whole lot about -- about Wikipedia's "funding," because that's NOT the most important part/innovation of that model.

We should better answer what is?

Re: Meta to release open-source commercial AI model

#103
post #47

From the recent story about the Sarah Silverman lawsuit: The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together…

Training and copyright is going to be interesting, people can be trained on “illegally obtained” books too yet you’ll probably going to be hard pressed to make an argument that any employee who downloaded a book or a paper from “libre library” could be used as fruit of the poisonous tree argument down the line.

If the company supplied the employee with the “illegally obtained” books, that could be reason to view the situation differently than an employee acting on their own.

Since the company is obtaining + providing these models with 100% of their input data, it could be argued they have some responsibility to verify the legality of their procurement of the data.

Re: Meta to release open-source commercial AI model

#104

Earlier quoted context omitted.

> It's not like Meta can remove these books from the training set without retraining from scratch (or at least the last checkpoint before they were used). They probably can: https://github.com/zjunlp/EasyEdit > I wonder if this is going to cause issues down the road. There are some popular Stable Diffusion models, being run in small businesses, that I am certain have CSAM in them because they have a particular 4chan…

I’ve been wondering when the landmark moral panic would start against Civit.AI and the coomer crowd. People have no idea just how much porn is being produced by this stuff. One of the top textual inversions right now is a… age slider… ( https://civitai.com/models/65214/age-slider ) ewww. It’s also extremely well rated and reviewed on there. I’m terrified at the impending backlash because depending on what happens the…

People have been saying this about underage hand drawn hentai forever, but its still around.

Not that I am disagreeing with you. What I find particularly disturbing are the paid services for this.

Also, I have seen 2 seperate OnlyFans pimps ask for help in a text generation chatroom. Something about automating "private" texting from their "girls."

Re: Meta to release open-source commercial AI model

#105

Earlier quoted context omitted.

Copyright laws should be amended to allow this scenario. If I read a book and write about it in a blog, it is considered review. Why shouldn’t we allow companies to do the same to train their models? Overall it will benefit society more than it hurts some rich authors.

Couldn’t they just buy the ebook and call it a day? The rich people are the people training LLMs not the authors lol

I doubt it makes a difference whether they purchase the ebook or not. And probably a bunch of them aren't even available as ebooks legitimately, people scan books and upload them to zlibrary etc.

Re: Meta to release open-source commercial AI model

#106

I have a 128 core Threadripper, a 2080 Ti and a 3080 Ti. How can I play with open source LLM's locally?

If you're just looking to play with something locally for the first time, this is the simplest project I've found and has a simple web UI: https://github.com/cocktailpeanut/dalai

It works for 7B/13B/30B/65B LLaMA and Alpaca (fine-tuned LLaMA which definitely works better). The smaller models at least should run on pretty much any computer.

Re: Meta to release open-source commercial AI model

#107

Earlier quoted context omitted.

I’ve been wondering when the landmark moral panic would start against Civit.AI and the coomer crowd. People have no idea just how much porn is being produced by this stuff. One of the top textual inversions right now is a… age slider… ( https://civitai.com/models/65214/age-slider ) ewww. It’s also extremely well rated and reviewed on there. I’m terrified at the impending backlash because depending on what happens the…

People have been saying this about underage hand drawn hentai forever, but its still around. Not that I am disagreeing with you. What I find particularly disturbing are the paid services for this. Also, I have seen 2 seperate OnlyFans pimps ask for help in a text generation chatroom. Something about automating "private" texting from their "girls."

It’s trivial to use these methods to produce real looking images, or even stuff in the likeness of real people…

Re: Meta to release open-source commercial AI model

#108
post #41

Earlier quoted context omitted.

Quite the opposite, this is great for Meta's competitors. Meta is not trying to get market share with this strategy, it's trying to commoditize their complements ( https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/ ) Content is a complement to a social network: the cheaper it is to create content, the more content is available, the easier it is to optimize a feed, the larger the time people spend in the pla…

kind of a dystopian nightmare world in which large corporations utilize AI to create low cost, infinite content that humans engage with (mostly content catering to the human tendency for tribalism, prestige, sexual desires etc...), sounds like we are creating a world similar to the Matrix.

Infinite flame wars. Can't wait!

Re: Meta to release open-source commercial AI model

#109
post #34
post #27

Maybe they've solved the fingerprinting problem and can identify text generated from their model, and this is a way of discovering the market they can sell more advanced models to directly. B2B leadgen...

I mean you could probably just train it on some sequence s.t. the model identifies itself, would be hard to detect that

That would prbably work to detect if e.g. OpenAI or Anthropic start using their weights directly. It wouldn't detect whether e.g. a blog was generated with their model or not.

Re: Meta to release open-source commercial AI model

#110
post #47

From the recent story about the Sarah Silverman lawsuit: The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together…

its not deemed illegal yet

its in a weird place imo, with japan ruling that anything goes for AI data, other countries are put under pressure to allow the same

ie,

you're allowed to scrape the web

you're allowed to take what you scrape and put it in a database

you're allowed to use your database to inform on decisions you might make, or content you might create

but once you put AI model in the mix, all of a sudden there's problems, despite the fact that making the model is 10000% harder than doing all of the points mentioned above, the problem of using someone else's work somehow becomes a problem when it never was before

and if truly free and open source LLMs come into the game, then might the corporate ones become crippled from copyright? that's bad for business

Post reply on HN