Live data from Hacker News

Hello Dolly: Democratizing the magic of ChatGPT with open models

databricks.com

111–120 of 194 posts

Re: Hello Dolly: Democratizing the magic of ChatGPT with open models

#111
post #69

I might be having a moment - but I can't find any links to a git repo, huggingface, or anything about the models/weights/checkpoints directly from the article. I just see a zip download that AFAIK also doesn't contain the weights/checkpoints. I find this a bit odd, the contents of the zip (from the gdrive preview) look like they should be in a git repo, and I assume they download the model from somewhere? GDrive usua…

The README also says this: > This fine-tunes the [GPT-J 6B]( https://huggingface.co/EleutherAI/gpt-j-6B ) model on the [Alpaca]( https://huggingface.co/datasets/tatsu-lab/alpaca ) dataset using a Databricks notebook. > Please note that while GPT-J 6B is Apache 2.0 licensed, the Alpaca dataset is licensed under Creative Commons NonCommercial (CC BY-NC 4.0). ...so, this cannot be used for commercial purposes

As far as I know the copyright situation for models is ambiguous and also depends on the region. In the US you can't copyright data made by an automated process but you can in the EU, or something to that effect.

Re: Hello Dolly: Democratizing the magic of ChatGPT with open models

#112
post #84
post #69

Earlier quoted context omitted.

The README also says this: > This fine-tunes the [GPT-J 6B]( https://huggingface.co/EleutherAI/gpt-j-6B ) model on the [Alpaca]( https://huggingface.co/datasets/tatsu-lab/alpaca ) dataset using a Databricks notebook. > Please note that while GPT-J 6B is Apache 2.0 licensed, the Alpaca dataset is licensed under Creative Commons NonCommercial (CC BY-NC 4.0). ...so, this cannot be used for commercial purposes

Essentially every model worth anything has been trained on a unfathomably large amount of data under copyright, with every possible licensing scheme you could imagine, under the assumption that it is fair use. While you can argue that it's all built on a house of cards (and a court may well agree with you) it's kind of arbitrary to draw a line here.

> under the assumption that it is fair use.

No, because you as a human looking at "art" over your lifetime and learning from it is not "fair use" of the copyright, it's no-use at all. This is the crux of every argument for both for language models and AI Art models, that their tools are learning how to draw, learning what styles and characteristics of input art correspond the most with words, and creating art with that knowledge just like any other human, not simply collaging together different pieces of art.

Re: Hello Dolly: Democratizing the magic of ChatGPT with open models

#113
post #15

This is the real risk to OpenAI's business model. If it turns out that you can get most of the same outcome with drastically smaller and cheaper models, then OpenAI is going to have a hell of a time keeping customers around as it will just be a race to the bottom on price and bigger, more expensive models will lose just from a hardware cost standpoint.

I wouldn't underestimate the power of momentum

Re: Hello Dolly: Democratizing the magic of ChatGPT with open models

#114
post #101

Earlier quoted context omitted.

Thought the same until I read this: > Contact us at hello-dolly@databricks.com if you would like to get access to the trained weights.

data transfer might actually be the problem there not something like trying to hide the model

bittorrent, come on

Re: Hello Dolly: Democratizing the magic of ChatGPT with open models

#115
post #103

Does anyone else find it ironic that all these ChatGPT "clones" are popping up when OpenAi is supposed to be the ones open sourcing and sharing their work? I guess: "You Either Die A Hero, Or You Live Long Enough To See Yourself Become The Villain"?

> when OpenAi is supposed to be the ones open sourcing and sharing their work? OpenAI renounced being open source. Don't let the name fool you.

I think all of the "AI alignment" talk is mostly fearmongering. It's a cunningly smart way to get ignorant people scared enough of AI so they have no choice but to trust the OpenAI overlords when they say they need AI to be closed. Then OpenAI gets a free pass to be the gatekeeper of the model, and people stop questioning the fact that they went from Open to Closed.

AI being tuned to be "safe" by an exceedingly small set of humans is the thing we should be afraid of. It's the effective altruism effect: if you bombard people enough with "safety" and "alignment" speak, they will look past the fact that you're mainly interested in being a monopoly. My bigger conspiracy theory is that Bill Gates getting behind "AI alignment" is a calculated move to get people to look past Microsoft's unilateral involvement.

Re: Hello Dolly: Democratizing the magic of ChatGPT with open models

#116
post #69

I might be having a moment - but I can't find any links to a git repo, huggingface, or anything about the models/weights/checkpoints directly from the article. I just see a zip download that AFAIK also doesn't contain the weights/checkpoints. I find this a bit odd, the contents of the zip (from the gdrive preview) look like they should be in a git repo, and I assume they download the model from somewhere? GDrive usua…

The README also says this: > This fine-tunes the [GPT-J 6B]( https://huggingface.co/EleutherAI/gpt-j-6B ) model on the [Alpaca]( https://huggingface.co/datasets/tatsu-lab/alpaca ) dataset using a Databricks notebook. > Please note that while GPT-J 6B is Apache 2.0 licensed, the Alpaca dataset is licensed under Creative Commons NonCommercial (CC BY-NC 4.0). ...so, this cannot be used for commercial purposes

> ...so, this cannot be used for commercial purposes

or you can raise $30,000,000 right now and worry about the copyright infringement lawsuit in 2026 or never.

Re: Hello Dolly: Democratizing the magic of ChatGPT with open models

#117
post #15

This is the real risk to OpenAI's business model. If it turns out that you can get most of the same outcome with drastically smaller and cheaper models, then OpenAI is going to have a hell of a time keeping customers around as it will just be a race to the bottom on price and bigger, more expensive models will lose just from a hardware cost standpoint.

On the flip-side, OpenAI is primed to destroy their competitors. Partnership with Microsoft means they can buy Azure compute at-cost if need be. Their current portfolio of models is diverse on the expensive and cheap ends of the spectrum, with thousands of people on Twitter and HN still giving them lip-service. With dozens of clones hitting the market, OpenAI is the only one staying consistently relevant. The widespr…

The difference between this and SaaS is that businesses have been moving their (end user) products to SaaS due to wider broadband availability, as well as greed (read: MRR), but on the LLM side, people are building new products with it, so the incentives are to keep your costs low (or free) so you can make more money once you release.

Re: Hello Dolly: Democratizing the magic of ChatGPT with open models

#119
post #105
post #103

Does anyone else find it ironic that all these ChatGPT "clones" are popping up when OpenAi is supposed to be the ones open sourcing and sharing their work? I guess: "You Either Die A Hero, Or You Live Long Enough To See Yourself Become The Villain"?

Sam Altman has turned into a megalomaniac.

[deleted]

Re: Hello Dolly: Democratizing the magic of ChatGPT with open models

#120
post #4

This is really great news and something I felt was missing from the market so far. It seems everyone wants to create `moats` or walled-gardens with some aspect of their models etc. Nice job DataBricks, nice numbers too. Looking forward to more improvements.

Thought the same until I read this: > Contact us at hello-dolly@databricks.com if you would like to get access to the trained weights.

https://github.com/databrickslabs/dolly it’s now available on GitHub
Post reply on HN