This is really great news and something I felt was missing from the market so far. It seems everyone wants to create `moats` or walled-gardens with some aspect of their models etc. Nice job DataBricks, nice numbers too. Looking forward to more improvements.
Thought the same until I read this: > Contact us at hello-dolly@databricks.com if you would like to get access to the trained weights.
Hello Dolly: Democratizing the magic of ChatGPT with open models
101–110 of 194 posts
Re: Hello Dolly: Democratizing the magic of ChatGPT with open models
#102Earlier quoted context omitted.
The README also says this: > This fine-tunes the [GPT-J 6B]( https://huggingface.co/EleutherAI/gpt-j-6B ) model on the [Alpaca]( https://huggingface.co/datasets/tatsu-lab/alpaca ) dataset using a Databricks notebook. > Please note that while GPT-J 6B is Apache 2.0 licensed, the Alpaca dataset is licensed under Creative Commons NonCommercial (CC BY-NC 4.0). ...so, this cannot be used for commercial purposes
Essentially every model worth anything has been trained on a unfathomably large amount of data under copyright, with every possible licensing scheme you could imagine, under the assumption that it is fair use. While you can argue that it's all built on a house of cards (and a court may well agree with you) it's kind of arbitrary to draw a line here.
Re: Hello Dolly: Democratizing the magic of ChatGPT with open models
#103I guess: "You Either Die A Hero, Or You Live Long Enough To See Yourself Become The Villain"?
Re: Hello Dolly: Democratizing the magic of ChatGPT with open models
#104Earlier quoted context omitted.
I figured it was a reference to the Dalai Lama (which doesn't invalidate your comment, since that's also pronounced like Dalí). LLM -> Llama -> Dalai Lama
I thought "Dalai" pronounced "Dall Eye" rhymes with "Shall I" "Dali" pronounced "Dahl eee" rhymes with "Carly"
Re: Hello Dolly: Democratizing the magic of ChatGPT with open models
#105Does anyone else find it ironic that all these ChatGPT "clones" are popping up when OpenAi is supposed to be the ones open sourcing and sharing their work? I guess: "You Either Die A Hero, Or You Live Long Enough To See Yourself Become The Villain"?
Re: Hello Dolly: Democratizing the magic of ChatGPT with open models
#106Does anyone else find it ironic that all these ChatGPT "clones" are popping up when OpenAi is supposed to be the ones open sourcing and sharing their work? I guess: "You Either Die A Hero, Or You Live Long Enough To See Yourself Become The Villain"?
OpenAI renounced being open source. Don't let the name fool you.
Re: Hello Dolly: Democratizing the magic of ChatGPT with open models
#107I’d like some clarification of terms - when they say it takes 3 hours to train, they’re not saying from scratch are they? There’s already a huge amount of training to get to that point, isn’t that correct? If so, then it’s pretty audacious to claim they’ve democratized an LLM because the original training likely cost an epic amount of money. Then who knows how much guidance their training has incorporated, and it cou…
The 3 hours is the instruction fine-tuning. The base foundational model is GPT-J which was already provided by Eleuther-AI and has been around for a couple of years. Note: I work at Databricks and am familiar with this project but didn't work on it.
Re: Hello Dolly: Democratizing the magic of ChatGPT with open models
#108Re: Hello Dolly: Democratizing the magic of ChatGPT with open models
#109It's immediately become difficult to untangle the licensing here. Is this safe for production use - I have no idea if I can expect a DMCA from Mark if I step out of bounds with this or other post-Alpaca models, unless I'm missing something important. Meta really botched the Llama release.
Yes it's nuanced, but will be simplified going forward. This uses a fully open source (liberally licensed) model and we also open sourced (liberally licensed) our own training code. However, the uptraining dataset of ~50,000 samples was generated with OpenAI's text-davinci-003 model, and depending on how one interprets their terms, commercial use of the resulting model may violate the OpenAI terms of use. For that re…
Re: Hello Dolly: Democratizing the magic of ChatGPT with open models
#110Earlier quoted context omitted.
AFAIK DALL-E is pronounced as Dalí, as in Salvador Dalí. https://en.wikipedia.org/wiki/Salvador_Dal%C3%AD
It's quite clearly a reference to WALL-E the environmentally conscious robot, which is pronounced as you'd expect. I like to think of it as DALL-E the surrealist robot painter.