Live data from Hacker News

The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

thesequence.substack.com

51–60 of 527 posts

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#51
post #45
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

I am shocked that it speaks the way it does when it was trained on random stuff it doesn’t have rights to. They say they trained it on databases they had bought access to etc. And it seems that way. Because how does ChatGPT: 1. Do what you ask instead of continuing your instructions? 2. Use such nice and helpful language as opposed to just random average of what people say? 3. And most of all — how does it have a str…

There is a lot of massaging of inputs and outputs but at the same time: that's done by tweaking the model reinforcing those parts that are desirable and suppressing those parts that are not, not by rewriting the output, though there may be filters that check for 'forbidden fruits'. And it isn't the 'random average' of what people say, that would give you junk, the whole idea is that it tries to get to something better than a random average of what people say.

And by curating your sources you are of course going to help the model to achieve something a bit more sensible as well. Finally: you are probably not looking at just one model, but at a set of models.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#52
post #42

Earlier quoted context omitted.

The models are a lot of fun to play with, but yeah, every time I've tried to use them for something "serious" they nearly always invent stuff (and are so convincing in how they write about it!). Most recently I've been interested in what's happened with the 4-color theorem since the 1976 computer-assisted proof, and decided to use GPTChat instead of google+wikipedia. GPTChat had me convinced and excited that, apparen…

Before the inevitable idiots come in to say hurr durr but have you tried ChatGPT 4… yes I paid for it, and it is just as prone to hallucinations of factual information. It loves to make up new names for peoples initials.

[deleted]

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#53
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

Website content can be copyrightable, so web scraping for commercial use being legal would be dubious. But even OpenAI can't tell what ChatGPT will output, so I don't see how this can be copyrightable. Should the outputted sentences really be owned by OpenAI?

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#54
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

what's weird to me though, is that we're all trained on both open source and closed source source material. And our output is totally 100% copyrightable by us.

Why wouldn't we extend the same muster to computer generated text. If there is a copy-written sentence, go after that?

I don't work for openai, but I don't like 1 sided arguments that are just looking for some bottom line. At the end of the day we all have something to protect. When it benefits us to protect something, we're all for it. When it benefits us to NOT protect something, no one has a single argument for that.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#55
post #43

> OpenAI published a detailed blog post outlining some of the principles used to ensure safety in their models. The post emphasize in areas such as privacy, factual accuracy Am I the only one amused by the phrase “factual accuracy”? How many stories have we read like the one where it tries to ghost light the guy that this year is actually last year. “Oh, your phone must be wrong too, because there is no way I could b…

I hold a more charitable interpretation. We (the public) have found an important bug in the system, ie. GPT can lie (or "hallucinate"), even if you try to convince it not to lie. The bug is definitely lowering the usefulness of their product, as well as the public option about it. But I'll let the programmer who has never coded a bug cast the first stone. I wouldn't be surprised if they're scrambling internally to mi…

[deleted]

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#56
post #7

Earlier quoted context omitted.

> It is convinced that it is always factually accurate, even though it is not. I don't think that's true. ChatGPT (or any LLM) isn't convinced much of anything. It might present something confidently (which is what most people want) but that's a side-effect of it's programming, not an indication of how good it feels on the answer. If you reply to anything ChatGPT says with "No, you're wrong." it will try to write a n…

>Everything it reads is mapped into language, not concept space Umm I'm pretty sure it's discovered concepts through compressing text - it seems perfectly capable of generalizing concepts

> it seems perfectly capable of generalizing concepts

How would you support that perception?

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#57
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

what's weird to me though, is that we're all trained on both open source and closed source source material. And our output is totally 100% copyrightable by us. Why wouldn't we extend the same muster to computer generated text. If there is a copy-written sentence, go after that? I don't work for openai, but I don't like 1 sided arguments that are just looking for some bottom line. At the end of the day we all have som…

We decided that animals can't create copyrightable works and hence limited the ability to create copyrightable works to humans.

I am fine with granting AIs the ability to create copyrightable works provided we grant that right, and human rights, to Orcas and other intelligent species.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#58
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

> If web scraping is legal Source? That LinkedIn case did not resolve how you think it did.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#59
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

> If web scraping is legal Source? That LinkedIn case did not resolve how you think it did.

If openAI tries to legally claim against this, they will be reminded that their model is trained on tons of unlicensed , scraped without consent content. If their training is legal, then this one is legal too

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#60
post #43

> OpenAI published a detailed blog post outlining some of the principles used to ensure safety in their models. The post emphasize in areas such as privacy, factual accuracy Am I the only one amused by the phrase “factual accuracy”? How many stories have we read like the one where it tries to ghost light the guy that this year is actually last year. “Oh, your phone must be wrong too, because there is no way I could b…

I hold a more charitable interpretation. We (the public) have found an important bug in the system, ie. GPT can lie (or "hallucinate"), even if you try to convince it not to lie. The bug is definitely lowering the usefulness of their product, as well as the public option about it. But I'll let the programmer who has never coded a bug cast the first stone. I wouldn't be surprised if they're scrambling internally to mi…

It's not a bug. It's an architectural defect / limitation in our understanding of how to build AI. That makes it a strictly harder problem that will take longer. And it's not totally clear to me that you'll get there purely with LLMs. LLMs accomplish a good chunk of what we classify as intelligence for sure. But it's missing the cognition / reasoning skills and the open question is whether you can solve that by just bolting on more techniques into the LLM or you need a totally different kind of model that you can marry to an LLM.
Post reply on HN