Live data from Hacker News

The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

thesequence.substack.com

41–50 of 527 posts

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#41
post #7

Earlier quoted context omitted.

> It is convinced that it is always factually accurate, even though it is not. I don't think that's true. ChatGPT (or any LLM) isn't convinced much of anything. It might present something confidently (which is what most people want) but that's a side-effect of it's programming, not an indication of how good it feels on the answer. If you reply to anything ChatGPT says with "No, you're wrong." it will try to write a n…

>Everything it reads is mapped into language, not concept space Umm I'm pretty sure it's discovered concepts through compressing text - it seems perfectly capable of generalizing concepts

Text compression isn't a deterministic process, unfortunately. It's "concept" of compression is clearly derived from token sampling, in the same way it's concept of "math" is based on guessing the number/token that comes next.

While I do agree that ChatGPT exhibits pattern-recognizing qualities, that's basically what it was built to do. I'm not arguing against emergent properties, just against emergent intelligence or even the idea of "understanding" in the first place.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#42

> OpenAI published a detailed blog post outlining some of the principles used to ensure safety in their models. The post emphasize in areas such as privacy, factual accuracy Am I the only one amused by the phrase “factual accuracy”? How many stories have we read like the one where it tries to ghost light the guy that this year is actually last year. “Oh, your phone must be wrong too, because there is no way I could b…

The models are a lot of fun to play with, but yeah, every time I've tried to use them for something "serious" they nearly always invent stuff (and are so convincing in how they write about it!). Most recently I've been interested in what's happened with the 4-color theorem since the 1976 computer-assisted proof, and decided to use GPTChat instead of google+wikipedia. GPTChat had me convinced and excited that, apparen…

Before the inevitable idiots come in to say hurr durr but have you tried ChatGPT 4… yes I paid for it, and it is just as prone to hallucinations of factual information. It loves to make up new names for peoples initials.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#43

> OpenAI published a detailed blog post outlining some of the principles used to ensure safety in their models. The post emphasize in areas such as privacy, factual accuracy Am I the only one amused by the phrase “factual accuracy”? How many stories have we read like the one where it tries to ghost light the guy that this year is actually last year. “Oh, your phone must be wrong too, because there is no way I could b…

I hold a more charitable interpretation.

We (the public) have found an important bug in the system, ie. GPT can lie (or "hallucinate"), even if you try to convince it not to lie. The bug is definitely lowering the usefulness of their product, as well as the public option about it. But I'll let the programmer who has never coded a bug cast the first stone.

I wouldn't be surprised if they're scrambling internally to minimize the problem (in the product, not in public perception). They have also recently added a note to ChatGPT: "ChatGPT may produce inaccurate information about people, places, or facts" which is an acknowledement that yes, watch out (I compare it to "caution: contents hot" labels).

On the topic of dealing with it, I like the stance that simonw recently took: "We need to tell people ChatGPT will lie to them, not debate linguistics" [0].

I don't attach intentions to a machine algorithm (to me, "gaslight" definitely implies an evil intent), and I don't think OpenAI people are evil, stupid, corrupted or something else because they put out a product that has a bug. But since the wide public can't handle nuances, I'd agree it's better to say "chatgpt lies, use it for things where it either doesn't matter or you can verify; don't use it for fact-finding" to get the point across.

[0] https://simonwillison.net/2023/Apr/7/chatgpt-lies/

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#44
post #16

> OpenAI published a detailed blog post outlining some of the principles used to ensure safety in their models. The post emphasize in areas such as privacy, factual accuracy Am I the only one amused by the phrase “factual accuracy”? How many stories have we read like the one where it tries to ghost light the guy that this year is actually last year. “Oh, your phone must be wrong too, because there is no way I could b…

Well, unless they designed it to have zero confidence in itself, we are bound to have situations like this. When I was trying to troll it, by saying that IPCC just released a report stating that climate change is not real, and that they were completely wrong after all, it properly said that it is not very likely and that I'm probably mistaken. It admitted that it doesn't have internet access, but still refused to bel…

For better or worse in the current age of the internet prose is a good first pass filter for quality.

Someone arguing incoherently is seen as less believable.

Unfortunately the prose for these chat models doesn't change based on how certain it is of the facts. So you can't tell based on how it is talking whether it is true or not.

Certainly people online speak well while lying either intentionally or unintentionally but usually well intentioned people will coach things they aren't as certain about helping to paint a more accurate picture.

I haven't taken a deep dive on the latest models but historically most AI haven't worried about "facts" as much as associating speech patterns. It knows how to talk about facts because other people have done so in the past kind of thing.

This means you need to patch in arbitrary rules to reintroduce some semblance of truth to the outputs which isn't an easy task.

False training is a whole different area IMO. Especially when there is a difference between responding to a particular user and responding to everyone based on new information.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#45
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

I am shocked that it speaks the way it does when it was trained on random stuff it doesn’t have rights to.

They say they trained it on databases they had bought access to etc. And it seems that way.

Because how does ChatGPT:

1. Do what you ask instead of continuing your instructions?

2. Use such nice and helpful language as opposed to just random average of what people say?

3. And most of all — how does it have a structure where it helpfully restates things, summarizes things, warns you against doing dangerous stuff… no way is it just continuing the most probable random Internet text!!

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#46
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

Yeah, I'm particularly curious about that -- there's already legal precedent in the US that an AI cannot author copyrighted nor patented work. OpenAI can try to curtail it through a clickwrap agreement, but those are notoriously weak.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#47

Earlier quoted context omitted.

It is vastly better than anything else so far though. The rest will catch up but openai is not sleeping and they are well funded.

I thought that was the case before trying Vicuna. I agree that LLaMA and Alpaca are inferior to ChatGPT but I'm really not sure Vicuna is. It even (unfortunately) copies some of ChatGPT's quirks, like getting prudish when asking it to write a love scene ("It would not be appropriate for me to write...")

I've tried Vicuna but it still seems inferior to ChatGPT imo. Maybe if it was applied to a version of LLaMA with a number of parameters matching GPT-4 but I'm not sure of that either

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#48
post #45
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

I am shocked that it speaks the way it does when it was trained on random stuff it doesn’t have rights to. They say they trained it on databases they had bought access to etc. And it seems that way. Because how does ChatGPT: 1. Do what you ask instead of continuing your instructions? 2. Use such nice and helpful language as opposed to just random average of what people say? 3. And most of all — how does it have a str…

It is steered by RLHF to give helpful, nice, structured continuations. it was totally trained on random text they never paid a dime for.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#49
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

definetly. I don't think it's right when openai scraped data without consent from other resources. I feel that if openai can get data from the internet bard or someone else too can do it. Now being that chatgpt is also a part of the internet it's a fair game IMHO.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#50

Earlier quoted context omitted.

It is vastly better than anything else so far though. The rest will catch up but openai is not sleeping and they are well funded.

I thought that was the case before trying Vicuna. I agree that LLaMA and Alpaca are inferior to ChatGPT but I'm really not sure Vicuna is. It even (unfortunately) copies some of ChatGPT's quirks, like getting prudish when asking it to write a love scene ("It would not be appropriate for me to write...")

I admittedly have not interacted with Vicuna yet.
Post reply on HN