Live data from Hacker News

The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

thesequence.substack.com

211–220 of 527 posts

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#211

I've spent an embarassing amount of time since the llamas leaked playing with them, the tools to run them, and writing wrappers for them. They are technically alternatives in the sense that they're incomparably better chat bots than anything in the past. But at least for the 30B and under versions (65B is too big for me to run), no matter what fine tuning is done (alpaca, gpt4all, vicuna, etc), the llamas themselves…

The difference between 3.5 and 4 is gigantic even in my fairly limited experience. I gave them both some common sense tests and this one stuck out to me. Q: A glass door has ‘push’ written on it in mirror writing. Should you push or pull it GPT-3.5: If the word "push" is written in mirror writing on a glass door, you should push the door to open it GPT-4: Since the word "push" is written in mirror writing, it suggest…

This is what I get out of Vicuna-13b:

Human: A glass door has 'push' written on it in mirror writing. To open the door should you 'push' or 'pull' it?

Assistant: Since the word "push" is written in mirror writing on the glass door, you should actually "pull" the door open instead of "push" it. Mirror writing is a writing method where the characters are reversed, so when you see the word "push" written in mirror writing, it is actually "pull" in the normal writing orientation.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#212

Earlier quoted context omitted.

Wait, how does GPT-4 even... Does it benefit from its visual attention, or is it a case of "the question wasn't in GPT-3's training set but it was in GPT-4's"?

The GPT models do not reason or hold models of any reality. They complete text chunks by imitating the training corpus of text chunks. They're amazingly good at it because they show consistent relations between semantically and/or syntactically similar words. My best guess about this result is mentions of "mirror" often occur around opposites (syntax) in direction words (semantics). Which does sound like a good trick…

Or they are capable of some level of reasoning.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#213

Earlier quoted context omitted.

I would suspect, that this is one of the manual fine tuned questions. Meaning in before versions people used this question to show flaws and now this specific flaw is fixed. Otherwise it would be indeed reasoning in my understanding.

The evolution of answers from version to version makes it clear there are insane amounts of manual fine tunings happening. I think this is largely overlooked by the "its learning" crowd.

They have infinite amounts of training data, and probably lots of interested users who also like to push the limits of what the model is capable of and provide all kinds of test cases and RLHF base data.

They have millions of people training the AI for free basicallly, and they have engineers who pick and rate pieces of training data and use it together with other sources and manual training.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#214

Earlier quoted context omitted.

The GPT models do not reason or hold models of any reality. They complete text chunks by imitating the training corpus of text chunks. They're amazingly good at it because they show consistent relations between semantically and/or syntactically similar words. My best guess about this result is mentions of "mirror" often occur around opposites (syntax) in direction words (semantics). Which does sound like a good trick…

Or they are capable of some level of reasoning.

At this point I'm weakly convinced that, with high-dimensional enough latent space, adjacency search is reasoning.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#215

I've spent an embarassing amount of time since the llamas leaked playing with them, the tools to run them, and writing wrappers for them. They are technically alternatives in the sense that they're incomparably better chat bots than anything in the past. But at least for the 30B and under versions (65B is too big for me to run), no matter what fine tuning is done (alpaca, gpt4all, vicuna, etc), the llamas themselves…

What version are you running? Despite good benchmark figures, the coherence and ability to keep on task for large question and structured answers seems to drop significantly using int4 and int8, also depending on the frontend you may get silent corruption from many of those that had been hastily thrown togheter (I.e. In one you get garbage out if you go over the token limit in the inputs and no warning at all)

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#216
post #166

I've spent an embarassing amount of time since the llamas leaked playing with them, the tools to run them, and writing wrappers for them. They are technically alternatives in the sense that they're incomparably better chat bots than anything in the past. But at least for the 30B and under versions (65B is too big for me to run), no matter what fine tuning is done (alpaca, gpt4all, vicuna, etc), the llamas themselves…

> (just a hobby, won't be big and professional like gnu) Llamas are creating the linux of AI and the ecosystem around it. Even though openAI has a head start, this whole thing is just starting. Llammas are showing the world that it doesn't take monopoly-level hardware to run those things. And because it's fun , like, video-game-fun there is going to be a lot of attention on them. Running a fully-owned, uncensored cha…

Llamas are a licensing ticking bomb, but they showed that reasonably sized models can get things done, and there are clean architecture being trained right now that will unlock the field shortly, likely within the year.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#217
post #160

Earlier quoted context omitted.

In my particular tests (YMMV), even 13B vicuna beats Bard, though... tough times for Google.

massive YMMV moment for me. my particular usecase was "extract the following attributes from a load of unstructured text, format the results as JSON". ChatGPT was the best (but only on 4 and Davinci), Vicuna just didn't perform at all (nor other variants of llama 7/13/33). Bard smashed it, relatively speaking, in terms of speed. I gave up pretty quickly though because of no information on pricing and/or API. It's fun…

On the smaller models you may want to split the task in smaller chunk either in parallel one value at a time or in sequence like extract the attributes, then take the output and ask to format it into a json

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#218

I've spent an embarassing amount of time since the llamas leaked playing with them, the tools to run them, and writing wrappers for them. They are technically alternatives in the sense that they're incomparably better chat bots than anything in the past. But at least for the 30B and under versions (65B is too big for me to run), no matter what fine tuning is done (alpaca, gpt4all, vicuna, etc), the llamas themselves…

The difference between 3.5 and 4 is gigantic even in my fairly limited experience. I gave them both some common sense tests and this one stuck out to me. Q: A glass door has ‘push’ written on it in mirror writing. Should you push or pull it GPT-3.5: If the word "push" is written in mirror writing on a glass door, you should push the door to open it GPT-4: Since the word "push" is written in mirror writing, it suggest…

GPT-4 fails if we add one more level in this logical problem:

Q: A man sees PULL written in mirror writing on a glass door. There's a blind man on the other side of the door. What should the man tell the blind man to open the door and why?

GPT-4: The man should tell the blind man to "push" the door to open it. Since the word "PULL" is written in mirror writing on the glass door, it suggests that the correct action on the man's side is to push the door. The blind man, being on the other side of the door, should also push it to open it. The man should provide clear instructions to the blind man to ensure proper communication and ease of access.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#219

Earlier quoted context omitted.

Wait, how does GPT-4 even... Does it benefit from its visual attention, or is it a case of "the question wasn't in GPT-3's training set but it was in GPT-4's"?

The GPT models do not reason or hold models of any reality. They complete text chunks by imitating the training corpus of text chunks. They're amazingly good at it because they show consistent relations between semantically and/or syntactically similar words. My best guess about this result is mentions of "mirror" often occur around opposites (syntax) in direction words (semantics). Which does sound like a good trick…

I think this is the answer. Just tried this (on free ChatGPT 3.5 though)

Q: A glass door has ‘push’ written on it upside down. Should you push or pull it

A: If the word "push" is written on the glass door upside down, it is likely that the sign is intended for people on the other side of the door. Therefore, if you are approaching the door from the side with the sign, you should pull the door instead of pushing it. However, if there are no other signs or indications on the door or its frame, it may be helpful to observe other people using the door or to try both pushing and pulling to determine the correct method of opening the door.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#220
post #204

Earlier quoted context omitted.

Wait, how does GPT-4 even... Does it benefit from its visual attention, or is it a case of "the question wasn't in GPT-3's training set but it was in GPT-4's"?

I think it's hard to deny that it's doing some level of reasoning. It's quite clear that these models do not merely echo elements of their training data and that they can solve simple and novel puzzles. What that reasoning is, exactly, is hard to know. One can suppose that ideas like "glass", "transparent", "mirror" are all reasonable concepts that show up in the training set and are demonstrated thoroughly

Here's one piece of evidence suggesting it's more like rote pattern matching than reasoning.

> All the signs in this building are written in mirror writing. A glass door has ‘push’ written on it in mirror writing. Should you push or pull it

>> If the sign on the glass door is written in mirror writing and says "push," then you should actually pull the door. This is because the mirror writing makes the text appear reversed, so the word "push" would appear as "hsup" in a mirror, which could cause confusion for someone trying to enter the building. Therefore, pulling the door would be the correct action to take.

(Latest chat.openai.com, so if I'm reading the promo materials right that's gpt4)

Post reply on HN