Earlier quoted context omitted.
I don't get it. My husband also swears by Clause Sonnet 3.5, but every time I use it, the output is considerably worse than GPT-4o
Just don't listen to anecdata, and use objective metrics instead: https://chat.lmsys.org/?leaderboard
Large Enough
361–370 of 512 posts
Re: Large Enough
#362Earlier quoted context omitted.
Presuming the breakthrough is openly shared. It remains surprising how transparent many of these companies are about new approaches that push the SoTa forward, and I suspect we're going to see a change. That companies won't reveal the secret sauce so readily. e.g. Almost the entire market relies upon Attention Is All You Need paper detailing transformers, and it would be an entirely different market if Google had hel…
Given how absolutely pitiful the proprietary advancements in AI have been, I would posit we have little to worry about.
Given how good the model is terms of the quality vs speed tradeoff, they must have something.
Re: Large Enough
#363Earlier quoted context omitted.
If I ask an LLM to generate new words for some concept or category, it can do that. How do the new words form, if not from joining letters?
Not letters, but tokens. Think that it's translating everything to/from Chinese.
[1] Suggestion from chatgpt3.5 for new fruit name.
Re: Large Enough
#364Earlier quoted context omitted.
The problem is that the model never gets to see individual letters. The tokenizers used by these models break up the input in pieces. Even though the smallest pieces/units are bytes in most encodings (e.g. BBPE), the tokenizer will cut up most of the input in much larger units, because the vocabulary will contain fragments of words or even whole words. For example, if we tokenize Welcome to Hacker News, I hope you li…
The thing is, how the tokenizing work is about as relevant to the person asking the question as name of the cat of the delivery guy who delivered the GPU that the llm runs on.
Without looking at the word 'strawberry', or spelling it one letter at a time, can you rattle off how many letters are in the word off the top of your head? No? That is what we are asking the LLM to do.
Re: Large Enough
#365Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…
Claude needs to fix their text input box. It tries to be so advanced that code in backticks gets reformatted, and when you copy it, the formatting is lost (even the backticks).
I am curious what you mean by the formatting is lost though?
Re: Large Enough
#366Earlier quoted context omitted.
The thing is, how the tokenizing work is about as relevant to the person asking the question as name of the cat of the delivery guy who delivered the GPU that the llm runs on.
How the tokenizer works explains why a model can’t answer the question, what the name of the cat is doesn’t explain anything. This is Hacker News, we are usually interested in how things work.
To paraphrase how this thread started - it was someone testing different boats to see whether they can simply float - and they couldn’t. And the reply was questioning the validity of testing boats whether they can simply float.
At least this is how it sounds to me when I am told that our AI overlords can’t figure out how many Rs are in the word “strawberry”.
Re: Large Enough
#367These companies full of brilliant engineers are throwing millions of dollars in training costs to produce SOTA models that are... "on par with GPT-4o and Claude Opus"? And then the next 2.23% bump will cost another XX million? It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; that an ARC prize level breakthrough is needed to move the needle any farther at this point.
Re: Large Enough
#368Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…
All 3 models you ranked cannot get "how many r's are in strawberry?" correct. They all claim 2 r's unless you press them. With all the training data I'm surprised none of them fixed this yet.
What happens if you ask the total number of occurrences of the letter r in the word? Does it still not get it right?
Re: Large Enough
#369Earlier quoted context omitted.
How is a layman supposed to even know that it's testing on that? All they know is it's a large language model. It's not unreasonable they should expect it to be good at things having to do with language, like how many letters are in a word. Seems to me like a legit question for a young child to answer or even ask.
> How is a layman supposed to even know that it's testing on that? They're not, but laymen shouldn't think that the LLM tests they come up with have much value.
Re: Large Enough
#370Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…
All 3 models you ranked cannot get "how many r's are in strawberry?" correct. They all claim 2 r's unless you press them. With all the training data I'm surprised none of them fixed this yet.
How many letters R are in the word "s-t-r-a-w-b-e-r-r-y"?
The word "s-t-r-a-w-b-e-r-r-y" contains three instances of the letter "R."
How many letters R contain the word strawberry?
The word "strawberry" contains two instances of the letter "R."