Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…
> Longer context window (1M+) What's your use case for this? Uploading multiple documents/books?
Large Enough
311–320 of 512 posts
Re: Large Enough
#312Sorry for the slightly off topic question, but can someone enlighten me which Claude model is more capable, Opus or Sonnet 3.5? I am confused because I see people fuzzing about Sonnet 3.5 being the best and yet somehow I seem to read again and again in factual texts and some benchmarks that Claude Opus is the most capable. Is there a simple answer to the question, what do I not understand? Please, thank you.
I.e. Opus is the largest and best model of each family but Sonnet is the first model of the 3.5 family and can beat 3's Opus in most tasks. When 3.5 Opus is released it will again outpace the 3.5 Sonnet model of the same family universally (in terms of capability) but until then it's a comparison of two different families without a universal guarantee, just a strong lean towards the newer model.
Re: Large Enough
#313Earlier quoted context omitted.
Testing models on their tokenization has always struck me as kinda odd. Like, that has nothing to do with their intelligence.
> Like, that has nothing to do with their intelligence. Because they don't have intelligence. If they did, they could count the letters in strawberry.
They fundamentally perceive the world in terms of tokens, not "letters".
Re: Large Enough
#314Earlier quoted context omitted.
I did it (fairly simple really) but found most of my (unsophisticated) coding these days to go through Aider [1] paired with Sonnet, for UX reasons mostly. It is easier to just prompt over the entire codebase, vs Cursor way of working with text selections. [1] https://aider.chat
Thanks for this suggestion. If anyone has other suggestions for working with large code context windows and changing code workflows, I would love to hear about them.
Re: Large Enough
#315Personally, language diversity should be the last thing on the list. If we had optimized every software from the get-go for a dozen languages our forward progress would have been dead in the water.
Language diversity means access to more training data, and you might also hope that by learning the same concept in multiple languages it does a better job of learning the underlying concept independent of the phrase structure... At least from a distance it seems like training a multilingual state of the art model might well be easier than a monolingual one.
Re: Large Enough
#316Earlier quoted context omitted.
Testing models on their tokenization has always struck me as kinda odd. Like, that has nothing to do with their intelligence.
I hear this a lot but there are vast sums of money thrown at where a model fails the strawberry cases. Think about math and logic. If a single symbol is off, it’s no good. Like a prompt where we can generate a single tokenization error at my work, by my very rough estimates, generates 2 man hours of work. (We search for incorrect model responses, get them to correct themselves, and if they can’t after trying, we tell…
In that case the tokenization is done at the appropriate level.
This is a complete non-issue for the use cases these models are designed for.
Re: Large Enough
#317Earlier quoted context omitted.
> Like, that has nothing to do with their intelligence. Because they don't have intelligence. If they did, they could count the letters in strawberry.
People have been over this. If you believe this, you don't understand how LLMs work. They fundamentally perceive the world in terms of tokens, not "letters".
Nor do they understand how intelligence works.
Humans don't read text a letter at a time. We're capable of deconstructing words into individual letters, but based on the evidence that's essentially a separate "algorithm".
Multi-model systems could certainly be designed to do that, but just like the human brain, it's unlikely to ever make sense for a text comprehension and generation model to work at the level of individual letters.
Re: Large Enough
#318Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…
Re: Large Enough
#319Earlier quoted context omitted.
> It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; I think you're just seeing the "make it work" stage of the combo "first make it work, then make it fast". Time to market is critical, as you can attest by the fact you framed the situation as "on par with GPT-4o and Claude Opus". You're seeing huge investments because being the first to get a working model stands to b…
ChatGPT is like Google now. It is the default. Even if Claude becomes as good as ChatGPT or even slightly better it won't make me switch. It has to be like a lot better. Way better. It feels like ChatGPT won the time to market war already.
Re: Large Enough
#320I'm building a ai coding assistant ( https://double.bot ) so I've tried pretty much all the frontier models. I added it this morning to play around with it and it's probably the worst model I've ever played with. Less coherent than 8B models. Worst case of benchmark hacking I've ever seen. example: https://x.com/WesleyYue/status/1816153964934750691