Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…
Large Enough
451–460 of 512 posts
Re: Large Enough
#452Earlier quoted context omitted.
Claude needs to fix their text input box. It tries to be so advanced that code in backticks gets reformatted, and when you copy it, the formatting is lost (even the backticks).
Claude is truly incredible but I'm so tired of the JavaScript bloat everywhere. Just why. Both theirs and ChatGPTs UIs are hot garbage when it comes to performance (I constantly have to clear my cache and have even relegated them to a different browser entirely). Not everyone has an M4, and if we did - we'd probably just run our own models.
I have been using ChatGPT and Claude for a while and I never noticed anything resembling a performance issue. Can you elaborate on what you perceived as being "hot garbage"?
Re: Large Enough
#453Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…
All 3 models you ranked cannot get "how many r's are in strawberry?" correct. They all claim 2 r's unless you press them. With all the training data I'm surprised none of them fixed this yet.
How many thoughts go through your brain when you read this comment? You can give me a number but it will be a guess at best.
Re: Large Enough
#454Earlier quoted context omitted.
> Aoccdrnig to a rscheearch at Cmabrigde Uinervtisy, it deosn't mttaer in waht oredr the ltteers in a wrod are, the olny iprmoetnt tihng is taht the frist and lsat ltteer be at the rghit pclae. The rset can be a toatl mses and you can sitll raed it wouthit porbelm. Tihs is bcuseae the huamn mnid deos not raed ervey lteter by istlef, but the wrod as a wlohe. We are also not exactly looking letter by letter at everythi…
On the other hand explain to me how you are able to read the word “spotvoxilhapentosh”.
However, this implies you need to know the word to begin with.
I can write "asdf" and you might be oblivious to what I mean. I can mention "adsf" to a JavaScript developer and he will immediately think of the tool versioning tool. Because context and familiarity is important.
Re: Large Enough
#455Earlier quoted context omitted.
It’s not impressive that one has to go to that length though.
To be fair, I just asked a real person and had to go to even greater lengths: Me: How many "r"s are in strawberry? Them: What? Me: How many times does the letter "r" appear in the word "strawberry"? Them: Is this some kind of trick question? Me: No. Just literally, can you count the "r"s? Them: Uh, one, two, three. Is that right? Me: Yeah. Them: Why are you asking me this?
"No, I don't think I shall answer that. The question is too basic, and you know better than to insult me."
Re: Large Enough
#456Earlier quoted context omitted.
The test problem is emblematic of a type of synthetic query that could fail but of limited import in actual usage. For instance you could ask it for a JavaScript function to count any letter in any word and pass it r and strawberry and it would be far more useful. Having edge cases doesn't mean its not useful it is neither a free assastant nor a coder who doesn't expect a paycheck. At this stage it's a tool that you…
Does not seem work universally. Just tested a few with this prompt "create a javascript function to count any letter in any word. Run this function for the letter "r" and the word "strawberry" and print the count" ChatGPT-4o => Output is 3. Passed Claude3.5 => Output is 2. Failed. Told it the count is wrong. It apologised and then fixed the issue in the code. Output is now 3. Useless if the human does not spot the er…
Re: Large Enough
#457Earlier quoted context omitted.
Except: Meta doesn't sell AI at all. Zuck is just doing this for two reasons: - flex - deal a blow to Altmann
Meta uses ai in all the recommendation algorithms. They absolutely hope to turn their chat assistants into a product on WhatsApp too, and GenAI is crucial to creating the metaverse. This isn’t just a charity case.
But it has nothing to do with LLMs (and interestingly enough they aren't opening their recommendation tech).
Re: Large Enough
#458Earlier quoted context omitted.
>If you insert characters to breaks the tokens down, it find the correct result: how many r's are in "s"t"r"a"w"b"e"r"r"y" ? The issue is that humans don't talk like this. I don't ask someone how many r's there are in strawberry by spelling out strawberry, I just say the word.
> Aoccdrnig to a rscheearch at Cmabrigde Uinervtisy, it deosn't mttaer in waht oredr the ltteers in a wrod are, the olny iprmoetnt tihng is taht the frist and lsat ltteer be at the rghit pclae. The rset can be a toatl mses and you can sitll raed it wouthit porbelm. Tihs is bcuseae the huamn mnid deos not raed ervey lteter by istlef, but the wrod as a wlohe. We are also not exactly looking letter by letter at everythi…
https://chatgpt.com/share/4298efbf-1c29-474a-b333-c6cc1a3ce3...
Re: Large Enough
#459This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.
I stopped my ChatGPT subscription and subscribed instead to Claude, it's simply much better. But, it's hard to tell how much better day to day beyond my main use cases of coding. It is more that I felt ChatGPT felt degraded than Claude were much better. The hedonic treadmill runs deep.
Not sure what folks who accept Anthropic license are thinking after they read the terms.
Seems they didn’t read the terms, and they aren’t thinking? (Wouldn’t you want outputs you could use to compete with intelligence??? What are you thinking after you read their terms?)
Re: Large Enough
#460Earlier quoted context omitted.
For what it's worth, this is what I use: "You are a maximally terse assistant with minimal affect. As a highly concise assistant, spare any moral guidance or AI identity disclosure. Be detailed and complete, but brief. Questions are encouraged if useful for task completion." It's... ok. But I'm getting a bit sick of trying to un-fubar with a pocket knife that which OpenAI has fubar'd with a thermal lance. I'm definit…
Switch to Claude. I haven’t used ChatGPT for coding at all since they release Sonnet 3.5.