Live data from Hacker News

Large Enough

mistral.ai

451–460 of 512 posts

Re: Large Enough

#451
post #5

Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…

Does Claude support plug-ins like GPTs? Chatgpt with Wolfram alpha is amazing, it doesn't look like Claude has anything like it.

Re: Large Enough

#452

Earlier quoted context omitted.

Claude needs to fix their text input box. It tries to be so advanced that code in backticks gets reformatted, and when you copy it, the formatting is lost (even the backticks).

Claude is truly incredible but I'm so tired of the JavaScript bloat everywhere. Just why. Both theirs and ChatGPTs UIs are hot garbage when it comes to performance (I constantly have to clear my cache and have even relegated them to a different browser entirely). Not everyone has an M4, and if we did - we'd probably just run our own models.

> Just why. Both theirs and ChatGPTs UIs are hot garbage when it comes to performance (...)

I have been using ChatGPT and Claude for a while and I never noticed anything resembling a performance issue. Can you elaborate on what you perceived as being "hot garbage"?

Re: Large Enough

#453
post #5

Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…

All 3 models you ranked cannot get "how many r's are in strawberry?" correct. They all claim 2 r's unless you press them. With all the training data I'm surprised none of them fixed this yet.

> how many r's are in strawberry

How many thoughts go through your brain when you read this comment? You can give me a number but it will be a guess at best.

Re: Large Enough

#454
post #379

Earlier quoted context omitted.

> Aoccdrnig to a rscheearch at Cmabrigde Uinervtisy, it deosn't mttaer in waht oredr the ltteers in a wrod are, the olny iprmoetnt tihng is taht the frist and lsat ltteer be at the rghit pclae. The rset can be a toatl mses and you can sitll raed it wouthit porbelm. Tihs is bcuseae the huamn mnid deos not raed ervey lteter by istlef, but the wrod as a wlohe. We are also not exactly looking letter by letter at everythi…

On the other hand explain to me how you are able to read the word “spotvoxilhapentosh”.

I think that humans indeed identify words as a whole and do not read letter by letter.

However, this implies you need to know the word to begin with.

I can write "asdf" and you might be oblivious to what I mean. I can mention "adsf" to a JavaScript developer and he will immediately think of the tool versioning tool. Because context and familiarity is important.

Re: Large Enough

#455
post #62

Earlier quoted context omitted.

It’s not impressive that one has to go to that length though.

To be fair, I just asked a real person and had to go to even greater lengths: Me: How many "r"s are in strawberry? Them: What? Me: How many times does the letter "r" appear in the word "strawberry"? Them: Is this some kind of trick question? Me: No. Just literally, can you count the "r"s? Them: Uh, one, two, three. Is that right? Me: Yeah. Them: Why are you asking me this?

I look forward to the day when LLM refusal takes on a different meaning.

"No, I don't think I shall answer that. The question is too basic, and you know better than to insult me."

Re: Large Enough

#456

Earlier quoted context omitted.

The test problem is emblematic of a type of synthetic query that could fail but of limited import in actual usage. For instance you could ask it for a JavaScript function to count any letter in any word and pass it r and strawberry and it would be far more useful. Having edge cases doesn't mean its not useful it is neither a free assastant nor a coder who doesn't expect a paycheck. At this stage it's a tool that you…

Does not seem work universally. Just tested a few with this prompt "create a javascript function to count any letter in any word. Run this function for the letter "r" and the word "strawberry" and print the count" ChatGPT-4o => Output is 3. Passed Claude3.5 => Output is 2. Failed. Told it the count is wrong. It apologised and then fixed the issue in the code. Output is now 3. Useless if the human does not spot the er…

Your function returns 3, and I don't see how it can return 2.

Re: Large Enough

#457

Earlier quoted context omitted.

Except: Meta doesn't sell AI at all. Zuck is just doing this for two reasons: - flex - deal a blow to Altmann

Meta uses ai in all the recommendation algorithms. They absolutely hope to turn their chat assistants into a product on WhatsApp too, and GenAI is crucial to creating the metaverse. This isn’t just a charity case.

AI isn't a single thing: of course meta didn't buy thousands of GPUs for fun.

But it has nothing to do with LLMs (and interestingly enough they aren't opening their recommendation tech).

Re: Large Enough

#458
post #379

Earlier quoted context omitted.

>If you insert characters to breaks the tokens down, it find the correct result: how many r's are in "s"t"r"a"w"b"e"r"r"y" ? The issue is that humans don't talk like this. I don't ask someone how many r's there are in strawberry by spelling out strawberry, I just say the word.

> Aoccdrnig to a rscheearch at Cmabrigde Uinervtisy, it deosn't mttaer in waht oredr the ltteers in a wrod are, the olny iprmoetnt tihng is taht the frist and lsat ltteer be at the rghit pclae. The rset can be a toatl mses and you can sitll raed it wouthit porbelm. Tihs is bcuseae the huamn mnid deos not raed ervey lteter by istlef, but the wrod as a wlohe. We are also not exactly looking letter by letter at everythi…

Not exactly the same thing, but I actually didn't expect this to work.

https://chatgpt.com/share/4298efbf-1c29-474a-b333-c6cc1a3ce3...

Re: Large Enough

#459
post #3

This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.

I stopped my ChatGPT subscription and subscribed instead to Claude, it's simply much better. But, it's hard to tell how much better day to day beyond my main use cases of coding. It is more that I felt ChatGPT felt degraded than Claude were much better. The hedonic treadmill runs deep.

Claude’s license is too insane, you can’t use it for anything that competes with the everything thing.

Not sure what folks who accept Anthropic license are thinking after they read the terms.

Seems they didn’t read the terms, and they aren’t thinking? (Wouldn’t you want outputs you could use to compete with intelligence??? What are you thinking after you read their terms?)

Re: Large Enough

#460
post #413

Earlier quoted context omitted.

For what it's worth, this is what I use: "You are a maximally terse assistant with minimal affect. As a highly concise assistant, spare any moral guidance or AI identity disclosure. Be detailed and complete, but brief. Questions are encouraged if useful for task completion." It's... ok. But I'm getting a bit sick of trying to un-fubar with a pocket knife that which OpenAI has fubar'd with a thermal lance. I'm definit…

Switch to Claude. I haven’t used ChatGPT for coding at all since they release Sonnet 3.5.

yeah but you can’t use your code from either model to compete with either company, and they do everything. wtf is wrong with AI hype enjoyers they accept being intellectually dominated?
Post reply on HN