Live data from Hacker News

Large Enough

mistral.ai

1–10 of 512 posts

Re: Large Enough

#3
This race for the top model is getting wild. Everyone is claiming to one-up each with every version.

My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away.

I'm not really sure how to even test/use Mistral or Llama for everyday use though.

Re: Large Enough

#4
I love how much AI is bringing competition (and thus innovation) back to tech. Feels like things were stagnant for 5-6 years prior because of the FAANG stranglehold on the industry. Love also that some of this disruption is coming at out of France (HuggingFace and Mistral), which Americans love to typecast as incapable of this.

Re: Large Enough

#5
Links to chat with models that released this week:

Large 2 - https://chat.mistral.ai/chat

Llama 3.1 405b - https://www.llama2.ai/

I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history.

I'd rank as:

1. Sonnet 3.5

2. Large 2 and Llama 405b (similar, no clear winner between the two)

If you're using Claude, stick with it.

My Claude wishlist:

1. Smarter (yes, it's the most intelligent, and yes, I wish it was far smarter still)

2. Longer context window (1M+)

3. Native audio input including tone understanding

4. Fewer refusals and less moralizing when refusing

5. Faster

6. More tokens in output

Re: Large Enough

#6
post #3

This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.

Sonnet 3.5 to me still seems far ahead. Maybe not on the benchmarks, but in everyday life I am finding it renders the other models useless. Even still, this monthly progress across all companies is exciting to watch. Its very gratifying to see useful technology advance at this pace, it makes me excited to be alive.

Re: Large Enough

#7
post #3

This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.

I stopped my ChatGPT subscription and subscribed instead to Claude, it's simply much better. But, it's hard to tell how much better day to day beyond my main use cases of coding. It is more that I felt ChatGPT felt degraded than Claude were much better. The hedonic treadmill runs deep.

Re: Large Enough

#8
post #3

This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.

Sonnet 3.5 to me still seems far ahead. Maybe not on the benchmarks, but in everyday life I am finding it renders the other models useless. Even still, this monthly progress across all companies is exciting to watch. Its very gratifying to see useful technology advance at this pace, it makes me excited to be alive.

I’ve stopped using anything else as a coding assistant. It’s head and shoulders above GPT-4o on reasoning about code and correcting itself.

Re: Large Enough

#9
post #3

This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.

Agree on Claude. I also feel like ChatGPT has gotten noticeably worse over the last few months.

Re: Large Enough

#10
These companies full of brilliant engineers are throwing millions of dollars in training costs to produce SOTA models that are... "on par with GPT-4o and Claude Opus"? And then the next 2.23% bump will cost another XX million? It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; that an ARC prize level breakthrough is needed to move the needle any farther at this point.
Post reply on HN