Live data from Hacker News

Large Enough

mistral.ai

251–260 of 512 posts

Re: Large Enough

#251

Earlier quoted context omitted.

This reminds me of when I had to supervise outsourced developers. I wanted to say "build a function that does X and returns Y". But instead I had to say "build a function that takes these inputs, loops over them and does A or B based on condition C, and then return Y by applying Z transformation" At that point it was easier to do it myself.

Exact instruction challenge https://www.youtube.com/watch?v=cDA3_5982h8

"What programming computers is really like."

EDIT: Although perhaps it's even more important when dealing with humans and contracts. Someone could deliberately interpret the words in a way that's to their advantage.

Re: Large Enough

#252

Earlier quoted context omitted.

Sonnet 3.5 to me still seems far ahead. Maybe not on the benchmarks, but in everyday life I am finding it renders the other models useless. Even still, this monthly progress across all companies is exciting to watch. Its very gratifying to see useful technology advance at this pace, it makes me excited to be alive.

Such a relief/contrast to the period between 2010 and 2020, when the top five Google, Apple, Facebook, Amazon, and Microsoft monopolized their own regions and refused to compete with any other player in new fields. Google : Search Facebook : social Apple : phones Amazon : shopping Microsoft : enterprise .. > Even still, this monthly progress across all companies is exciting to watch. Its very gratifying to see useful…

Google refused to compete with Apple in phones?

Microsoft also competes in search, phones

Microsoft, Amazon and Google compete in cloud too

Re: Large Enough

#253
post #11

Earlier quoted context omitted.

I stopped my ChatGPT subscription and subscribed instead to Claude, it's simply much better. But, it's hard to tell how much better day to day beyond my main use cases of coding. It is more that I felt ChatGPT felt degraded than Claude were much better. The hedonic treadmill runs deep.

Have you (or anyone) swapped on Cursor with Anthropic API Key? For coding assistant, it's on my to do list to try. Cursor needs some serious work on model selection clarity though so I keep putting off.

One big advantage Claude artifacts have is that they maintain conversation context, versus when I am working with Cursor I have to basically repeat a bunch of information for each prompt, there is no continuity between requests for code edits.

If Cursor fixed that, the user experience would become a lot better.

Re: Large Enough

#254

Earlier quoted context omitted.

All 3 models you ranked cannot get "how many r's are in strawberry?" correct. They all claim 2 r's unless you press them. With all the training data I'm surprised none of them fixed this yet.

4o will get the answer right on the first go if you ask it "Search the Internet to determine how many R's are in strawberry?" which I find fascinating

I didn't even need to do that. 4o got it right straight away with just:

"how many r's are in strawberry?"

The funny thing is, I replied, "Are you sure?" and got back, "I apologize for the mistake. There are actually two 'r's in the word strawberry."

Re: Large Enough

#255
post #5

Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…

All 3 models you ranked cannot get "how many r's are in strawberry?" correct. They all claim 2 r's unless you press them. With all the training data I'm surprised none of them fixed this yet.

sonate 3.5 thinks 2

Re: Large Enough

#256
post #78

Earlier quoted context omitted.

I sell widgets. I promise the incalculable power of widgets has yet to be unleashed on the world, but it is tremendous and awesome and we should all be very afraid of widgets taking over the world because I can't see how they won't. Anyway here's the sales page. the widget subscription is so premium you won't even miss the subscription fee.

Except: Meta doesn't sell AI at all. Zuck is just doing this for two reasons: - flex - deal a blow to Altmann

[deleted]

Re: Large Enough

#257
post #68

Earlier quoted context omitted.

E = T/A! [0] A faster evolving approach to AI is coming out this year that will smoke anyone who still uses the term "license" in regards to ideas [1]. [0] https://breckyunits.com/eta.html [1] https://breckyunits.com/freedom.html

So it's made up?

I do what I say and I say what I do.

https://github.com/breck7/breckyunits.com/blob/afe70ad66cfbb...

Re: Large Enough

#258
post #78

Earlier quoted context omitted.

I sell widgets. I promise the incalculable power of widgets has yet to be unleashed on the world, but it is tremendous and awesome and we should all be very afraid of widgets taking over the world because I can't see how they won't. Anyway here's the sales page. the widget subscription is so premium you won't even miss the subscription fee.

Except: Meta doesn't sell AI at all. Zuck is just doing this for two reasons: - flex - deal a blow to Altmann

Meta uses ai in all the recommendation algorithms. They absolutely hope to turn their chat assistants into a product on WhatsApp too, and GenAI is crucial to creating the metaverse. This isn’t just a charity case.

Re: Large Enough

#259
post #198
post #3

This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.

I don't get it. My husband also swears by Clause Sonnet 3.5, but every time I use it, the output is considerably worse than GPT-4o

I don't see how that's possible. I decided to give GPT-4o a second chance after reaching my daily use on Sonnet 3.5, after 10 prompts GTP-4o failed to give me what Claude did in a single prompt (game-related programming). And with fragments and projects on top of that, the UX is miles ahead of anything OpenAI offers right now.

Re: Large Enough

#260
I like Claude 3.5 Sonnet, but despite paying for a plan, I run out of tokens after about 10 minutes. Text only, I'm typing everything in myself.

It's almost useless because I literally can't use it.

Update: https://support.anthropic.com/en/articles/8325612-does-claud...

45 messages per 5 hours is the limit for Pro users, less if Claude is wordy in its responses—which it always is. I hit that limit so fast when I'm investigating something. So annoying.

They used to let you select another, worse model but I don't see that option anymore. le sigh

Post reply on HN