Live data from Hacker News

Mistral 3 family of models released

mistral.ai

241–243 of 243 posts

Re: Mistral 3 family of models released

#241
post #239

Earlier quoted context omitted.

Yeah - things are easy when you can objectively score an output, otherwise as you said you'll probably need another LLM to score it. For summaries you can try to make that somewhat more objective, like length and "8/10 key points are covered in this summary." This is a real training method (like Group Relative Policy Optimization), so it's a legitimate approach.

Thank you. I will google Group Relative Policy Optimization to learn about that and the other training methods. If you have any resources handy that I should be reading, that would be appreciated as well. Have a great weekend.

Nothing off the top of my head! If you find anything good let me know. GRPO is a training technique likely not exactly what you'd do for benchmarking, but it's interesting to read about anyway. Glad I cuold help

Re: Mistral 3 family of models released

#242

Earlier quoted context omitted.

Thanks for sharing your use case of the mistral models, which are indeed top-notch ! I had a look at phrasing.app, and while a nice website, I found the copy of "Hand-crafted. Phrasing was designed & developed by humans, for humans." somewhat of a false virtue given your statements here of advanced lllm usage.

I don't see the contention. I do not use llms in the design, development, copywriting, marketing, blogging, or any other aspect of the crafting of the application. I labor over every word, every button, every line of code, every blog post. I would say it is as hand-crafted as something digital can be.

It's interesting. I've been tinkering with an article summarizing/highlighting browser extension, and realized that I don't want the end-user to have read AI-generated content because it's not as high-quality as I'd hoped. But on the flip side, I'm loving having the AI write most of the code for me.

Re: Mistral 3 family of models released

#243
post #110

Earlier quoted context omitted.

Yep, Gemini is my least favorite and I’m convinced that the hype around it isn’t organic because I don’t see the claimed “superiority”, quite the opposite.

I think a lot of the hype around Gemini comes down to people who aren't using it for coding but for other things maybe. Frankly, I don't actually care about or want "general intelligence" -- I want it to make good code, follow instructions, and find bugs. Gemini wasn't bad at the last bit, but wasn't great at the others. They're all trying to make general purpose AI, but I just want really smart augmentation / tools.

I exclusively use Gemini Pro for coding, and it's been writing ~100% of the code I produce since July.

It's great.

Post reply on HN