Live data from Hacker News

Mistral releases Devstral2 and Mistral Vibe CLI

mistral.ai

251–260 of 363 posts

Re: Mistral releases Devstral2 and Mistral Vibe CLI

#251

Look interesting, eager to play around with it! Devstral was a neat model when it released and one of the better ones to run locally for agentic coding. Nowadays I mostly use GPT-OSS-120b for this, so gonna be interesting to see if Devstral 2 can replace it. I'm a bit saddened by the name of the CLI tool, which to me implies the intended usage. "Vibe-coding" is a fun exercise to realize where models go wrong, but for…

High quality code is a thing from the past What matters is high quality specifications including test cases

"high quality specifications" have _always_ been a thing that matters.

In my mind, it's somewhat orthogonal to code quality.

Waterfall has always been about "high quality specifications" written by people who never see any code, much less write it. Agile make specs and code quality somewhat related, but in at least some ways probably drives lower quality code in the pursuit of meeting sprint deadlines and producing testable artefacts at the expense of thoroughness/correctness/quality.

Re: Mistral releases Devstral2 and Mistral Vibe CLI

#252

Earlier quoted context omitted.

yeahhhhhhh, that's not how this works. Unless this authority has some ownership over the term and can prevent its misuse (e.g. with lawsuits or similar), it is not actually the authority of the term, and people will continue to use it how they see fit. Indeed, I am not part of a movement (nor would I want to be) which focuses more on what words are used rather than what actions are taken.

> people will continue to use it how they see fit. People can also say 2+2=5, and they're wrong. And people will continue to call them out on it. And we will keep doing so, because stopping lets people move the Overton window and try to get away with even more.

2+2 is a mathematical concept. Definitions do not need to be agreed upon beyond fundamental axioms.

The same is not true for "open source", which is a purely linguistic construct.

Re: Mistral releases Devstral2 and Mistral Vibe CLI

#253
post #196

Earlier quoted context omitted.

> I've personally decided to just rent systems with GPUs from a cloud provider and setup SSH tunnels to my local system. That's a good idea! Curious about this, if you don't mind sharing: - what's the stack ? (Do you run like llama.cpp on that rented machine?) - what model(s) do you run there? - what's your rough monthly cost? (Does it come up much cheaper than if you called the equivalent paid APIs)

I ran ollama first because it was easy, but now download source and build llama.cpp on the machine. I don't bother saving a file system between runs on the rented machine, I build llama.cpp every time I start up. I am usually just running gpt-oss-120b or one of the qwen models. Sometimes gemma? These are mostly "medium" sized in terms of memory requirements - I'm usually trying unquantized models that will easily run…

I don't suppose you have (or would be interested in writing) a blog post about how you set that up? Or maybe a list of links/resources/prompts you used to learn how to get there?

Re: Mistral releases Devstral2 and Mistral Vibe CLI

#254
post #133

Earlier quoted context omitted.

TIL: https://garlicmodel.com/ That looks like the next flagship rather than the fast distillation, but thanks for sharing.

Lol, someone vibecoded an entire website for OpenAI's model, that's some dedication.

"GPT, please make me a website about OpenAI's 'Garlic' model."

Re: Mistral releases Devstral2 and Mistral Vibe CLI

#255

Earlier quoted context omitted.

We are getting to the point that its not unreasonable to think that "Generate an SVG of a pelican riding a bicycle" could be included in some training data. It would be a great way to ensure an initial thumbs up from a prominent reviewer. It's a good benchmark but it seems like it would be a good idea to include an additional random or unannounced similar test to catch any benchmaxxing.

It would be easy to out models that train on the bike pelican, because they would probably suck at the kayaking bumblebee. So far though, the models good at bike pelican are also good at kayak bumblebee, or whatever other strange combo you can come up with. So if they are trying to benchmaxx by making SVG generation stronger, that's not really a miss, is it?

That depends on if "SVG generation" is a particularly useful LLM/coding model skill outside of benchmarking. I.e., if they make that stronger with some params that otherwise may have been used for "rust type system awareness" or somesuch, it might be a net loss outside of the benchmarks.

Re: Mistral releases Devstral2 and Mistral Vibe CLI

#256

Earlier quoted context omitted.

People have been doing this for literally every anticipated model release, and I presume skimming some amount of legitimate interest since their sites end up being top indexed until the actual model is released. Google should be punishing these sites but presumably it's too narrow of a problem for them to care.

Black SEO in the age of LLMs

It would need outbound links to be SEO

Or at least a profit model. I don't see either on that page but maybe I'm missing something

Re: Mistral releases Devstral2 and Mistral Vibe CLI

#258
post #153

Earlier quoted context omitted.

It's open source; the price is up to the provider, and I do not see any on openrouter yet. ̶G̶i̶v̶e̶n̶ ̶t̶h̶a̶t̶ ̶d̶e̶v̶s̶t̶r̶a̶l̶ ̶i̶s̶ ̶m̶u̶c̶h̶ ̶s̶m̶a̶l̶l̶e̶r̶,̶ ̶I̶ ̶c̶a̶n̶ ̶n̶o̶t̶ ̶i̶m̶a̶g̶i̶n̶e̶ ̶i̶t̶ ̶w̶i̶l̶l̶ ̶b̶e̶ ̶m̶o̶r̶e̶ ̶e̶x̶p̶e̶n̶s̶i̶v̶e̶,̶ ̶l̶e̶t̶ ̶a̶l̶o̶n̶e̶ ̶5̶x̶.̶ ̶I̶f̶ ̶a̶n̶y̶t̶h̶i̶n̶g̶ ̶D̶e̶e̶p̶S̶e̶e̶k̶ ̶w̶i̶l̶l̶ ̶b̶e̶ ̶5̶x̶ ̶t̶h̶e̶ ̶c̶o̶s̶t̶.̶ edit: Mea culpa. I missed the active vs dense differe…

Deepseek v3.2 is that cheap because its attention mechanism is ridiculously efficient.

Yeah, DeepSeek Sparse Attention. Section 2: https://arxiv.org/abs/2512.02556

Re: Mistral releases Devstral2 and Mistral Vibe CLI

#259
post #233

Earlier quoted context omitted.

> We are getting to the point that its not unreasonable to think that "Generate an SVG of a pelican riding a bicycle" could be included in some training data. I may be stupid, but _why_ is this prompt used as a benchmark? I mean, pelicans _can't_ ride a bicycle, so why is it important for "AI" to show that they can (at least visually)? The "wine glass problem"[0] - and probably others - seems to me to be a lot more r…

It's not nessessarily the best benchmark, it's a popular one, probably because it's funny. Yes it's like the wine glass thing. Also it's kind of got depth. Does it draw the pelican and the bicycle? Can the penguin reach the peddles? How? I can imagine a really good AI finding a funny or creative or realistic way for the penguin to reach the peddles. An slightly worse AI will do an OK job, maybe just making the bike s…

> It's not nessessarily the best benchmark, it's a popular one, probably because it's funny.

> Yes it's like the wine glass thing.

No, it's not!

That's part of my point; the wine glass scenario is a _realistic_ scenario. The pelican riding a bike is not. It's a _huge_ difference. Why should we measure intelligence (...) in regards to something that is realistic and something that is unrealistic?

I just don't get it.

Re: Mistral releases Devstral2 and Mistral Vibe CLI

#260
post #232

Earlier quoted context omitted.

> We are getting to the point that its not unreasonable to think that "Generate an SVG of a pelican riding a bicycle" could be included in some training data. I may be stupid, but _why_ is this prompt used as a benchmark? I mean, pelicans _can't_ ride a bicycle, so why is it important for "AI" to show that they can (at least visually)? The "wine glass problem"[0] - and probably others - seems to me to be a lot more r…

The fact that pelicans can't ride bicycles is pretty much the point of the benchmark! Asking an LLM to draw something that's physically impossible means it can't just "get it right" - seeing how different models (especially at different sizes) handle the problem is surprisingly interesting. Honestly though, the benchmark was originally meant to be a stupid joke. I only started taking it slightly more seriously about…

> If a model draws a really good picture of a pelican riding a bicycle there's a solid chance it will be great at all sorts of other things.

Why?

If I hired a worker that was really good at drawing pelicans riding a bike, it wouldn't tell me anything about his/her other qualities?!

Post reply on HN