Earlier quoted context omitted.
Hell, I would consider myself graced that simonw, yes, THAT simonw, the LLM whisperer, took time out of his busy schedule to send me to a discussion I might have expressed interest in.
> send me to a discussion I might have expressed interest in No, no, remember? Points to the blog you were already reading! Working diligently to build a brand: podcast, paid newsletter, the works.
Mistral releases Devstral2 and Mistral Vibe CLI
271–280 of 363 posts
Re: Mistral releases Devstral2 and Mistral Vibe CLI
#272Earlier quoted context omitted.
> But where are the professional tools, meant to be used for people who don't want to do vibe-coding, but be heavily assisted by LLMs? Something that is meant to augment the human intellect, not replace it? Claude Code not good enough for ya?
Claude Code has absolutely zero features that help me review code or do anything else than vibe-coding and accept changes as they come in. We need diff-comparisons between different executions, tailored TUI for that kind of work and more. Claude Code is basically a MVP of that. Still, I do use Claude Code and Codex daily as there is nothing better out there currently. But they still feel tailored towards vibe-coding…
Re: Mistral releases Devstral2 and Mistral Vibe CLI
#273Earlier quoted context omitted.
The fact that pelicans can't ride bicycles is pretty much the point of the benchmark! Asking an LLM to draw something that's physically impossible means it can't just "get it right" - seeing how different models (especially at different sizes) handle the problem is surprisingly interesting. Honestly though, the benchmark was originally meant to be a stupid joke. I only started taking it slightly more seriously about…
> If a model draws a really good picture of a pelican riding a bicycle there's a solid chance it will be great at all sorts of other things. Why? If I hired a worker that was really good at drawing pelicans riding a bike, it wouldn't tell me anything about his/her other qualities?!
It's not a human intelligence - it's a totally different thing, so why would the same test that you use to evaluate human abilities apply here?
Also more directly the "all sorts of other things" we want llms to be good at often involve writing code/spatial reasoning/world understanding which creating an svg of a pelican riding a bicycle very very directly evaluates so it's not even that surprising?
Re: Mistral releases Devstral2 and Mistral Vibe CLI
#274Earlier quoted context omitted.
It's not nessessarily the best benchmark, it's a popular one, probably because it's funny. Yes it's like the wine glass thing. Also it's kind of got depth. Does it draw the pelican and the bicycle? Can the penguin reach the peddles? How? I can imagine a really good AI finding a funny or creative or realistic way for the penguin to reach the peddles. An slightly worse AI will do an OK job, maybe just making the bike s…
> It's not nessessarily the best benchmark, it's a popular one, probably because it's funny. > Yes it's like the wine glass thing. No, it's not! That's part of my point; the wine glass scenario is a _realistic_ scenario. The pelican riding a bike is not. It's a _huge_ difference. Why should we measure intelligence (...) in regards to something that is realistic and something that is unrealistic? I just don't get it.
Re: Mistral releases Devstral2 and Mistral Vibe CLI
#275Earlier quoted context omitted.
No this is comparable to Deepseek-v3.2 even on their highlight task, with significantly worse general ability. And it's priced 5x of that.
It's open source; the price is up to the provider, and I do not see any on openrouter yet. ̶G̶i̶v̶e̶n̶ ̶t̶h̶a̶t̶ ̶d̶e̶v̶s̶t̶r̶a̶l̶ ̶i̶s̶ ̶m̶u̶c̶h̶ ̶s̶m̶a̶l̶l̶e̶r̶,̶ ̶I̶ ̶c̶a̶n̶ ̶n̶o̶t̶ ̶i̶m̶a̶g̶i̶n̶e̶ ̶i̶t̶ ̶w̶i̶l̶l̶ ̶b̶e̶ ̶m̶o̶r̶e̶ ̶e̶x̶p̶e̶n̶s̶i̶v̶e̶,̶ ̶l̶e̶t̶ ̶a̶l̶o̶n̶e̶ ̶5̶x̶.̶ ̶I̶f̶ ̶a̶n̶y̶t̶h̶i̶n̶g̶ ̶D̶e̶e̶p̶S̶e̶e̶k̶ ̶w̶i̶l̶l̶ ̶b̶e̶ ̶5̶x̶ ̶t̶h̶e̶ ̶c̶o̶s̶t̶.̶ edit: Mea culpa. I missed the active vs dense differe…
Re: Mistral releases Devstral2 and Mistral Vibe CLI
#276I'm sure I'm not the only one that thinks "Vibe CLI" sounds like an unserious tool. I use Claude Code a lot and little of it is what I would consider Vibe Coding.
Even the Gemini 3 announcement page had some bit like "best model for vibe coding".
Re: Mistral releases Devstral2 and Mistral Vibe CLI
#277Earlier quoted context omitted.
The point stands. Whether or not the standard is current has no relevance for the ability of the "AI" to produce the requested content. Either it can or can't.
https://news.ycombinator.com/item?id=46183673
Browsers are able to parse a webpage from 1996. I don't know what the argument in the linked comment is about, but in this one, we discuss the relevance of creating a 1996 page vs a pelican on a a bicycle in SVG.
Here is Gemini when asked how to build a webpage from 1996. Seems pretty correct. In general I dislike grand statements that are difficult to back up. In your case, if models have only a cursory knowledge of something (what does this mean in the context of LLMs anyway), what exactly they were trained on etc.
The shortened Gemini answer, the detailed version you can ask for yourself:
Layout via Tables: Without modern CSS, layouts were created using complex, nested HTML tables and invisible "spacer GIFs" to control white space.
Framesets: Windows were often split into independent sections (like a static sidebar and a scrolling content window) using Frames.
Inline Styling: Formatting was not centralized; fonts and colors were hard-coded individually on every element using the tag.
Low-Bandwidth Design: Visuals relied on tiny tiled background images, animated GIFs, and the limited "Web Safe" color palette.
CGI & Java: Backend processing was handled by Perl/CGI scripts, while advanced interactivity used slow-loading Java Applets.
Re: Mistral releases Devstral2 and Mistral Vibe CLI
#278Earlier quoted context omitted.
> I do not mind having a license like that, my gripe is with using the terms "permissive" and "open source" like that because such use dilutes them. I cannot think of any reason to do that aside from trying to dilute the term (especially when some laws, like the EU AI Act, are less restrictive when it comes to open source AIs specifically). Good. In this case, let it be diluted! These extra "restrictions" don't affec…
> I couldn't care less that the term is "diluted" and that makes it harder It also makes life harder for individuals and small companies, because this is not Open Source . It's incompatible with Open Source, it can't be reused in other Open Source projects. Terms have meanings. This is not Open Source, and it will never be Open Source.
I'm amazed at the social engineering that the megacorps have done with the whole Open Source (TM) thing. They engineered a whole generation of engineers to advocate not in their own self-interest, nor for the interest of the little people, but instead for the interest of the megacorps.
As soon as there is even the tiniest of restrictions, one which doesn't affect anyone besides a bunch of richiest corporations in the world, a bunch of people immediately come out of the woodwork, shout "but it's not open source!" and start bullying everyone else to change their language. Because if you even so much as inconvenience a megacorporation even a little bit it's not Open Source (TM) anymore.
If we're talking about ideals then this is something I find unsettling and dystopian.
I hard disagree with your "It also makes life harder for individuals and small companies" statement. It's the opposite. It gives them a competitive advantage vs megacorps, however small it may be.
Re: Mistral releases Devstral2 and Mistral Vibe CLI
#279Earlier quoted context omitted.
the question is, how do you want to provide instructions for what the AI is to do? You might not like calling it "chat" but somehow you need to communicate that, right? With aider you can write a comment for a function and then instruct it to finish the function inline (see other comments). But unless you just want pure autocomplete based on it guessing things, you need to provide guidance to it somehow.
I don't know exactly, but I guess in a more declarative manner rather than anything. Maybe we set goals/milestones/concrete objectives, or similar, rather than imperatively steer it, give it space to experiment yet make it very easy to understand exactly what important tradeoffs everything is doing. It's all very fluffy and theoretical of course.
Re: Mistral releases Devstral2 and Mistral Vibe CLI
#280Earlier quoted context omitted.
It's not nessessarily the best benchmark, it's a popular one, probably because it's funny. Yes it's like the wine glass thing. Also it's kind of got depth. Does it draw the pelican and the bicycle? Can the penguin reach the peddles? How? I can imagine a really good AI finding a funny or creative or realistic way for the penguin to reach the peddles. An slightly worse AI will do an OK job, maybe just making the bike s…
> It's not nessessarily the best benchmark, it's a popular one, probably because it's funny. > Yes it's like the wine glass thing. No, it's not! That's part of my point; the wine glass scenario is a _realistic_ scenario. The pelican riding a bike is not. It's a _huge_ difference. Why should we measure intelligence (...) in regards to something that is realistic and something that is unrealistic? I just don't get it.
It is unrealistic because if you go to a restaurant, you don't get served a glass like that. It is frowned upon (alcohol is a drug, after all) and impractical (wine stains are annoying) to fill a glass of wine as such.
A pelican riding a bike, on the other hand, is realistic in a scenario because of TV for children. Example from 1950's animation/comic involving a pelican [1].
[1] https://en.wikipedia.org/wiki/The_Adventures_of_Paddy_the_Pe...