This doesn't really explain what "reasoning" means in the context of genAI, or how it's done by this product. Are there any good sources to learn more about what "reasoning model" means outside of marketing-speak?
It's pure marketing. See the recent paper by Apple called "The Illusion of Thinking". https://ml-site.cdn-apple.com/papers/the-illusion-of-thinkin...
Magistral — the first reasoning model by Mistral AI
251–260 of 444 posts
Re: Magistral — the first reasoning model by Mistral AI
#252Earlier quoted context omitted.
Honestly the US approach to AI is incredibly irresponsible. As an American, I'm glad that someone somewhere is thinking about regulation. Not sure it will be enough though: https://xcancel.com/ESYudkowsky/status/1922710969785917691#m
There's nothing the regulation could meaningfully hope to accomplish other than slow down people willing to play by the rules.
Re: Magistral — the first reasoning model by Mistral AI
#253Earlier quoted context omitted.
Europe isn't going to catch up in tech as long as its market is open to US tech giants. Tech doesn't have marginal costs, so you want to have one of it in one place and sell it everywhere and when the infra and talent is already in US, EU tech is destined to do niche products. UK has a bit of it, France has some and that's it. The only viable alternatives are countries who have issues with US and that is China and Ru…
The problem is, CONSUMER level tech The EU is doing a lot of enterprise level shit and it's great The biggest company in Europe sells B2B software (SAP)
Re: Magistral — the first reasoning model by Mistral AI
#254Earlier quoted context omitted.
With how amazing the first R1 model was and how little compute they needed to create it, I'm really wondering how the new R1 model isn't beating o3 and 2.5 Pro on every single benchmark. Magistral Small is only 24B and scores 70.7% on AIME2024 while the 32B distill of R1 scores 72.6%. And with majority voting @64 the Magistral Small manages 83.3%, which is better than the full R1. Since I can run a 24B model on a reg…
It's because DeepSeek was a fast copy. That was the easy part and it's why they didn't have to use so much compute to get near the top. Going well beyond o3 or 2.5 Pro is drastically more expensive than fast copy. China's cultural approach to building substantial things produces this sort of outcome regularly, you see the same approach in automobiles, planes, Internet services, industrial machinery, military, et al.…
> That was the easy part
Is a bit hand-wavy in that it doesn't explain why it's only DeepSeek who can do this "easy" thing, but still not Meta, Mistral or anyone else really. There are many other players who have way more compute than DeepSeek (even inside China, not even considering rest of the world), and I can assure you more or less everyone trains on synthetic data/distillation from whatever bigger model they can access.
Re: Magistral — the first reasoning model by Mistral AI
#255Earlier quoted context omitted.
Jm2c but I feel conflicted about this arms race. You can be 6/12 months later, and have not burned tens of billions compared to the best in class, I see it an engineering win. I absolutely understand those that say "yeah, but customers will only use the best", I see it, but is market share of forever money losing businesses that valuable?
Indeed, and with the technology plateau-ing, being 6-12 months late with less debt is just long term thinking. Also, Europe being in the race is a big deal for consumers.
Re: Magistral — the first reasoning model by Mistral AI
#256Earlier quoted context omitted.
If so we need to fix the benchmarks.
I think there's a fundamental limit to benchmarks when it comes to real-world utility. The best option would be more like a user survey.
Re: Magistral — the first reasoning model by Mistral AI
#257Earlier quoted context omitted.
Money. No, really - EU doesn't have the VCs and the megacorps. People laugh at EU sponsoring projects, but there is no private money to sponsor them. There are plenty of US companies with sites in the EU though, so you have people working the problems, but no branding.
Ok, just a quick question… why does Europe not have the money actual/people?
Re: Magistral — the first reasoning model by Mistral AI
#258Earlier quoted context omitted.
Are we sure more time butt in office equates to more productivity?
$89,000 GDP per capita vs $46,000 rather proves the point about productivity per butt. US office workers are extraordinarily productive in terms of what their work generates (thanks to numerous well understood things like the outsized US scaling abilities). Measuring beyond that is very difficult due to the variance of every business.
Re: Magistral — the first reasoning model by Mistral AI
#259As a quick test of logical reasoning and basic Wikipedia-level knowledge, I asked Mistral AI the following question: A Brazilian citizen is flying from Sao Paulo to Paris, with a connection in Lisbon. Does he need to clear immigration in Lisbon or in Paris or in both cities or in neither city? Mistral AI said that "immigration control will only be cleared in Paris," which I think is wrong. After I pointed it to the W…
doing some reason.. uhh intuitioning i imagine brazil and portugal might have some sort of a visa-free deal going on in which case llama 4 might actually be right here?
Re: Magistral — the first reasoning model by Mistral AI
#260Benchmarks suggest this model loses to Deepseek-R1 in every one-shot comparison. Considering they were likely not even pitting it against the newer R1 version (no mention of that in the article) and at more than double the cost, this looks like the best AI company in the EU is struggling to keep up with the state-of-the-art.