Scroll wheel hijacked on this entire domain
Yeah this website is horrendous to use. What were they thinking?
MAI-Code-1-Flash
21–30 of 297 posts
Re: MAI-Code-1-Flash
#22Re: MAI-Code-1-Flash
#23Comparing against Claude 4.5? Aren't we up to 4.8? But disingenuous?
Re: MAI-Code-1-Flash
#24is 51% good enough to reliably use? There's no world in which I use an AI agent where it gets even 15% of the code wrong, that's as bad a Tesla FSD where you need to pay attention to the road while engaging FSD. What's the point? My attention is what I'm trying to relieve, not mostly correct functionality. The only thing that matters is whether you can one-shot code like Claude or Codex, I'm not interested in a small…
Re: MAI-Code-1-Flash
#25Comparing against Claude 4.5? Aren't we up to 4.8? But disingenuous?
Even if it were Opus, comparing to a version number makes for an interesting snapshot of time comparison: if you knew how a model performed at whatever time in was in vogue, you can say "well, it looks like Model X is about 6 months/1 year/etc. behind the frontier SOTA" - which is exactly the discussion that happens in the open-weight/local LLM space. (interesting, MAI-Code-1-Flash does not appear to be such an open-weight model, following the western trend of locking models up)
Re: MAI-Code-1-Flash
#26Seems like the work from a good system design to code is practically solved.
Now it’s a matter of the design of the system. Or is that represented in these evals?
Re: MAI-Code-1-Flash
#27It's so weird to me that the benchmarks remain so low, but the models are marketed as revolutionary. And if you say that low coding capabilities aren't a problem, say that to the token price hike and 'general use' model setup. Why not sell it as a math agent? Why do I have to set up 4 agents to check each others' work?
It is my belief that smaller models will get better and better, and even cloud SOTA models will shrink.
Yet another reason the current buildout will feel like the railroads.
Re: MAI-Code-1-Flash
#28Re: MAI-Code-1-Flash
#29It is good to se big companies like Microsoft launching LLMs. They have large amount of compute power and good scientists to create useful models.