Live data from Hacker News

MAI-Thinking-1

microsoft.ai

41–50 of 90 posts

Re: MAI-Thinking-1

#41
> MAI-Thinking-1 is a 35B-active, ~1T-total parameters, sparse Mixture of Experts model, a smaller inference footprint than much larger models.

This seemingly nonsensical sentence (of course this will have a smaller inference footprint than larger models) suggests this model's competitors have larger inference footprints and total parameter sizes.

Re: MAI-Thinking-1

#42
post #13

Earlier quoted context omitted.

I'm interested how much "Clean Data" is synthetic data from "unclean" models...

> with AI-generated content excluded from pre-training. > without distillation from third-party models sounds like zero unless they are lying.

"how many of those shapes are rectangles?" "sounds like zero unless they are squares"

Adding "unless" to a statement makes it vacuous if the latter clause is weaker than the first clause. I find it hard to believe that a company willing to violate licenses would have scruples about lying about it.

Re: MAI-Thinking-1

#43
post #5

> Second, clean data. MAI-Thinking-1 was trained on clean and appropriately licensed data, with AI-generated content excluded from pre-training. This matters for quality, provenance, and control. If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it. Shots fired? It would be interesting to see how far "clean data" can go on the scaling laws.

I would really like to see what "appropriately licensed data" means. Cannot imagine they didn't copy all open repo's on GitHub, and can't imagine they asked for permission, or are reproducing license texts from these repo's now. It sounds hand wavy.

P.S. A fairly basic website otherwise, but it unfortunately seems to be hacking scroll for no good reason.

Re: MAI-Thinking-1

#44

Looks like the OAI divergence is finally taking place. Seems like the comparisons are mainly with Opus 4.6 and GPT 5.4 though. Still, exciting to see a new frontier player.

Post 4.6 Anthropic models do not exactly have a stellar reputation, so that choice is smart.

Re: MAI-Thinking-1

#45
post #5

> Second, clean data. MAI-Thinking-1 was trained on clean and appropriately licensed data, with AI-generated content excluded from pre-training. This matters for quality, provenance, and control. If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it. Shots fired? It would be interesting to see how far "clean data" can go on the scaling laws.

I would really like to see what "appropriately licensed data" means. Cannot imagine they didn't copy all open repo's on GitHub, and can't imagine they asked for permission, or are reproducing license texts from these repo's now. It sounds hand wavy. P.S. A fairly basic website otherwise, but it unfortunately seems to be hacking scroll for no good reason.

I assume they took the actual repos’ licenses info account. I don’t understand why they should ask for permission when the license would already allow for it.

Re: MAI-Thinking-1

#47

> MAI-Thinking-1 is built with enterprise readiness in mind. It supports long context with a 256k token window Isn’t 1M becoming the norm?

Yes it is, but I can imagine that they want to start out a bit smaller to see how well things scale, and/or did not yet have the time to work on optimizing for the large context windows.

Re: MAI-Thinking-1

#48

It's good there is a new player on the market, I take benchmark tables with a grain of salt, however. Speaking about model presentation it's funny to see how clearly their website is inspired by other AI company blogs with extra innovation of hijacked scrollbar.

[deleted]

Re: MAI-Thinking-1

#49

> MAI-Thinking-1 is built with enterprise readiness in mind. It supports long context with a 256k token window Isn’t 1M becoming the norm?

Yes it is, but I can imagine that they want to start out a bit smaller to see how well things scale, and/or did not yet have the time to work on optimizing for the large context windows.

I struggle to get quality results from the frontier models at contexts > 256k anyway.

Re: MAI-Thinking-1

#50
post #42
post #13

Earlier quoted context omitted.

> with AI-generated content excluded from pre-training. > without distillation from third-party models sounds like zero unless they are lying.

"how many of those shapes are rectangles?" "sounds like zero unless they are squares" Adding "unless" to a statement makes it vacuous if the latter clause is weaker than the first clause. I find it hard to believe that a company willing to violate licenses would have scruples about lying about it.

Adding "unless" to a statement makes it vacuous if the latter clause is weaker than the first clause

I think that's the point. "How do I say they're lying without outright saying they're lying?"

It's a common rhetorical trick.

Post reply on HN