Live data from Hacker News

Xiaomi MiMo Reasoning Model

github.com

41–50 of 203 posts

Re: Xiaomi MiMo Reasoning Model

#42
post #20

Earlier quoted context omitted.

But who will keep them updated and what incentive they would have? That's I can't imagine. Bit vague.

Who keeps open source projects maintained and what incentive do they have?

For the bigger open source projects, companies who use that code for making money. Such as Microsoft and Google and IBM (and many others) supporting Linux because they use it extensively. The same answer may end up applying to these models though - if they really become something that gets integrated into products and internal workflows, there will be a market for companies to collaborate on maintaining a good implementation rather than competing needlessly.

Re: Xiaomi MiMo Reasoning Model

#44
post #30

Earlier quoted context omitted.

Last time I did that I was also impressed, for a start. Problem was that of a top ten book recommendations only the first 3 existed and the rest was a casually blended hallucination delivered in perfect English without skipping a beat. "You like magic? Try reading the Harlew Porthouse series by JRR Marrow, following the orphan magicians adventures in Hogwesteros" And the further towards the context limit it goes the…

LLMs are not search engines…

Exactly, I think all those base models should be weeded out from this nonsense, kardashian-like labyrinths of knowledge complexities that just makes them dumber by taking space and compute time. If you can google out some nonsense news, it should stay there in search engines for retrieval. Models should be good at using search tools, not at trying to replicate their results. They should start from logic, math, programming, physics and so on, similar to how education system is suppose to equip you with. IMHO small models can give this speed advantage (faster to experiment ie. with parallel diverging results, ability to munch through more data etc). Stripped to this bare minimum they can likely be much smaller with impressive results, tunable, allow for huge context etc.

Re: Xiaomi MiMo Reasoning Model

#45
post #24

Earlier quoted context omitted.

The smaller models have been creeping upward. They don't make headlines because they aren't leapfrogging the mainline models from the big companies, but they are all very capable. I loaded up a random 12B model on ollama the other day and couldn't believe how good it competent it seemed and how fast it was given the machine I was on. A year or so ago, that would have not been the case.

What model? I have been using api's mostly since ollama was too slow for me.

I really like Gemma 3. Some quantized version of the 27B will be good enough for a lot of things. You can also take some abliterated version[0] with zero (like zero zero) guardrails and make it write you a very interesting crime story without having to deal with the infamous "sorry but I'm a friendly and safe model and cannot do that and also think about the children" response.

[0]: https://huggingface.co/mlabonne/gemma-3-12b-it-abliterated

Post reply on HN