Mercury: Commercial-scale diffusion language model
inceptionlabs.ai
Mercury: Commercial-scale diffusion language model
1–10 of 189 posts
Re: Mercury: Commercial-scale diffusion language model
#2Saw another on Twitter past few days that looked like a better contender to Mercury, doesn't look like it got posted to LocalLLaMa, and I can't find it now. Very exciting stuff
Re: Mercury: Commercial-scale diffusion language model
#3It feels like models are becoming fungible apart from the hyperscaler frontier models from OpenAI, Google, Anthropic, et al.
I suppose VCs won't be funding many more "labs"-type companies or "we have a model" as the core value prop companies? Unless it has a tight application loop or is truly unique?
Disregarding the team composition, research background, and specific problem domain - if you were starting an AI company today, what part of the stack would you focus on? Foundation models, AI/ML infra, tooling, application layer, ...?
Where does the value accrue? What are the most important problems to work on?
Re: Mercury: Commercial-scale diffusion language model
#4Diffusion is an alternative but I am having a hard time understanding the whole "built in error correction" that sounds like marketing BS. Both approaches replicate probability distributions which will be naturally error-prone because of variance.
Re: Mercury: Commercial-scale diffusion language model
#5You have 2 minutes to cool down a cup of coffee to the lowest temp you can
You have two options:
1. Add cold milk immediately, then let it sit for 2 mins.
2. Let it sit for 2 mins, then add the cold milk.
Which one cools the coffee to the lowest temperature and why?
And Mercury gets this right - while as of right now ChatGPT 4o get it wrong.
So that’s pretty impressive.
Re: Mercury: Commercial-scale diffusion language model
#6Re: Mercury: Commercial-scale diffusion language model
#7Yes, it's incredible boring to wait for the AI Agents in IDEs to finish their job. I get distracted and open YouTube. Once I gave a prompt so big and complex to Cline it spent 2 straight hours writing code.
But after these 2 hours I spent 16 more tweaking and fixing all the stuff that wasn't working. I now realize I should have done things incrementally even when I have a pretty good idea of the final picture.
I've been more and more only using the "thinking" models of o3 in ChatGPT, and Gemini / Claude in IDEs. They're slower, but usually get it right.
But at the same time I am open to the idea that speed can unlock new ways of using the tooling. It would still be awesome to basically just have a conversation with my IDE while I am manually testing the app. Or combine really fast models like this one with a "thinking background" one, that would runs for seconds/minutes but try to catch the bugs left behind.
I guess only giving a try will tell.
Re: Mercury: Commercial-scale diffusion language model
#8Speed is great but it doesn't seem like other text-based model trends are going to work out of the box, like reasoning. So you have to get dLLMs up to the quality of a regular autoregressive LLM and then you need to innovate more to catch up to reasoning models, just to match the current state of the art. It's possible they'll get there, but I'm not optimistic.
Re: Mercury: Commercial-scale diffusion language model
#9Super happy to see something like this getting traction. As someone that is trying to reduce my carbon footprint sometimes I feel bad about asking any model to do something trivial. With something like that perhaps the guilt will lessen
Re: Mercury: Commercial-scale diffusion language model
#10Ok. My go to puzzle is this: You have 2 minutes to cool down a cup of coffee to the lowest temp you can You have two options: 1. Add cold milk immediately, then let it sit for 2 mins. 2. Let it sit for 2 mins, then add the cold milk. Which one cools the coffee to the lowest temperature and why? And Mercury gets this right - while as of right now ChatGPT 4o get it wrong. So that’s pretty impressive.
Okay, let's break this down using the principle of heat transfer (specifically Newton's Law of Cooling):
Heat loss is faster when the temperature difference is greater. A very hot object loses heat to cool surroundings much faster than a warm object.
Option 1: Add milk immediately.
You instantly lower the coffee's temperature by mixing it with cold milk.
Now, this warmer (not hot) mixture sits for 2 minutes. Because the temperature difference between the mixture and the room is smaller, it cools down more slowly over those 2 minutes.
Option 2: Let it sit for 2 mins, then add milk.
The very hot coffee sits for 2 minutes. Because the temperature difference between the hot coffee and the room is large, it loses heat more quickly during these 2 minutes.
After 2 minutes of rapid cooling, you then add the cold milk, lowering the temperature further.
Conclusion:To get the coffee to the lowest temperature, you should choose Option 2: Let it sit for 2 mins, then add the cold milk.