Live data from Hacker News

Mercury: Commercial-scale diffusion language model

inceptionlabs.ai

1–10 of 189 posts

Re: Mercury: Commercial-scale diffusion language model

#2
There are some open weight attempts at this around too: https://old.reddit.com/r/LocalLLaMA/search?q=diffusion&restr...

Saw another on Twitter past few days that looked like a better contender to Mercury, doesn't look like it got posted to LocalLLaMa, and I can't find it now. Very exciting stuff

Re: Mercury: Commercial-scale diffusion language model

#3
There are so many models. Every single day half a dozen new models land. And even more papers.

It feels like models are becoming fungible apart from the hyperscaler frontier models from OpenAI, Google, Anthropic, et al.

I suppose VCs won't be funding many more "labs"-type companies or "we have a model" as the core value prop companies? Unless it has a tight application loop or is truly unique?

Disregarding the team composition, research background, and specific problem domain - if you were starting an AI company today, what part of the stack would you focus on? Foundation models, AI/ML infra, tooling, application layer, ...?

Where does the value accrue? What are the most important problems to work on?

Re: Mercury: Commercial-scale diffusion language model

#4
Interesting approach. However, I never thought of auto regression being _the_ current issue with language modeling. If anything it seems the community was generally surprised just how far next "token" prediction took us. Remember back when we did char generating RNNs and were impressed they could make almost coherent sentences?

Diffusion is an alternative but I am having a hard time understanding the whole "built in error correction" that sounds like marketing BS. Both approaches replicate probability distributions which will be naturally error-prone because of variance.

Re: Mercury: Commercial-scale diffusion language model

#5
Ok. My go to puzzle is this:

You have 2 minutes to cool down a cup of coffee to the lowest temp you can

You have two options:

1. Add cold milk immediately, then let it sit for 2 mins.

2. Let it sit for 2 mins, then add the cold milk.

Which one cools the coffee to the lowest temperature and why?

And Mercury gets this right - while as of right now ChatGPT 4o get it wrong.

So that’s pretty impressive.

Re: Mercury: Commercial-scale diffusion language model

#7
Not sure if I would tradeoff speed for accuracy.

Yes, it's incredible boring to wait for the AI Agents in IDEs to finish their job. I get distracted and open YouTube. Once I gave a prompt so big and complex to Cline it spent 2 straight hours writing code.

But after these 2 hours I spent 16 more tweaking and fixing all the stuff that wasn't working. I now realize I should have done things incrementally even when I have a pretty good idea of the final picture.

I've been more and more only using the "thinking" models of o3 in ChatGPT, and Gemini / Claude in IDEs. They're slower, but usually get it right.

But at the same time I am open to the idea that speed can unlock new ways of using the tooling. It would still be awesome to basically just have a conversation with my IDE while I am manually testing the app. Or combine really fast models like this one with a "thinking background" one, that would runs for seconds/minutes but try to catch the bugs left behind.

I guess only giving a try will tell.

Re: Mercury: Commercial-scale diffusion language model

#8
This sounds like a neat idea but it seems like bad timing. OpenAI just released token-based that beats the best diffusion image generation. If diffusion isn't even the best at generating images, I don't know if I'm going to spend a lot of time evaluating it for text.

Speed is great but it doesn't seem like other text-based model trends are going to work out of the box, like reasoning. So you have to get dLLMs up to the quality of a regular autoregressive LLM and then you need to innovate more to catch up to reasoning models, just to match the current state of the art. It's possible they'll get there, but I'm not optimistic.

Re: Mercury: Commercial-scale diffusion language model

#9

Super happy to see something like this getting traction. As someone that is trying to reduce my carbon footprint sometimes I feel bad about asking any model to do something trivial. With something like that perhaps the guilt will lessen

If you live in the U.S., marginal electricity demand during the day is almost invariably met with solar or wind (solar typically runs at a huge surplus on sunny days). Go forth and AI in peace, marcyb5st.

Re: Mercury: Commercial-scale diffusion language model

#10

Ok. My go to puzzle is this: You have 2 minutes to cool down a cup of coffee to the lowest temp you can You have two options: 1. Add cold milk immediately, then let it sit for 2 mins. 2. Let it sit for 2 mins, then add the cold milk. Which one cools the coffee to the lowest temperature and why? And Mercury gets this right - while as of right now ChatGPT 4o get it wrong. So that’s pretty impressive.

Token-based Gemini is 4 seconds of thinking:

Okay, let's break this down using the principle of heat transfer (specifically Newton's Law of Cooling):

    Heat loss is faster when the temperature difference is greater. A very hot object loses heat to cool surroundings much faster than a warm object.

    Option 1: Add milk immediately.

        You instantly lower the coffee's temperature by mixing it with cold milk.

        Now, this warmer (not hot) mixture sits for 2 minutes. Because the temperature difference between the mixture and the room is smaller, it cools down more slowly over those 2 minutes.

    Option 2: Let it sit for 2 mins, then add milk.

        The very hot coffee sits for 2 minutes. Because the temperature difference between the hot coffee and the room is large, it loses heat more quickly during these 2 minutes.

        After 2 minutes of rapid cooling, you then add the cold milk, lowering the temperature further.
Conclusion:

To get the coffee to the lowest temperature, you should choose Option 2: Let it sit for 2 mins, then add the cold milk.

Post reply on HN