Look into groq.com guys. some good models at similar speed to inception labs
Mercury: Commercial-scale diffusion language model
131–140 of 189 posts
Re: Mercury: Commercial-scale diffusion language model
#132Earlier quoted context omitted.
Unfortunately the intuition and the math proofs so far suggest that autoregressive training is learning the joint distribution of probabilistic streams of tokens much better than diffision models do or will ever do. My intuitive take is that the conditional probability distribtion of decoder-only autoregressive models is at just the right level of complexity for probabilistic models to learn accurately enough. Intuit…
This is tremendously interesting! Could you point me to some literature? Especially regarding mathematical proofs of your intuition? I’d like to recalibrate my priors to align better with current research results.
Re: Mercury: Commercial-scale diffusion language model
#133It's nice to see a team doing something different. The cost[1] is US$1.00 per million output tokens and US$0.25 per million input tokens. By comparison, Gemini 2.5 Flash Preview charges US$0.15 per million tokens for text input and $0.60 (non-thinking) output[2]. Hmmm... at those prices they need to focus on markets where speed is especially important, eg high-frequency trading, transcription/translation services and…
Chinese companies will be similarly eager for market share, but not everyone has the access to the same raw capital.
Re: Mercury: Commercial-scale diffusion language model
#134Earlier quoted context omitted.
So is my knowledge of newtons law of cooling
If an LLM has only that knowledge and nothing else (pieces of text saying that heat transfer is proportional to some function of the temp difference) such that is not trained on any texts that give problems and solutions in this area, it will not work this out, since it has nothing to generate tokens from. Also, your knowledge doesn't come from anywhere near having scanned terabytes of text, which would take you mult…
Re: Mercury: Commercial-scale diffusion language model
#135Not sure if I would tradeoff speed for accuracy. Yes, it's incredible boring to wait for the AI Agents in IDEs to finish their job. I get distracted and open YouTube. Once I gave a prompt so big and complex to Cline it spent 2 straight hours writing code. But after these 2 hours I spent 16 more tweaking and fixing all the stuff that wasn't working. I now realize I should have done things incrementally even when I hav…
I think speed and convenience are essential. I use chat gpt desktop for coding. Not because it's the best but because it's fast and easy and doesn't interrupt my flow too much. I mostly stick to the 4o model. I only use the o3 model when I really have to. Because at that point getting an answer is slooooow. 4o is more than good enough most of the time. And more importantly it's a simple option+shift+1 away. I simply…
Re: Mercury: Commercial-scale diffusion language model
#136Look into groq.com guys. some good models at similar speed to inception labs
Groq is heading to a dead end.
Re: Mercury: Commercial-scale diffusion language model
#137Earlier quoted context omitted.
AI field desperately needs smarter models - not faster models.
Definitely needs faster and cheaper models. Fast and cheap models could replace software in tons of situations. Imagine a vending machine or a mobile game or a word processor where basically all logic is implemented as a prompt to an llm. It would serve as the ultimate high level programming language.
Re: Mercury: Commercial-scale diffusion language model
#138Earlier quoted context omitted.
If an LLM has only that knowledge and nothing else (pieces of text saying that heat transfer is proportional to some function of the temp difference) such that is not trained on any texts that give problems and solutions in this area, it will not work this out, since it has nothing to generate tokens from. Also, your knowledge doesn't come from anywhere near having scanned terabytes of text, which would take you mult…
We get way more info than llms do, just not solely from text
Re: Mercury: Commercial-scale diffusion language model
#139I'd hope that with diffusion, it would be able to go back and forth between parts of the output to adjust issues with part of the output which it had previously generated. This would not be possible with a purely sequential model. However, > Prompt: Write a sentence with ten words which has exactly as many r’s in the first five words as in the last five > > Response: Rapidly running, rats rush, racing, racing.
Why not possible with autoregressive model? o4 mini https://chatgpt.com/share/681315c2-aa90-800d-b02d-c3ba653281...
Re: Mercury: Commercial-scale diffusion language model
#140High tech US service industry exports are cooked.