> "Note that this s1 dataset is distillation. Every example is a thought trace generated by another model, Qwen2.5" The traces are generated by Gemini Flash Thinking. 8 hours of H100 is probably more like $24 if you want any kind of reliability, rather than $6.
"You can train a SOTA LLM for $0.50" (as long as you're distilling a model that cost $500m into another pretrained model that cost $5m)
S1: A $6 R1 competitor?
191–200 of 430 posts
Re: S1: A $6 R1 competitor?
#192I found the discussion around inference scaling with the 'Wait' hack so surreal. The fact such an ingeniously simple method can impact performance makes me wonder how many low-hanging fruit we're still missing. So weird to think that improvements on a branch of computer science is boiling down to conjuring the right incantation words, how you even change your mindset to start thinking this way?
Now imagine where we are in 12 months from now. This article from February 5 2025 will feel quaint by then. The acceleration keeps increasing. It seems likely we will soon have recursive self-improving AI -- reasoning models which do AI research. This will accelerate the rate of acceleration itself. It sounds stupid to say it, but yes, the singularity is near. Vastly superhuman AI now seems to arrive within the next…
I went from accepting I wouldn't see a true AI in my lifetime, to thinking it is possible before I die, to thinking it is possible in in the next decade, to thinking it is probably in the next 3 years to wondering if we might see it this year.
Just 6 months ago people were wondering if pre-training was stalling out and if we hit a wall. Then deepseek drops with RL'd inference time compute, China jumps from being 2 years behind in the AI race to being neck-and-neck and we're all wondering what will happen when we apply those techniques to the current full-sized behemoth models.
It seems the models that are going to come out around summer time may be jumps in capability beyond our expectations. And the updated costs means that there may be several open source alternatives available. The intelligence that will be available to the average technically literate individual will be frightening.
Re: S1: A $6 R1 competitor?
#193I found the discussion around inference scaling with the 'Wait' hack so surreal. The fact such an ingeniously simple method can impact performance makes me wonder how many low-hanging fruit we're still missing. So weird to think that improvements on a branch of computer science is boiling down to conjuring the right incantation words, how you even change your mindset to start thinking this way?
I think the fact alone that distillation and quantization are techniques that can produce substantial improvements is a strong sign that we still have no real comprehensive understanding how the models work. If we had, there would be no reason to train a model with more parameters than are strictly necessary to represent the space's semantic structure. But then it should be impossible for distilled models with less p…
We do understand how they work, we just have not optimised their usage.
For example someone who has a good general understanding of how an ICE or EV car works. Even if the user interface is very unfamiliar, they can figure out how to drive any car within a couple of minutes.
But that does not mean they can race a car, drift a car or drive a car on challenging terrain even if the car is physically capable of all these things.
Re: S1: A $6 R1 competitor?
#194If an LLM output is like a sculpture, then we have to sculpt it. I never did sculpting, but I do know they first get the clay spinning on a plate. Whatever you want to call this “reasoning” step, ultimately it really is just throwing the model into a game loop. We want to interact with it on each tick (spin the clay), and sculpt every second until it looks right. You will need to loop against an LLM to do just about…
We need to cluster the AI's insights on a spatial grid hash, give it a minimap with the ability to zoom in and out, and give it the agency to try and find its way to an answer and build up confidence and tests for that answer.
coarse -> fine, refine, test, loop.
Maybe a parallel model that handles the visualization stuff. I imagine its training would look more like computer vision. Mind palace generation.
If you're stuck or your confidence is low, wander the palace and see what questions bubble up.
Bringing my current context back through the web is how I think deeply about things. The context has the authority to reorder the web if it's "epiphany grade".
I wonder if the final epiphany at the end of what we're creating is closer to "compassion for self and others" or "eat everything."
Re: S1: A $6 R1 competitor?
#195Earlier quoted context omitted.
I think you're missing the point: H100 isn't going to remain useful for a long time, would you consider Tesla or Pascal graphic cards a collateral? That's what those H100 will look like in just a few years.
Not sure I do tbh. Any asset depreciates over time. But they usually get replaced. My 286 was replaced by a faster 386 and that by an even faster 468. I’m sure you see a naming pattern there.
Re: S1: A $6 R1 competitor?
#196If chain of thought acts as a scratch buffer by providing the model more temporary "layers" to process the text, I wonder if making this buffer a separate context with its own separate FNN and attention would make sense; in essence, there's a macroprocess of "reasoning" that takes unbounded time to complete, and then there's a microprocess of describing this incomprehensible stream of embedding vectors in natural lan…
I reflected on the pop-psychology idea of consciousness and subconsciousness. I thought of each as an independent stream of tokens, like stream of consciousness poetry. But along the stream there were joining points between these two streams, points where the conscious stream was edited by the subconscious stream. You could think of the subconscious stream as performing CRUD like operations on the conscious stream. The conscious stream would act like a buffer of short-term memory while the subconscious stream would act like a buffer of long-term memory. Like, the subconscious has instructions related to long-term goals and the conscious stream has instructions related to short-term goals.
You can imagine perception as input being fed into the conscious stream and then edited by the subconscious stream before execution.
It seems entirely possible to actually implement this idea in this current day and age. I mean, it was a fever dream as a kid, but now it could be an experiment!
Re: S1: A $6 R1 competitor?
#197> I doubt that OpenAI has a realistic path to preventing or even detecting distealing outside of simply not releasing models. Couldn't they just start hiding the thinking portion? It would be easy for them to do this. Currently, they already provide one sentence summaries for each step of the thinking I think users would be fine or at least stay if it were changed to provide only that.
They hid it and deepseek came up with R1 anyway, with RL on only results and not even needing any of the thinking tokens that OpenAI hid.
Re: S1: A $6 R1 competitor?
#198> "Note that this s1 dataset is distillation. Every example is a thought trace generated by another model, Qwen2.5" The traces are generated by Gemini Flash Thinking. 8 hours of H100 is probably more like $24 if you want any kind of reliability, rather than $6.
"You can train a SOTA LLM for $0.50" (as long as you're distilling a model that cost $500m into another pretrained model that cost $5m)
Re: S1: A $6 R1 competitor?
#199Earlier quoted context omitted.
I think you're missing the point: H100 isn't going to remain useful for a long time, would you consider Tesla or Pascal graphic cards a collateral? That's what those H100 will look like in just a few years.
Not sure I do tbh. Any asset depreciates over time. But they usually get replaced. My 286 was replaced by a faster 386 and that by an even faster 468. I’m sure you see a naming pattern there.
Given that inference time will soon be extremely valuable with agents and models, H100s may yet be worth something in a couple years.
Re: S1: A $6 R1 competitor?
#200Earlier quoted context omitted.
This is pure speculation on my part but I think at some point a company's valuation became tied to how big their compute is so everybody jumped on the bandwagon.
I don't think you need to speculate too hard. On CNBC they are not tracking revenue, profits or technical breakthroughs, but how much the big companies are spending (on gpus). That's the metric!
Burn rate based valuations!
The 2000's are back in full force!