--
They should call this the siphon/sifter model of RL.
You siphon only the initial domains, then sift to the solution....
51–60 of 178 posts
--
They should call this the siphon/sifter model of RL.
You siphon only the initial domains, then sift to the solution....
I love that emphasizing math learning and coding leads to general reasoning skills. Probably works the same in humans, too. 20x smaller than Deep Seek! How small can these go? What kind of hardware can run this?
Overall though quite impressive if you're not in a hurry.
Chinese strategy is open-source software part and earn on robotics part. And, They are already ahead of everyone in that game. These things are pretty interesting as they are developing. What US will do to retain its power? BTW I am Indian and we are not even in the race as country. :(
If I had to guess, more tariffs and sanctions that increase the competing nation's self-reliance and harm domestic consumers. Perhaps my peabrain just can't comprehend the wisdom of policymakers on the sanctions front, but it just seems like all it does is empower the target long-term.
This makes it even better!
Note the massive context length (130k tokens). Also because it would be kinda pointless to generate a long CoT without enough context to contain it and the reply. EDIT: Here we are. My first prompt created a CoT so long that it catastrophically forgot the task (but I don't believe I was near 130k -- using ollama with fp16 model). I asked one of my test questions with a coding question totally unrelated to what it say…
i've also been experimenting with different chunking strategies to see if that helps maintain coherence over larger contexts. it's a tricky problem.
I guess I won’t be needing that 512GB M3 Ultra after all.
I love that emphasizing math learning and coding leads to general reasoning skills. Probably works the same in humans, too. 20x smaller than Deep Seek! How small can these go? What kind of hardware can run this?