Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
1–10 of 29 posts
Re: Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
#2Man, I find SOTA deep learning somewhat hilarious. We scale models to absurd proportions, burning through a shitload of resources just to achieve (slightly above) human intelligence. The human brain has a few billion neurons and uses as much power as a light bulb.
Re: Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
#3> scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; Man, I find SOTA deep learning somewhat hilarious. We scale models to absurd proportions, burning through a shitload of resources just to achieve (slightly above) human intelligence. The human brain has a few billion neurons and uses as much power as a light bulb.
However, a neuron is much more than a single parameter. The brain is estimated to have from 10^14 to 5x10^14 synapses.
Re: Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
#4> scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; Man, I find SOTA deep learning somewhat hilarious. We scale models to absurd proportions, burning through a shitload of resources just to achieve (slightly above) human intelligence. The human brain has a few billion neurons and uses as much power as a light bulb.
Re: Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
#5> scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; Man, I find SOTA deep learning somewhat hilarious. We scale models to absurd proportions, burning through a shitload of resources just to achieve (slightly above) human intelligence. The human brain has a few billion neurons and uses as much power as a light bulb.
huge parameter models with many small but efficient layers can work quickly on low resource hardware
similar to how neurons experience chemical spiking to activate small portions of the brain at once
Re: Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
#6> scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; Man, I find SOTA deep learning somewhat hilarious. We scale models to absurd proportions, burning through a shitload of resources just to achieve (slightly above) human intelligence. The human brain has a few billion neurons and uses as much power as a light bulb.
Re: Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
#7> scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; Man, I find SOTA deep learning somewhat hilarious. We scale models to absurd proportions, burning through a shitload of resources just to achieve (slightly above) human intelligence. The human brain has a few billion neurons and uses as much power as a light bulb.
Re: Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
#8> scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; Man, I find SOTA deep learning somewhat hilarious. We scale models to absurd proportions, burning through a shitload of resources just to achieve (slightly above) human intelligence. The human brain has a few billion neurons and uses as much power as a light bulb.
Do you find cars similarly hilarious?
Re: Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
#9> scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; Man, I find SOTA deep learning somewhat hilarious. We scale models to absurd proportions, burning through a shitload of resources just to achieve (slightly above) human intelligence. The human brain has a few billion neurons and uses as much power as a light bulb.
When for the training part you have to consider brains had like billions of years to develop. Maybe one of the reasons llms seem to be so expensive to train is because we are "compressing" in far less time that learning part
Re: Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
#10> scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; Man, I find SOTA deep learning somewhat hilarious. We scale models to absurd proportions, burning through a shitload of resources just to achieve (slightly above) human intelligence. The human brain has a few billion neurons and uses as much power as a light bulb.
When for the training part you have to consider brains had like billions of years to develop. Maybe one of the reasons llms seem to be so expensive to train is because we are "compressing" in far less time that learning part