Live data from Hacker News

S1: A $6 R1 competitor?

timkellogg.me

361–370 of 430 posts

Re: S1: A $6 R1 competitor?

#361
post #358

Earlier quoted context omitted.

And what is compression but finding the minimum amount of information required to reproduce a phenomena? I.e. discovering natural laws.

Finding minimum complexity explanations isn't what finding natural laws is about, I'd say. It's considered good practice (Occam's razor), but it's often not really clear what the minimal model is, especially when a theory is relatively new. That doesn't prevent it from being a natural law, the key criterion is predictability of natural phenomena, imho. To give an example, one could argue that Lagrangian mechanics req…

Maybe I'm just a filthy computationalist, but the way I see it, the most accurate model of the universe is the one which makes the most accurate predictions with the fewest parameters.

The Newtonian model makes provably less accurate predictions than Einsteinian (yes, I'm using a different example), so while still useful in many contexts where accuracy is less important, the number of parameters it requires doesn't much matter when looking for the one true GUT.

My understanding, again as a filthy computationalist, is that an accurate model of the real bonafide underlying architecture of the universe will be the simplest possible way to accurately predict anything. With the word "accurately" doing all the lifting.

As always: https://www.sas.upenn.edu/~dbalmer/eportfolio/Nature%20of%20...

I'm sure there are decreasingly accurate, but still useful, models all the way up the computational complexity hierarchy. Lossy compression is, precisely, using one of them.

Re: S1: A $6 R1 competitor?

#362
post #287
post #204

Earlier quoted context omitted.

Your example is somewhat inadequate. We _fundamentally_ don’t understand how deep learning systems works in the sense that they are more or less black boxes that we train and evaluate. Innovations in ML are a whole bunch of wizards with big stacks of money changing “Hmm” to “Wait” and seeing what happens. Would a different sampler help you? I dunno, try it. Would a smaller dataset help? I dunno, try it. Would trainin…

> _fundamentally_ don’t understand how deep learning systems works. It's like saying we don't understand how quantum chromodynamics works. Very few people do, and it's the kind of knowledge not easily distilled for the masses in an easily digestible in a popsci way. Look into how older CNNs work -- we have very good visual/accesible/popsci materials on how they work. I'm sure we'll have that for LLM but it's not wort…

> The kind of progress being made leads me to believe there absolutely ARE people who absolutely know how the LLMs work

Just like alchemists made enormous strides in chemistry, but their goal was to turn piss into gold.

Re: S1: A $6 R1 competitor?

#363

>it can run on my laptop Has anyone run it on a laptop (unquantized)? Disk size of the 32B model appears to be 80GB. Update: I'm using a 40GB A100 GPU. Loading the model took 30GB vRAM. I asked a simple question "How many r in raspberry". After 5 minutes nothing got generated beyond the prompt. I'm not sure how the author ran this on a laptop.

32B models are easy to run on 24GB of RAM at a 4-bit quant.

It sounds like you need to play with some of the existing 32B models with better documentation on how to run them if you're having trouble, but it is entirely plausible to run this on a laptop.

I can run Qwen2.5-Instruct-32B-q4_K_M at 22 tokens per second on just an RTX 3090.

Re: S1: A $6 R1 competitor?

#364
post #308

Earlier quoted context omitted.

As a person who has trained a number of computer vision deep networks, I can tell you that we have some cool-looking visualizations on how lower layers work but no idea how later layers work. The intuition is built over training numerous networks and trying different hyperparameters, data shuffling, activations, etc. it’s absolutely brutal over here. If the theory was there, people like Karpathy who have great teache…

It may be as simple as this: https://youtube.com/shorts/7GrecDNcfMc Many many layers of that. It’s not a profound mechanism. We can understand how that works, but we’re dumbfounded how such a small mechanism is responsible for all this stuff going on inside a brain. I don’t think we don’t understand, it’s a level beyond that. We can’t fathom the implications, that it could be that simple, just scaled up.

> Many many layers of that. It’s not a profound mechanism

Bad argument. Cavemen understood stone, but they could not build the aqueducts. Medieval people understood iron, water and fire but they could not make a steam engine

Finally we understand protons, electrons, and neutrons and the forces that government them but it does not mean we understand everything they could mossibly make

Re: S1: A $6 R1 competitor?

#365

Earlier quoted context omitted.

I've had an idea since I was a kid which I can share. I was contemplating AI and consciousness generally, probably around the time I read "The Minds I". I reflected on the pop-psychology idea of consciousness and subconsciousness. I thought of each as an independent stream of tokens, like stream of consciousness poetry. But along the stream there were joining points between these two streams, points where the conscio…

Have you read Jaynes' "The Origin of Consciousness in the Breakdown of the Bicameral Mind"?

I haven't read the original but I am familiar with the broad stroke view. There are similarities (perhaps vague) in the more recent work of someone like McGilchrist and his The Master and His Emissary (another book which I only have a broad stroke view of).

At the time I had this idea I did not know of either of these. I think I was drawing explicitly on the conscious / subconscious vocabulary.

Re: S1: A $6 R1 competitor?

#366
post #337

Earlier quoted context omitted.

This is something I have been suppressing since I don't want to become chicken little. Anyone who isn't terrified by the last 3 months probably doesn't really understand what is happening. I went from accepting I wouldn't see a true AI in my lifetime, to thinking it is possible before I die, to thinking it is possible in in the next decade, to thinking it is probably in the next 3 years to wondering if we might see i…

This frightens mostly people whose identity is built around "intelligence", but without grounding in the real world. I've yet to see really good articulations of what, precisely we should be scared of. Bedroom superweapons? Algorithmic propaganda? These things have humans in the loop building them. And the problem of "human alignment" is one unsolved since Cain and Abel. AI alone is words on a screen. The sibling thr…

Some of the scariest horror movies are the ones where the monster isn't shown. Often once the monster is shown, it is less terrifying.

In a general sense, uncertainty causes anxiety. Once you know the properties of the monster you are dealing with you can start planning on how to address it.

Some people have blind and ignorant confidence. A feeling they can take on literally anything, no matter how powerful. Sometimes they are right, sometimes they are wrong.

I'm reminded by the scene in No Country For Old Men where the good guy bad-ass meets the antagonist and immediately dies. I have little faith in blind confidence.

edit: I'll also add that human adaptability (which is probably the trait most confidence in humans would rest) has shown itself capable of saving us from many previous civilization changing events. However, this change with AI is happening much, much faster than any before it. So part of the anxiety is whether or not our species reaction time is enough to avoid the cliff we are accelerating towards.

Re: S1: A $6 R1 competitor?

#367
post #341
post #337

Earlier quoted context omitted.

This frightens mostly people whose identity is built around "intelligence", but without grounding in the real world. I've yet to see really good articulations of what, precisely we should be scared of. Bedroom superweapons? Algorithmic propaganda? These things have humans in the loop building them. And the problem of "human alignment" is one unsolved since Cain and Abel. AI alone is words on a screen. The sibling thr…

> without grounding in the real world. > I've yet to see really good articulations of what, precisely we should be scared of. Bedroom superweapons? Loss of paid employment opportunities and increasing inequality are real world concerns. UBI isn't coming by itself.

Worst case scenario humans mostly go back to manual labor, which would fix a lot of modern day ailments such as obesity and (some) mental health struggles, with added enormous engineering advancements based on automatic research.

Re: S1: A $6 R1 competitor?

#368

Earlier quoted context omitted.

It may be as simple as this: https://youtube.com/shorts/7GrecDNcfMc Many many layers of that. It’s not a profound mechanism. We can understand how that works, but we’re dumbfounded how such a small mechanism is responsible for all this stuff going on inside a brain. I don’t think we don’t understand, it’s a level beyond that. We can’t fathom the implications, that it could be that simple, just scaled up.

> Many many layers of that. It’s not a profound mechanism Bad argument. Cavemen understood stone, but they could not build the aqueducts. Medieval people understood iron, water and fire but they could not make a steam engine Finally we understand protons, electrons, and neutrons and the forces that government them but it does not mean we understand everything they could mossibly make

"Cavemen understood stone"

How far removed are you from a caveman is the better question. There would be quite some arrogance coming out of you to suggest the several million years gap is anything but an instant in the grand timeline. As in, you understood stone just yesterday ...

The monkey that found the stone is the monkey that built the cathedral. It's only a delusion the second monkey creates to separate it from the first monkey (a feeling of superiority, with the only tangible asset being "a certain amount of notable time passed since point A and point B").

"Finally we understand protons, electrons, and neutrons and the forces that government them but it does not mean we understand everything they could mossibly make"

You and I agree. That those simple things can truly create infinite possibilities. That's all I was saying, we cannot fathom it (either because infinity is hard to fathom, or that it's origins are humble - just a few core elements, or both, or something else).

Anyway, this can discussion can head into any direction.

Re: S1: A $6 R1 competitor?

#369

The part about taking control of a reasoning model's output length using tags is interesting. > In s1, when the LLM tries to stop thinking with " ", they force it to keep going by replacing it with "Wait". I had found a few days ago that this let you 'inject' your own CoT and jailbreak it easier. Maybe these are related? https://pastebin.com/G8Zzn0Lw https://news.ycombinator.com/item?id=42891042#42896498

It's weird that you need to do that at all, couldn't you just reject that token and use the next most probable?

Re: S1: A $6 R1 competitor?

#370

Earlier quoted context omitted.

Except I can run R1 1.5b on a GPU-less and NPU-less Intel NUC from four-five years ago using half its cores and the reply speed is…functional. As the models have gotten more efficient and distillation better the minimum viable hardware for really cooking with LLMs has gone from a 4090 to suddenly something a lot of people already probably own. I definitely think a Digits box would be nice, but honestly I’m not sure I…

R1 1.5b won’t do what most people want at all.

No, it won't. But that's not the point I was making
Post reply on HN