Live data from Hacker News

S1: A $6 R1 competitor?

timkellogg.me

101–110 of 430 posts

Re: S1: A $6 R1 competitor?

#101
post #70

Earlier quoted context omitted.

can you please elaborate on the wait tokens? what's that? how do they work? is that also from the R1 paper?

The same idea is in both the R1 and S1 papers ( tokens are used similarly). Basically they're using special tokens to mark in the prompt where the LLM should think more/revise the previous response. This can be repeated many times until some stop criteria occurs. S1 manually inserts these with heuristics, R1 learns the placement through RL I think.

? theyre not special tokens really

Re: S1: A $6 R1 competitor?

#102
post #97

> having 10,000 H100s just means that you can do 625 times more experiments than s1 did I think the ball is very much in their court to demonstrate they actually are using their massive compute in such a productive fashion. My BigTech experience would tend to suggest that frugality went out the window the day the valuation took off, and they are in fact just burning compute for little gain, because why not...

This claim is mathematically nonsensical. It implies a more-or-less linear relationship, that more is always better. But there's no reason to limit that to H100s. Conventional servers are, if anything, rather more established in their ability to generate value, by which I mean, however much potential AI servers may have to be more important than conventional servers that they may manifest in the future, we know how t…

[dead]

Re: S1: A $6 R1 competitor?

#103

Hmmm, 1 + 1 equals 3. Alternatively, 1 + 1 equals -3. Wait, actually 1 + 1 equals 1.

As one with teaching experience, the idea of asking a student "are you sure about that?" is to get them to think more deeply rather than just blurting a response. It doesn't always work, but it generally does.

Re: S1: A $6 R1 competitor?

#104
post #53

If chain of thought acts as a scratch buffer by providing the model more temporary "layers" to process the text, I wonder if making this buffer a separate context with its own separate FNN and attention would make sense; in essence, there's a macroprocess of "reasoning" that takes unbounded time to complete, and then there's a microprocess of describing this incomprehensible stream of embedding vectors in natural lan…

Here's a paper your idea reminds me of. https://arxiv.org/abs/2501.19201

It's also so not far from Meta's large concept model idea.

Re: S1: A $6 R1 competitor?

#105

> having 10,000 H100s just means that you can do 625 times more experiments than s1 did I think the ball is very much in their court to demonstrate they actually are using their massive compute in such a productive fashion. My BigTech experience would tend to suggest that frugality went out the window the day the valuation took off, and they are in fact just burning compute for little gain, because why not...

Mainly it points to a non-scientific "bigger is better" mentality, and the researchers probably didn't mind playing around with the power because "scale" is "cool". Remember that the Lisp AI-labs people were working on non-solved problems on absolute potatoes of computers back in the day, we have a semblance of progress solution but so much of it has been brute-force (even if there has been improvements in the field)…

[deleted]

Re: S1: A $6 R1 competitor?

#106
post #67

Earlier quoted context omitted.

When you're only used to ollama, how do I go about using this model?

I think we need to wait for someone to convert it into a GGUF file format. However, once that happens, you can run it (and any GGUF model) from Hugging Face![0] [0] https://huggingface.co/docs/hub/en/ollama

you can load the safetensors with ollama, you just have to provide a modelfile. or wait for someone to do it. It will in theory also quantize it for you, as I guess most ppl cannot load a 129 GB model...

Re: S1: A $6 R1 competitor?

#107
post #8

> If you believe that AI development is a prime national security advantage, then you absolutely should want even more money poured into AI development, to make it go even faster. This, this is the problem for me with people deep in AI. They think it’s the end all be all for everything. They have the vision of the ‘AI’ they’ve seen in movies in mind, see the current ‘AI’ being used and to them it’s basically almost t…

I couldn't agree more. If we're not talking about cyber war exclusively, such as finding and exploiting vulnerabilities, for the time being national security will still be based on traditional army. Just a few weeks ago, italy announced a 16bln€ plan to buy >1000 rheinmetall ifv vehicles. That alone would make italy's army one of the most equipped in Europe. I can't imagine what would happen with a 500$bln investment…

> I can't imagine what would happen with a 500$bln investment in defense,lol.

The $90,000 bag of bushings becomes a $300,000 bag?

Re: S1: A $6 R1 competitor?

#108
post #8

> If you believe that AI development is a prime national security advantage, then you absolutely should want even more money poured into AI development, to make it go even faster. This, this is the problem for me with people deep in AI. They think it’s the end all be all for everything. They have the vision of the ‘AI’ they’ve seen in movies in mind, see the current ‘AI’ being used and to them it’s basically almost t…

You can choose to be somewhat ignorant of the current state in AI, about which I could also agree that at certain moments it appears totally overhyped, but the reality is that there hasn't been a bigger technology breakthrough probably in the last ~30 years. This is not "just" machine learning because we have never been able to do things which we are today and this is not only the result of better hardware. Better ha…

> the first one being from DeepMind in 2017

? what paper are you talking about

Re: S1: A $6 R1 competitor?

#109
post #37
post #8

> If you believe that AI development is a prime national security advantage, then you absolutely should want even more money poured into AI development, to make it go even faster. This, this is the problem for me with people deep in AI. They think it’s the end all be all for everything. They have the vision of the ‘AI’ they’ve seen in movies in mind, see the current ‘AI’ being used and to them it’s basically almost t…

Agreed. I was working on some haiku things with ChatGPT and it kept telling me that busy has only one syllable. This is a trivially searchable fact.

link a chat please

Re: S1: A $6 R1 competitor?

#110
post #57

Earlier quoted context omitted.

> But I don’t see why that will turn into Data from Star Trek. "Is Data genuinely sentient or is he just a machine with this impression" was a repeated plot point in TNG. https://en.wikipedia.org/wiki/The_Measure_of_a_Man_(Star_Tre... https://en.wikipedia.org/wiki/The_Offspring_(Star_Trek:_The_... https://en.wikipedia.org/wiki/The_Ensigns_of_Command https://en.wikipedia.org/wiki/The_Schizoid_Man_(Star_Trek:_T... Simi…

The main computer does not make choices stochastically and always understands what people ask it. I do not think that resembles the current crop of LLMs. On voyager the ships computer is some kind of biological computing entity that they eventually give up on as a story topic but there is an episode where the bio computing gel packs get sick. I believe data and the doctor both would be people to me. But is minuet? Th…

> The main computer does not make choices stochastically and always understands what people ask it.

The mechanism is never explained, but no, it doesn't always understand correctly — and neither does Data. If hologram-Moriarty is sentient (is he?), then the capability likely exceeds what current LLMs can do, but the cause of the creation is definitely a misunderstanding.

Even the episode where that happens, the script for Dr. Pulaski leading up to Moriarty's IQ boost was exactly the same arguments used against LLMs: https://www.youtube.com/watch?v=4pYDy7vsCj8

(Common trope in that era being that computers (including Data) are too literal, so there was also: https://www.youtube.com/watch?v=HiIlJaSDPaA)

Similar with every time the crew work iteratively to create something in the holodeck. And, of course: https://www.youtube.com/watch?v=srO9D8B6dH4

> I do not think that resembles the current crop of LLMs. On voyager the ships computer is some kind of biological computing entity that they eventually give up on as a story topic but there is an episode where the bio computing gel packs get sick.

"Take the cheese to sickbay" is one of my favourite lines from that series.

> But is minuet?

I would say the character was a puppet, with the Bynars pulling the strings, because the holo-character was immediately seen as lacking personhood the moment they stopped fiddling with the computer.

Vic Fontaine was more ambiguous in that regard. Knew he was "a lightbulb", but (acted like) he wanted to remain within that reality in a way that to me felt like he was *programmed* to respond as if the sim around him was the only reality that mattered rather than having free will in that regard.

(But who has total free will? Humans are to holograms as Q is to humans, and the main cast were also written to reject "gifts" from Riker that time he briefly became a Q).

The villagers of Fair Haven were, I think, not supposed to be sentient (from the POV of the crew), but were from the POV of the writers: https://en.wikipedia.org/wiki/Fair_Haven_(Star_Trek:_Voyager... and https://en.wikipedia.org/wiki/Spirit_Folk_(Star_Trek:_Voyage...

> does eqtransformer have awareness?

There's too many different definitions for a single answer.

We don't know what part of our own brains gives us the sensation of our own existence; and even if we did, we wouldn't know if it was the only mechanism to do so.

To paraphrase your own words:

At what point does chemical pipelines doing some kind of stochastic transformation and electrochemical integration of sensory input become an individual that presents a desire for autonomy like data or the doctor?

I don't know. Like you, I'd say:

> I think there’s lots of questions here to answer and I don’t know the answers to them.

Post reply on HN