Live data from Hacker News

S1: A $6 R1 competitor?

timkellogg.me

181–190 of 430 posts

Re: S1: A $6 R1 competitor?

#181
Anyone else wants more articles on how those benchmarks are created and how they work?

Those models can be trained in way tailored to have good results on specific benchmarks, making them way less general than it seems. No accusation from me, but I'm skeptical on all the recent so called 'breakthroughs'.

Re: S1: A $6 R1 competitor?

#182

Earlier quoted context omitted.

oh god, this is terrible! I just said "Hello!" and it went off the rails.

why how what? can you add a sample prompt with output ?

In my original prompt, it gave me back multiple greetings in various languages. In this example, it gave me back just multiple greetings:

>>> hello!

Hello there! I'm here to help with any questions you might have. How can I assist you today? If you're asking a question, feel free to provide some details about what you need help with. I'll do my best to answer your question or guide you through the process. Just let me know!

think

The user has asked "hello! " and provided no further information. Answer here is a direct response to the user's request for information, which was essentially just a greeting. The appropriate response is to greet back and offer assistance.

answer

Answer:

Hello! How can I assist you today? If you have a question or need help with something specific, feel free to ask. I'm here to support you!

think

answer

The user's message was "hello!" and there was no other information provided. The appropriate response is a greeting back and an offer of assistance.

answer

Answer:

Hello! How can I assist you today? If you have a question or need help with something specific, feel free to ask. I'm here to support you!

think

The user has greeted with "hello!" An appropriate response involves acknowledging the greeting and offering assistance. This could involve asking how we can help them or what they'd like guidance on.

answer

Answer:

Hello! How can I assist you today? If you have a question or need help with something specific, feel free to ask. I'm here to support you!

Re: S1: A $6 R1 competitor?

#183
post #56

Earlier quoted context omitted.

This is pure speculation on my part but I think at some point a company's valuation became tied to how big their compute is so everybody jumped on the bandwagon.

So, "No one was ever fired for ... buying more server infrastructure."

Walmart has massive, idle datacenters full of running machines doing nothing.

Re: S1: A $6 R1 competitor?

#184

Earlier quoted context omitted.

Is this a clever rhetorical trick to make it appear that your prior claim was correct? If not: what am I intended to take away from this? What is its relevance to my comment?

You made a joke about questioning reality, I simply entertained it. You can do whatever you want with it, wasn’t a slight at all.

It may have been in the form of a joke, but I certainly wasn't joking.

I think it is interesting what actions cannot be done by humans.

Re: S1: A $6 R1 competitor?

#185
post #133

I found the discussion around inference scaling with the 'Wait' hack so surreal. The fact such an ingeniously simple method can impact performance makes me wonder how many low-hanging fruit we're still missing. So weird to think that improvements on a branch of computer science is boiling down to conjuring the right incantation words, how you even change your mindset to start thinking this way?

I've noticed that R1 says "Wait," a lot in its reasoning. I wonder if there's something inherently special in that token.

Re: S1: A $6 R1 competitor?

#186

I have a bunch of questions, would love for anyone to explain these basics: * The $5M DeepSeek-R1 (and now this cheap $6 R1) are both based on very expensive oracles (if we believe DeepSeek-R1 queried OpenAI's model). If these are improvements on existing models, why is this being reported as decimating training costs? Isn't fine-tuning already a cheap way to optimize? (maybe not as effective, but still) * The R1 pap…

If what you say is true, and distilling LLMs is easy and cheap, and pushing the SOTA without a better model to rely on is dang hard and expensive, then that means the economics of LLM development might not be attractive to investors - spending billions to have your competitors come out with products that are 99% as good, and cost them pennies to train, does not sound like a good business strategy.

Re: S1: A $6 R1 competitor?

#187
post #133

I found the discussion around inference scaling with the 'Wait' hack so surreal. The fact such an ingeniously simple method can impact performance makes me wonder how many low-hanging fruit we're still missing. So weird to think that improvements on a branch of computer science is boiling down to conjuring the right incantation words, how you even change your mindset to start thinking this way?

I've noticed that R1 says "Wait," a lot in its reasoning. I wonder if there's something inherently special in that token.

Semantically, wait is a bit of a stop-and-breathe point.

Consider the text:

I think I'll go swimming today. Wait, ___

what comes next? Well, not something that would usually follow without the word "wait", probably something entirely orthogonal that impacts the earlier sentence in some fundamental way, like:

Wait, I need to help my dad.

Re: S1: A $6 R1 competitor?

#188
post #139
post #56

Earlier quoted context omitted.

This is pure speculation on my part but I think at some point a company's valuation became tied to how big their compute is so everybody jumped on the bandwagon.

I don't think you need to speculate too hard. On CNBC they are not tracking revenue, profits or technical breakthroughs, but how much the big companies are spending (on gpus). That's the metric!

They absolutely are tracking revenues/profits on CNBC, what are you talking about?

Re: S1: A $6 R1 competitor?

#189

I think a lot of people in the ML community were excited for Noam Brown to lead the O series at OpenAI because intuitively, a lot of reasoning problems are highly nonlinear i.e. they have a tree-like structure. So some kind of MCTS would work well. O1/O3 don’t seem to use this, and DeepSeek explicitly mentioned difficulties training such a model. However, I think this is coming. DeepSeek mentioned it was hard to lear…

Do you have a reference for us to check? - "DeepSeek explicitly mentioned difficulties training such a model."

Section 4.2: Unsuccessful attempts

https://arxiv.org/pdf/2501.12948

Re: S1: A $6 R1 competitor?

#190
> having 10,000 H100s just means that you can do 625 times more experiments than s1 did

The larger the organisation, the less experiments you can afford to do. Employees are mostly incentivised by getting something done quick enough to not to be fired in this job market. They know that the higher-ups would get them off for temporary gains. Rush this deadline, ship that feature, produce something that looks OK enough.

Post reply on HN