Live data from Hacker News

S1: A $6 R1 competitor?

timkellogg.me

111–120 of 430 posts

Re: S1: A $6 R1 competitor?

#111
post #53

If chain of thought acts as a scratch buffer by providing the model more temporary "layers" to process the text, I wonder if making this buffer a separate context with its own separate FNN and attention would make sense; in essence, there's a macroprocess of "reasoning" that takes unbounded time to complete, and then there's a microprocess of describing this incomprehensible stream of embedding vectors in natural lan…

Here's a paper your idea reminds me of. https://arxiv.org/abs/2501.19201 It's also so not far from Meta's large concept model idea.

Previous discussion:

[41 comments, 166 points] https://news.ycombinator.com/item?id=42919597

Re: S1: A $6 R1 competitor?

#112

Earlier quoted context omitted.

Or there is no objective reality (well there isn’t, check out the study), and reality is just a rendering of the few state variables that keep track of your simple life. A little context about you: - person - has hands, reads HN These few state variables are enough to generate a believable enough frame in your rendering. If the rendering doesn’t look believable to you, you modify state variables to make the render mo…

Is this a clever rhetorical trick to make it appear that your prior claim was correct? If not: what am I intended to take away from this? What is its relevance to my comment?

You made a joke about questioning reality, I simply entertained it. You can do whatever you want with it, wasn’t a slight at all.

Re: S1: A $6 R1 competitor?

#113
post #53

If chain of thought acts as a scratch buffer by providing the model more temporary "layers" to process the text, I wonder if making this buffer a separate context with its own separate FNN and attention would make sense; in essence, there's a macroprocess of "reasoning" that takes unbounded time to complete, and then there's a microprocess of describing this incomprehensible stream of embedding vectors in natural lan…

Once we train models on the chain of thought outputs, next token prediction can solve the halting problem for us (eg, this chain of thinking matches this other chain of thinking).

Re: S1: A $6 R1 competitor?

#114

Earlier quoted context omitted.

You can choose to be somewhat ignorant of the current state in AI, about which I could also agree that at certain moments it appears totally overhyped, but the reality is that there hasn't been a bigger technology breakthrough probably in the last ~30 years. This is not "just" machine learning because we have never been able to do things which we are today and this is not only the result of better hardware. Better ha…

> the first one being from DeepMind in 2017 ? what paper are you talking about

https://arxiv.org/abs/1706.03762

Re: S1: A $6 R1 competitor?

#115
post #87

Earlier quoted context omitted.

Year over year gains in computing continue to slow. I think we keep forgetting that when talking about these things as assets. The thing controlling their value is the supply which is tightly controlled like diamonds.

They have a fairly limited lifetime even if progress stands still.

Last I checked AWS 1-year reserve pricing for an 8x H100 box more than pays for the capital cost of the whole box, power, and NVIDIA enterprise license, with thousands left over for profit. On demand pricing is even worse. For cloud providers these things pay for themselves quickly and print cash afterwards. Even the bargain basement $2/GPU/hour pays it off in under two years.

Re: S1: A $6 R1 competitor?

#118
post #37

Earlier quoted context omitted.

Agreed. I was working on some haiku things with ChatGPT and it kept telling me that busy has only one syllable. This is a trivially searchable fact.

link a chat please

It wasn't just busy that it failed on. I was feeding it haikus and wanted them broken into a list of 17 words/fragments. Certain 2 syllable words weren't split and certain 1 syllable words were split into two.
Post reply on HN