Live data from Hacker News

DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

arxiv.org

11–20 of 37 posts

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#11
post #8
post #2

Not jus being impressed that every paper coming out is SOTA, but also leads the way in being Open-Source in the pure definition of OSS, even with permissible licensing. Let's not confuse the company with the country by over-fitting a narrative. Popular media is reenforcing hatred or anything that sponsors them, especially to weaker groups. Less repercussions and more clicks/money to be made I guess. While Politicians…

> Let's not confuse the company with the country What's wrong with China? They're wonderful in the OSS ecosystem.

It varies on a company to company basis. BOOX, for instance, are notorious GPL violators.

There's also significant alpha in releasing open weights models. You get to slow down the market leaders to make sure they don't have runaway success. It reduces moats, slows funding, creates a wealth of competition, reduces margin. It's a really smart move if you want to make sure there's a future where you can compete with Google, OpenAI, etc. There's even a chance it makes those companies bleed a little. The value chain moves to differently shaped companies (tools, infra) leaving space for consumer and product to not necessarily be won by the "labs" companies.

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#12
post #6
post #3

DeepSeek R1 is by far the best at writing prose of any model, including Grok-3, GPT-4o, o1-pro, o3, claude, etc. Paste in a snippet from a book and ask the model to continue the story in the style of the snippet. It's surprising how bad most of the models are. Grok-3 comes in a close second, likely because it is actually DeepSeek R1 with a few mods behind the scenes.

why do you think that grok 3 is deepseek, out of curiosity?

Yes that’s a pretty giant accusation, especially given they’re buying boatloads of GPUs and have previous versions as well (it’s not like they’re starting with 3).

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#13
post #2

Not jus being impressed that every paper coming out is SOTA, but also leads the way in being Open-Source in the pure definition of OSS, even with permissible licensing. Let's not confuse the company with the country by over-fitting a narrative. Popular media is reenforcing hatred or anything that sponsors them, especially to weaker groups. Less repercussions and more clicks/money to be made I guess. While Politicians…

I love open source and the general vibe of good vibes you're bringing, but...this isn't SOTA, or close, even on the papers own terms. (i.e. excluding models released the last 6 months, including their own, which is a strange, yet understandable, choice given the results they report)

Quickest way to show this:

- Table 2, top of page 7

- Gemma 2 27B, 0 interventions, has 94.1/56.6/60.2

- Gemma 2 27B, with all their interventions, has 86/64/69.

- Gemma 2 27B, with all their interventions, sampled 32 times, is at 90.4/67.2/70.3.

- Gemma 2 27B came out in...June 2024. :/

Quick heuristics employed here:

- What models did they compare against? (this isn't strictly an issue, the big screaming tell is "What models did they compare against compared to their last N papers?"

- How quickly does the paper have to move towards N samples, and how big does N get before they're happy enough to conclude? (32). How much does that improve performance on their chosen metric? (1.8%)

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#15

[dead]

> The idea of role-playing as different characters is novel.

It is not. I remember Karpathy being really excited about the "1 million gpt personas" dataset and highlighted it as a way to avoid reward hacking in RLAIF. That was 3-6 months ago I believe.

Of course paper / code / weights beats idea, and it's exciting to see how far this can go.

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#16
post #7

Earlier quoted context omitted.

When OpenAI surged ahead Meta ended up giving away its incredibly expensive to make llama model to reduce the OpenAI valuations. Is DeepSeeks openness in part to reduce the big American tech companies?

Correlation isn't causation, I hate to say this, but here's really applicable. Facebook aka Meta has always been very opensource. Let's not talk about the license though. :) Why do you imply malice in OSS companies? Or for profit companies opensourcing their models and sourcecode?

Meta is decidedly not an "OSS company" no matter how much they put out.

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#17
post #7

Earlier quoted context omitted.

Correlation isn't causation, I hate to say this, but here's really applicable. Facebook aka Meta has always been very opensource. Let's not talk about the license though. :) Why do you imply malice in OSS companies? Or for profit companies opensourcing their models and sourcecode?

Meta is decidedly not an "OSS company" no matter how much they put out.

In this case there are very few truly "OSS companies" except for Red Hat and few other Linux distribution maintainers. Even companies centered around open source like Gitlab are usually generate most of their revenue of proprietary products or use liceses like BSL.

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#19
post #6

Earlier quoted context omitted.

why do you think that grok 3 is deepseek, out of curiosity?

Yes that’s a pretty giant accusation, especially given they’re buying boatloads of GPUs and have previous versions as well (it’s not like they’re starting with 3).

1) Grok-2 was akin to GPT-3.5

2) Grok-3 comes out a month after DeepSeek R1 was open sourced. I think Grok-3 is DeepSeek R1 with some added params and about a month of training on the giant cluster, possibly a bit of in-house secret sauce added to the model or training methodology.

What are the chances that XAI just happened to have a thinking model close to as good as revolutionary DeepSeek but happened to launch it 30 days later?

It was both smart and pragmatic for XAI to simply use the best available open source stuff and layer their own stuff on top of it. Imagine they doubled the parameter count and trained it for 30 days, that would not even use half of the GPU power!

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#20
post #6
post #3

DeepSeek R1 is by far the best at writing prose of any model, including Grok-3, GPT-4o, o1-pro, o3, claude, etc. Paste in a snippet from a book and ask the model to continue the story in the style of the snippet. It's surprising how bad most of the models are. Grok-3 comes in a close second, likely because it is actually DeepSeek R1 with a few mods behind the scenes.

why do you think that grok 3 is deepseek, out of curiosity?

I replied to the child of your comment
Post reply on HN