Live data from Hacker News

DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

arxiv.org

31–37 of 37 posts

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#31
post #2

Not jus being impressed that every paper coming out is SOTA, but also leads the way in being Open-Source in the pure definition of OSS, even with permissible licensing. Let's not confuse the company with the country by over-fitting a narrative. Popular media is reenforcing hatred or anything that sponsors them, especially to weaker groups. Less repercussions and more clicks/money to be made I guess. While Politicians…

I love open source and the general vibe of good vibes you're bringing, but...this isn't SOTA, or close, even on the papers own terms. (i.e. excluding models released the last 6 months, including their own, which is a strange, yet understandable, choice given the results they report) Quickest way to show this: - Table 2, top of page 7 - Gemma 2 27B, 0 interventions, has 94.1/56.6/60.2 - Gemma 2 27B, with all their int…

Good vibes, I mean yeah, we need more breakthroughs and AI isn't here to take our jobs if WE can own the AI too and not just a super-corperation.

I think what we are all really excited about having finally AI at home and being unchained and freed from a central SaaS controlling all the AI is ever going to tell you.

So, 6-7y ago google had these AI Chats internally and never intended to release it, a friendly googler told me.

Then ChatGPT came along and locked you into their SaaS. That was fantastic in the beginning, but the more you used the AI, the more you felt helpless, swound by anyone who may have access to an AI at OpenAI that is unfiltered and uses the full power of the model. Then came the jailbreak and accounts being banned for using it.

Then came the freedom by LLAMA and DeepSeek and waves of otheres. It rolled into your laptop real quick and this freedom is priceless! Something we should be really thankful for that it happened and support more OSS.

Google and Facebook would never share their trove of data with us ever and very few people have enough storage and compute to even attempt to replicate them. But their Data Dominance doesn't protect them anymore. Once the models became intelligent enough to slurp up large chunks of the web, they became a better search, a better teacher and a better experience than sponsored ads, with ads with internal google/bing products listed up, then SEO websites and somewhere hidden what we really were looking for. Or often.. just being deleted for copyright and other reasons.

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#32
post #21
post #18

Any idea why I lost interest in deep seek? I used it and grok3 a whole bunch when they first came out but now I’ve fallen back to Claude for everything.

For coding, I‘m finding Claude‘s responses most to the point and on-task. While many other models try to extrapolate or lecture or patronize. DeepSeek is pretty good though. Maybe it’s the high latency (probably due to prompt processing)?

I personally have no good reason why I don't always ask Claude and DeepSeek the same prompt.

Thinking about it more, I think a big part of it is that correct answers are not a limiting factor for me it feels like. Claude is good enough and it is more what to do with all these correct answers is my problem. I am also naturally biased to a model if paying for it.

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#33
post #26
post #3

DeepSeek R1 is by far the best at writing prose of any model, including Grok-3, GPT-4o, o1-pro, o3, claude, etc. Paste in a snippet from a book and ask the model to continue the story in the style of the snippet. It's surprising how bad most of the models are. Grok-3 comes in a close second, likely because it is actually DeepSeek R1 with a few mods behind the scenes.

If it was Elon is even more stupid than he lets on because DS3: 5M training run Grok3: 400M training run for 2% difference in the benchmarks.

They probably pulled the plug at the last minute to switch to DeepSeek.

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#34
post #18

Any idea why I lost interest in deep seek? I used it and grok3 a whole bunch when they first came out but now I’ve fallen back to Claude for everything.

Deepseek is super bad at personal advice. I can tell it was trained on an oddly stodgy data set. It gives advice that would suit a hyper conservative world view. Like IBM 1950s middle management training course level advice.

Gemma is by far the best at giving advice and planning ones days and life priorities. Not sure how to benchmark that.

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#35
post #19

Earlier quoted context omitted.

1) Grok-2 was akin to GPT-3.5 2) Grok-3 comes out a month after DeepSeek R1 was open sourced. I think Grok-3 is DeepSeek R1 with some added params and about a month of training on the giant cluster, possibly a bit of in-house secret sauce added to the model or training methodology. What are the chances that XAI just happened to have a thinking model close to as good as revolutionary DeepSeek but happened to launch it…

> What are the chances that XAI just happened to have a thinking model close to as good as revolutionary DeepSeek but happened to launch it 30 days later? Extremely, extremely good. That was in fact the real point of the deepseek paper - it was extremely cheap to turn a frontier(ish?) model into a reasoning model. There is nothing suspicious about this timeline from an ML Ops point of view. In fact DeepSeek themselve…

Perhaps Grok-3 used the reasoning methodology from DeepSeek more than the underlying model, but the similarity of Grok-3 results to DeepSeek suggests that XAI used more than that.

Re: DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

#37
post #8

Earlier quoted context omitted.

> Let's not confuse the company with the country What's wrong with China? They're wonderful in the OSS ecosystem.

I didn't want to be politically correct, but also not insensitive. Many countries produce great things, but if we measure these countries rigerously, just a few stand out. Unfortunately from here on it get's messy, political, unsubstantiated or backed by data that is inherently biased due to selection criteria and weight. It's very difficult to be truly unbiased and neutral and it's not my goal, I just think it's a c…

No one's saying that you should take an absolute standpoint. I'm just asking, what's wrong with China since you made that comment as if China is bad for some reason.
Post reply on HN