Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
311–320 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#312Earlier quoted context omitted.
Chinese companies smuggling embargo'ed/controlled GPUs and using OpenAI outputs violating their ToS is considered cheating. As I see it, this criticism comes from a fear of USA losing its first mover advantage as a nation. PS: I'm not criticizing them for it nor do I really care if they cheat as long as prices go down. I'm just observing and pointing out what other posters are saying. For me if China cheating means t…
> using OpenAI outputs violating their ToS is considered cheating I fail to see how that is any different than any other training data scraped from the web. If someone shares a big dump of outputs from OpenAI models and I train my model on that then I'm not violating OpenAI's terms of service because I haven't agreed to them (so I'm not violating contract law), and everyone in the space (including OpenAI themselves)…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#313Earlier quoted context omitted.
There's both. With the web interface it clearly has stopwords or similar. If you run it locally and ask about e.g. Tienanmen square, the cultural revolution or Winnie-the-Pooh in China, it gives a canned response to talk about something else, with an empty CoT. But usually if you just ask the question again it starts to output things in the CoT, often with something like "I have to be very sensitive about this subjec…
This is super interesting. I am not an expert on the training: can you clarify how/when the censorship is "baked" in? Like is the a human supervised dataset and there is a reward for the model conforming to these censored answers?
There are multiple ways to do this: humans rating answers (e.g. Reinforcement Learning from Human Feedback, Direct Preference Optimization), humans giving example answers (Supervised Fine-Tuning) and other prespecified models ranking and/or giving examples and/or extra context (e.g. Antropic's "Constitutional AI").
For the leading models it's probably mix of those all, but this finetuning step is not usually very well documented.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#314we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#315 “OpenAI stole from the whole internet to make itself richer, DeepSeek stole from them and give it back to the masses for free I think there is a certain british folktale about this”Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#316I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.
Great as long as you’re not interested in Tiananmen Square or the Uighurs.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#317Earlier quoted context omitted.
Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.
I guess all that leetcoding and stack ranking didn't in fact produce "the cream of the crop"...
At least engineers have some code to show for, unlike managerial class...
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#318Genuinly curious, what is everyone using reasoning models for? (R1/o1/o3)
In my experience GPT is still the number one for code, but Deepseek is not that far away. I haven't used it much for the moment, but after a thousand coding queries i hope to have a much better picture of it's coding abilities. Really curious about that, but GPT is hard to beat.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#319Earlier quoted context omitted.
> What was the Tianamen Square Massacre? > I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. hilarious and scary
It may be due to their chat interface than in the model or their system prompt, as kagi's r1 answers it with no problems. Or maybe it is because of adding the web results. https://kagi.com/assistant/98679e9e-f164-4552-84c4-ed984f570... edit: it is due to adding the web results or sth about searching the internet vs answering on its own, as without internet access it refuses to answer https://kagi.com/assistant/3ef6d8…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#320I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…
Agree. The "need to build new buildings, new power plants, buy huge numbers of today's chips from one vendor" never made any sense considering we don't know what would be done in those buildings in 5 years when they're ready.
As AI or whatever gains more capability, I'm sure it will do more useful things, but I just see it displacing more non-physical jobs, and now will expand the reach of individual programmers, removing some white color jobs (hardly anyone uses an agent to buy their ticket), but that will result is less need for programmers. Less secretaries, even less humans doing actual tech support.
This just feels like radio stocks in the great depression in the us.