Live data from Hacker News

Deepseek R1-0528

huggingface.co

231–240 of 264 posts

Re: Deepseek R1-0528

#231
post #114

Earlier quoted context omitted.

Isn't it basically not possible for the input data set list to be listed? It's an open secret all these labs are using immense amounts of copyrighted material. There's a few efforts at full open data / open weight / open code models, but none of them have gotten to leading-edge performance.

My brain was largely trained using immense amounts of copyrighted material as well. Some of it I can even regurgitate almost exactly. I could list the names of many of the copyrighted works I have read/watched/listened to. I suppose my brain isn't open source, although I don't think it would currently be illegal to take a snapshot of my brain and publish it if the technology existed and open-source that. Granted, thi…

> Some of it I can even regurgitate almost exactly

If you (or any human) violate copyright law, legal redress can be sought. The amount of damage you can do is limited because there's only one of you vs the marginal cost of duplicating AI instances.

There are many other differences between humans and AI in terms of capabilities and motivations to f the legal persons making decisions.

Re: Deepseek R1-0528

#232

Well that didn't take long, available from 7 providers through openrouter. https://openrouter.ai/deepseek/deepseek-r1-0528/providers May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass. Fully open-source model.

Open weights.

Re: Deepseek R1-0528

#233

Earlier quoted context omitted.

Aren't the datasets mostly shared in torrents? They probably won't bitrot for some time.

...no? They also use web crawlers.

The datasets are collected using web crawlers, but that doesn’t tell us anything about how they are stored and re-distributed, right?

Re: Deepseek R1-0528

#234

Earlier quoted context omitted.

Sure you can. It's often legally protected activity. You're just limited to distributing your modifications without the original work.

For some games maybe, but software often has a clause forbidding reverse engineering

ChatGPT says that such clauses are typically void in the EU, though they may apply in some cases in the US. Even in the US, the triennial DMCA rule-making has granted broader exemptions for good-faith security research every cycle since 2016.

https://chatgpt.com/share/6838c070-705c-8005-9a88-83c9a5550a...

Re: Deepseek R1-0528

#235
post #218

Earlier quoted context omitted.

Ah FireShip, I forgot that channel existed at all. I asked YouTube to not recommend that channel after every vaguely AI-related news was "BIG NEWS!!!", the videos were also thin on actual content, and there were repeated factual errors over multiple videos too. At that point, the only thing it's good for is to make yourself (falsely) feel like you're keeping up.

Fireship consistently makes some of the most entertaining tech content out there

Hey if you enjoy it, go for it. I used to like it a couple of years ago too, but I found that more and more lately, it was neither entertaining nor reliably informative. The jokes/memes were lazy and recycled a lot, the tech content was often poorly researched, and it started feeling like content produced for the sake of having content.

Re: Deepseek R1-0528

#236

Earlier quoted context omitted.

Anyone who does not want to leak their data? I am actually surprised that people are ok with trusting their secrets to a random foreign company.

But what do you do with these secrets? Like tagging emails, summarizing documents?

a document management system is an easy example. Let’s say medical, legal, and tax documents.

Re: Deepseek R1-0528

#237

Earlier quoted context omitted.

Anyone who does not want to leak their data? I am actually surprised that people are ok with trusting their secrets to a random foreign company.

No one cares about your 'secrets' as much as you think. They're only potentially valuable if you're doing unpatented research or they can tie them back to you as an individual. The rest is paranoia. Having said that, I'm paranoid too. But if I wasn't they'd have got me by now.

step back for a bit. some people actually work with sensitive documents as part of their JOB. Like accountants, lawyers, people in medical industry, etc.

Sending a document with a social security number to OpenAI is just a dumb idea. As an example.

Re: Deepseek R1-0528

#238

Earlier quoted context omitted.

Much more preferred to what OpenAI always did and Anthropic recently started doing. Just write some complicated narrative about how scary this new model is and how it tried to escape and deceive and hack the mainframe while telling the alignment operators bed time stories.

Really? I missed this. The new hype trick is implying the new LLM releases are almost AGI? Love it.

Anthropic "warned" Claude 4 is so smart that it will try to use the terminal (if using Claude Code) or any other tools available (depending on where you're invoking it from) to contact local authorities if you're doing something very immoral.

Re: Deepseek R1-0528

#239

Out of sheer curiosity: What’s required for the average Joe to use this, even at a glacial pace, in terms of hardware? Or is it even possible without using smart person magic to append enchanted numbers and make it smaller for us masses?

I'm using GPT4All with DeepSeek-R1-Distill-QWen-7B (which is not R1-0528) on a Ryzen 5 3600 with 32Gb ram.

With an average of 3.6 tokens/sec, answers usually take 150-200 seconds.

Re: Deepseek R1-0528

#240
post #140
post #94

Earlier quoted context omitted.

No it doesn't, it has exactly the same source, zero. It has more downloadable binary.

That’s the ‘source’ for what the model spits out though, if not the source for what spits out the model.

The "source" for something is all the stuff that makes you able to build and change that something. The source for a model is all the stuff that makes you able to train and change the model.

Just because the model produces stuff doesn't mean that's the model's source, just like the binary for a compiler isn't the compiler's source.

Post reply on HN