Earlier quoted context omitted.
...no? They also use web crawlers.
The datasets are collected using web crawlers, but that doesn’t tell us anything about how they are stored and re-distributed, right?
Deepseek R1-0528
241–250 of 264 posts
Re: Deepseek R1-0528
#242Earlier quoted context omitted.
The datasets are collected using web crawlers, but that doesn’t tell us anything about how they are stored and re-distributed, right?
Why would you store the data after training?
I’d just assume they did because—why scrape again if you want to train a new model? But if you know otherwise, I’m not tied to this idea.
Re: Deepseek R1-0528
#243Earlier quoted context omitted.
Isn't it basically not possible for the input data set list to be listed? It's an open secret all these labs are using immense amounts of copyrighted material. There's a few efforts at full open data / open weight / open code models, but none of them have gotten to leading-edge performance.
My brain was largely trained using immense amounts of copyrighted material as well. Some of it I can even regurgitate almost exactly. I could list the names of many of the copyrighted works I have read/watched/listened to. I suppose my brain isn't open source, although I don't think it would currently be illegal to take a snapshot of my brain and publish it if the technology existed and open-source that. Granted, thi…
However, once these words are broadcast—once they’re read, and the ideas expressed here enter someone else’s mind—I believe it’s only fair that the person on the receiving end has the right to use, replicate, or create something from them. After all, they lent me their brain—ideas that originated in my mind now live in theirs.
This uses up their mental "meat space," their blood sugar, and their oxygen—resources they provide. So, they have rights too: the right to do as they please with those ideas, including creating any and all data derived from them. Denying them that right feels churlish, as if it isn’t the most natural thing in the world.
(Before people jump on me:- Yes, creators need to be compensated—they deserve to make a living from their work. But this doesn’t extend to their grandchildren. Copyright laws should incentivize creation, not provide luxury for the descendants of the original creator a century later.)
Re: Deepseek R1-0528
#244Earlier quoted context omitted.
But what do you do with these secrets? Like tagging emails, summarizing documents?
a document management system is an easy example. Let’s say medical, legal, and tax documents.
Re: Deepseek R1-0528
#245Earlier quoted context omitted.
My brain was largely trained using immense amounts of copyrighted material as well. Some of it I can even regurgitate almost exactly. I could list the names of many of the copyrighted works I have read/watched/listened to. I suppose my brain isn't open source, although I don't think it would currently be illegal to take a snapshot of my brain and publish it if the technology existed and open-source that. Granted, thi…
> Some of it I can even regurgitate almost exactly If you (or any human) violate copyright law, legal redress can be sought. The amount of damage you can do is limited because there's only one of you vs the marginal cost of duplicating AI instances. There are many other differences between humans and AI in terms of capabilities and motivations to f the legal persons making decisions.
But enough about whether it should be legal to own a Xerox machine. It's what you do with the machine that matters.
Re: Deepseek R1-0528
#246Well that didn't take long, available from 7 providers through openrouter. https://openrouter.ai/deepseek/deepseek-r1-0528/providers May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass. Fully open-source model.
Is there a downloadable model? (Not familiar with openrouter and not seeing the model on ollama.)
> ollama run deepseek-r1
Re: Deepseek R1-0528
#247Earlier quoted context omitted.
> But those benchmarks are in general fairly narrow. They don't really measure the "broader" intelligence we are after. I think a general model that can - finish nethack, doom, zelda and civilization, - solve the hardest codeforces/atcoder problems, - formally prove putnam solution with high probability, not given the answer - write a PR to close a random issue on github is likely to have some broader intelligence. I…
A couple of things: I wasn't trying to invent anything. Just describing what you would obviously have to do if you were to take a "scientific" or "objective" approach: Sound experiments, reproducible, free of financial incentives. As far as I can tell, no one is doing that at a significant scale. Everything is buried in hype and marketing. Now for that broad set of benchmarks (PRs to GitHub, Putnam, Zelda). There is…
No, that's not actually a good description of the mixture-of-experts methodology. It was poorly named. There is no conscious division of the weights into "This subset is good for poetry, this one is best for programming, this one for math, this one for games, this one for language translation, etc."
Re: Deepseek R1-0528
#248Earlier quoted context omitted.
a document management system is an easy example. Let’s say medical, legal, and tax documents.
Thank you, but what do you use the llm for? Writing new documents based on previous ones? Tagging/categorization/summarization/lookup? RAG? Extracting structured data from them?
i use ollama to generate a document title, with 8 words or less. I then go through and make any manual edits at my leisure. Saves me time which i appreciate!
Paperless-ngx already does a pretty good job auto-tagging, i think it uses some built in classifiers? not 100% sure.
Re: Deepseek R1-0528
#24974% smaller 713GB to 185GB.
Use the magic incantation -ot ".ffn_.*_exps.=CPU" to offload MoE layers to RAM, allowing non MoEs to fit < 24GB VRAM on 16K context! The rest sits in RAM & disk.
Re: Deepseek R1-0528
#250Earlier quoted context omitted.
> Some of it I can even regurgitate almost exactly If you (or any human) violate copyright law, legal redress can be sought. The amount of damage you can do is limited because there's only one of you vs the marginal cost of duplicating AI instances. There are many other differences between humans and AI in terms of capabilities and motivations to f the legal persons making decisions.
The amount of damage you can do is limited because there's only one of you vs the marginal cost of duplicating AI instances But enough about whether it should be legal to own a Xerox machine. It's what you do with the machine that matters.
The capabilities of a machine matter a lot under law. See current US gun legislation[1], or laws banning export of dual-use technology for examples of laws that have inherent capabilities - not just the use of the thing- as core considerations.
1. It's illegal to possess a new, automatic weapon with some grandfathering prior to 1986