Earlier quoted context omitted.
So OpenAI should just give up and release all their trade secrets? To what purpose?
I dunno, maybe the Open part of OpenAI should hint at it. The problem people are having is that OpenAI marketed themselves as supposedly democratizing AI, but it does the opposite.
OpenAI’s policies hinder reproducible research on language models
331–340 of 394 posts
Re: OpenAI’s policies hinder reproducible research on language models
#332Earlier quoted context omitted.
Two notes: 1) less than half of Americans vote in each election (less than 63% if you restrict to the voting-age population, less than 70% if you apply the scummy rules that restrict to the voting-eligible population) And 2) it's a false dichotomy to say that US elections have ever been "whatever we have now VS communism". Maybe you could say socialism was on the ballot all those times Eugene Debs ran for the preside…
> it sounds like you would struggle to define communism if pressed Having read the Communist Manifesto, I think that description of me is both totally fair and would also apply to Karl Marx. Darn thing read like an unhinged run-on blog rant.
Re: OpenAI’s policies hinder reproducible research on language models
#333Earlier quoted context omitted.
> Here's one example: GPT-4 was actually done back in August of last year. If their goal was to maximize profit, the obvious thing to do would be to release API access to it as soon as possible. They did that, that's how Reid Hoffman got early access to write his book, that's how Microsoft got access to start working on Bing/ChatGPT4 for a cool $10 billion and that's how countless of other got early access. Got the m…
Besides wasn't releasing GPT3 supposed to have caused major harm to society? Which is why they held off for so long. Still waiting for evidence of that harm (mass fake news, Google being ruined by even more low ranking spam sites, etc). It must be nice thinking that a small group withholding the keys R&D (for a short while until other R&D groups catch up) will somehow help the problem. Do these few months to a year r…
Re: OpenAI’s policies hinder reproducible research on language models
#334Earlier quoted context omitted.
Research into systemically important infrastructure cannot be damned because that infrastructure isn't public. It's a cheap moralizing argument to say "pfff, this was predictable". Maybe so, but there isn't an alternative. Much like research on Twitter. Once these companies start to drift into providing what become broadscale social utilities and public services it doesn't matter that they're private. There are(/shou…
Replying to both responses because they're all good points. My argument boils down to the fact that some private companies end up becoming social utilities and once that happens, the rules (should) change as part of the social contract which means, yeah, they can't simply "pull the rug". The research is important precisely because its into systemically significant systems. I get that it's difficult to define the line…
That said, what if OpenAI shut down codex because it has dangerous possibilities and amoral “researchers” started figuring out how to exploit them? What if it was fundamentally buggy or encouraging misleading research? What if codex was accidentally leaking or distributing export-controlled or other illegal (copyright, etc.) information? I’m explicitly speculating on possibilities, while you’re making unstated assumptions, so entertain the question of whether OpenAI is already doing a public service by shutting it down.
Re: OpenAI’s policies hinder reproducible research on language models
#335Earlier quoted context omitted.
I can't believe anyone considers a single number, which would work about equally well if it were 10% higher or lower, to be a trade secret.
This is not really true. The Chinchilla paper showed that a 4% difference in loss between Chinchilla and Gopher led Chinchilla to blow Gopher out of the water at most tasks, including 30x performance in physics. Empirically, LLMs have shown to have emergent abilities appear at different loss levels. So, a 10% difference could really matter.
Re: OpenAI’s policies hinder reproducible research on language models
#336Earlier quoted context omitted.
"GPT-3 175B model required 3.14E23 flops" according to their marketing material. Seti at home was about 1PetaFlops iirc so about 3 years training, possibly less if you can generate enough attention to the project that the people with the beefy devices will partecipate. The problem is that you need to train the full model you can't train aspect of it and even with each node doing independent tiny batches the network b…
How does something of this scale impact climate change? Like when there are 5-6-7 OpenAIs, what does that look like, is this just a huge amount of energy consumption ?
Re: OpenAI’s policies hinder reproducible research on language models
#337Earlier quoted context omitted.
The danger the AI alignment folk are afraid of is completely impossible with current tech, but they want to put up barriers because we have no idea what future tech might look like and there’s the possibility some future advance could be very dangerous. When anti-GMO or anti-nuclear folk used this same standard to put up barriers to research into nuclear or GMO research, they get lambasted for being anti-science, but…
The anti-gmo/nuclear people have no explanation for how things can go wrong. The AI alignment people do. You might not agree with it, but tons of AI researchers, including many at openAI, do.
Re: OpenAI’s policies hinder reproducible research on language models
#338Earlier quoted context omitted.
> We might look at 2022-2023 as a brief golden age when regular people could use stuff like GPT-4 Not sure about that since it seems to being baked into a lot of products at places like Microsoft. However, I'd change your statement a bit: We might look at 2023 as a brief golden age when regular people could access trained parameters (the LLaMA params) and run these models on their own machines (such as with alpaca.cp…
> However, I'd change your statement a bit: We might look at 2023 as a brief golden age when regular people could access trained parameters (the LLaMA params) and run these models on their own machines (such as with alpaca.cpp). I doubt we'll get access to LLM params again unless some kind of non-profit, actual open source organization is formed to produce them and put them out into the public domain. There are a lot…
The risk for these startups you describe as working on this as you type is the same thing happening to them that happened to Meta when they released their LLaMA params: they started getting copied all over the place. And it's not clear that Meta can do anything about this. It seems that params aren't copyrightable.
Re: OpenAI’s policies hinder reproducible research on language models
#339Earlier quoted context omitted.
I typed the query into chat-gpt3.5 (turbo and legacy), and 4, and they all said that there's 0.5 beb per bob. Did you use the quoted prompt exactly?
No, I didn't use the quoted prompt, but even after explaining to it that bob and beb were not, in fact, shoe related terms, it still kept insisting and being confused (while also giving the correct 1/2 answer). It can do it, but its not deterministic, and it doesnt really do it well. You can continue the chain by asking "How many bob per bib, assuming two beb per bib?", and see if it chokes then. It sometimes does, s…
Re: OpenAI’s policies hinder reproducible research on language models
#340Earlier quoted context omitted.
> it sounds like you would struggle to define communism if pressed Having read the Communist Manifesto, I think that description of me is both totally fair and would also apply to Karl Marx. Darn thing read like an unhinged run-on blog rant.
It’s literally a manifesto.