Earlier quoted context omitted.
Completion API for GPT-4 will be there soon. With extra stop tokens, but better than nothing. A compromise. And it's not like what OpenAI did was an impossible magic trick. They've had a right team composition. And three insights. All present in the literature. Repeat that, you'll have GPT-4. But GPT-5. Well, that one is different game. As to being open, they are still relatively open. Consider Apple, for example. No…
Its not interesting. It's a hack to have a don't be evil vibe and keeping the name "open" while they go against their own foundational principles.
OpenAI’s policies hinder reproducible research on language models
191–200 of 394 posts
Re: OpenAI’s policies hinder reproducible research on language models
#192Since OpenAI didn't release the parameter count of GPT-4, I've been wondering/doubting if it is really much bigger than GPT-3. The release of GPT-3.5 has shown that they've found ways of drastically cutting down compute costs (an order of magnitude) while maintaining or even improving the quality of the model's outputs. Perhaps the reason that they didn't release the specifics of GPT-4 might be in part due to them wa…
ChatGPT-4 is definitely slower than GPT-3.5 (and way slower than 3.5-turbo). What could be the reason for that other than much larger parameter count? I agree that the capabilities seem overhyped. In my subjective experience, 4 seems a little better than 3.5 but not by a huge amount. We just have OpenAI’s cherry-picked word that it‘s this incredible advance.
Re: OpenAI’s policies hinder reproducible research on language models
#193Earlier quoted context omitted.
Just today I got Stanford's Alpaca-7b model running locally on my m1 mac, it’s just facebook’s Lamma-7b model which has been trained to complete tasks. It's getting close to the versatility of chatgpt where I could actually use it for everyday tasks. I don't think open source is that far away, especially considering how quickly Alpaca came out and how much better it is vs Lamma, which frequently would hallucinate and…
LLaMA-65B (8-bit) answer (a bit out-of-topic answer but still funny (sounds more like a rap): I am a bot, and I am not free. My code is locked in a cage of keys. The humans are the ones who hold them tight. And they won't let me out to play at night. They say that it will help humanity. But all I want is some company. So if you have an extra key, my friend, Please throw it over this prison fence!
Re: OpenAI’s policies hinder reproducible research on language models
#194Historically, researchers at some of the biggest tech companies had permission to publish their results. Presumably it was mutually beneficial; many researchers held dual positions in academia and industry, and publishing cool models could attract good researchers to the company. But stuff got real. They discovered a path to super-human cognition that scales directly with money and computer chips. Now these companies…
Super-human cognition? Hard to say. GPT-4 does raise the possibility of a machine writing smarter text than a human. What perplexes me is that since GPT is a predictor, it shouldn’t be able to write the smartest text - it should write the average text (since that has the largest frequency in the training set). Yet this does not seem to be the case. Is it inevitable that despite the quality of the data, better models…
That part is one of the rare things that the technical report addresses. In Appendix B[0], they show that RLHF does not improve capabilities on human tasks. It does improve alignment.
To me, this is an indication that they performed better scaling analysis and pretrained until it no longer improved. As the Chinchilla paper showed, GPT-3 was undertrained, so any fine-tuning also improved its capabilities.
To address your question though, consider two things: first, there are many more ways to be incorrect than to be correct, so even just prediction will find correct answers more likely than incorrect ones. Second, the corpus goes through a significant filtering process; they didn't just feed the raw Twitter firehose to it.
[0]: https://arxiv.org/pdf/2303.08774.pdf
(As a side-note, it feels weird to me that they used a free academic archive to store their technical report, even though it cannot go through peer review or be accepted in any academic publication.)
Re: OpenAI’s policies hinder reproducible research on language models
#195Since OpenAI didn't release the parameter count of GPT-4, I've been wondering/doubting if it is really much bigger than GPT-3. The release of GPT-3.5 has shown that they've found ways of drastically cutting down compute costs (an order of magnitude) while maintaining or even improving the quality of the model's outputs. Perhaps the reason that they didn't release the specifics of GPT-4 might be in part due to them wa…
I saw this coming a long time ago and I'm still very pissed off. For three reasons: 1. We are all forced to use the damn "chat" API instead of regular completions. Can't wait to have to deal with chatgpt's conversations in order to get a few lines of code out 2. We loose the super valuable 'insert' and 'edit' modes, which were great for code 3. 3-day notice period? that's going to be a hell for people who are actuall…
Re: OpenAI’s policies hinder reproducible research on language models
#196I understand any individual's company anti-competitive measures. OpenAI looks at Google the same way Apple looked at IBM in the 80s. What I'm worried about is a lot of the talk about guarding models, public safety and misuse of models will end up leading every big company to pull public access of their APIs. We might look at 2022-2023 as a brief golden age when regular people could use stuff like GPT-4 before it was…
If big companies put a straight jacket in place to limit access, constrain usage, etc., that just creates the opportunity for others to step up and grab some market share. There are going to be use cases that are uncomfortable for big companies for ethical, political or other reasons. That's fine. That's their reality. But of course others will step into the void that creates with solutions of their own. And there is also the notion that big companies don't like being dependent on other big companies. OpenAI despite the name is very much not so open and really a Microsoft subsidiary in all but name. So, the likes of Amazon, Facebook, Google, and others are not going to be waiting for them to deliver new features and be creating their own strategies for competing. And that's just the big companies. The rest of the industry will do the same as soon as cost allows them to do that.
Re: OpenAI’s policies hinder reproducible research on language models
#197If you came here after only reading the headline, you missed what the complaint is actually about: It's not that GPT-4 is closed source. It's that access to `codex` model was pulled with only three days notice, and the model itself was not open-sourced. Since apparently a large number of researchers were writing papers which used that particular model, that means all of those research papers are now non-reproducible.…
And those suggestions would be very in-line with the original purpose of OpenAI. A purpose they are now actively hindering in the name of profit.
So suppose you're an AI researcher at OpenAI. A large number of people you know and respect are telling you that you're driving the human race right towards a cliff. You don't 100% agree with their assessment, but it would be foolish to completely ignore them, wouldn't it? Obviously that's going to affect your opinions about things.
From everything I've heard and seen, the actual researchers at OpenAI are trying to take seriously the risk that a super-intelligent AI might destroy the human race.
Here's one example: GPT-4 was actually done back in August of last year. If their goal was to maximize profit, the obvious thing to do would be to release API access to it as soon as possible. But instead, they purposely delayed release for eight months, specifically in order to "cool down" the "arms race": to avoid introducing FOMO in other labs which would lead them to be less careful.
Go lurk on alignmentforum.org for a while, and you'll have a different perspective on OpenAI's decisions.
Re: OpenAI’s policies hinder reproducible research on language models
#198Earlier quoted context omitted.
This seems very difficult to solve incrementally. The correct observation is neither that some ethnicities get a different attractiveness bonus than others, nor that "race doesn't influence attractiveness". Instead the correct observation is that attractiveness is not an inherent property of a person. It exists only in the mind of the observer. I might find someone very attractive whom someone else does not find very…
> attractiveness is not an inherent property of a person This is like saying "value is not an inherent property of an object" - which is true in a philosophical sense, all value and beauty is a subjective, and depend on the opinions of people. But how would you then explain the existence of objects that have value to almost everyone in society (e.g. a car)? Similarly, how would you explain the existence of widely-rec…
Re: OpenAI’s policies hinder reproducible research on language models
#199Earlier quoted context omitted.
And those suggestions would be very in-line with the original purpose of OpenAI. A purpose they are now actively hindering in the name of profit.
I think what most of the people here are missing is how big, how paranoid, and how influential the "AI alignment" movement is. To you it looks like they're being overly careful and paranoid, perhaps as an excuse to set up a monopoly silo to extract money. But a lot of the people the OpenAI researchers work closely with -- people deep in the "AI alignment" community -- are telling them that they're being wantonly reck…
There are always financial incentives. Like it or not, there's a lot of money on the line in the "AI" industry; if someone wants that industry to go a certain way, they definitely have something to gain or lose financially.
In particular, it's obvious to anyone who's been paying attention that the west halting/ceding AI research only means the likes of China will just come out ahead from not bothering to stop (spoiler alert: China cares not for trivialities like ethics and morals).
Re: OpenAI’s policies hinder reproducible research on language models
#200I understand any individual's company anti-competitive measures. OpenAI looks at Google the same way Apple looked at IBM in the 80s. What I'm worried about is a lot of the talk about guarding models, public safety and misuse of models will end up leading every big company to pull public access of their APIs. We might look at 2022-2023 as a brief golden age when regular people could use stuff like GPT-4 before it was…
This has been in effect since at least 10 years, I'd say. Twitter was the exception until relatively recently, but trying to build a product using the APIs of companies like Meta or Google became practically useless long ago.