Live data from Hacker News

OpenAI’s policies hinder reproducible research on language models

aisnakeoil.substack.com

171–180 of 394 posts

Re: OpenAI’s policies hinder reproducible research on language models

#171

Earlier quoted context omitted.

I typed the query into chat-gpt3.5 (turbo and legacy), and 4, and they all said that there's 0.5 beb per bob. Did you use the quoted prompt exactly?

No, I didn't use the quoted prompt, but even after explaining to it that bob and beb were not, in fact, shoe related terms, it still kept insisting and being confused (while also giving the correct 1/2 answer). It can do it, but its not deterministic, and it doesnt really do it well. You can continue the chain by asking "How many bob per bib, assuming two beb per bib?", and see if it chokes then. It sometimes does, s…

GPT-4:

   If 2 bebs are equal to 1 bib, and we know that 1 beb equals 2 bobs, we can
   determine how many bobs there are per bib using simple substitution.
   
   1 bib = 2 bebs
   1 beb = 2 bobs
   
   Therefore,
   
   1 bib = 2 bebs × 2 bobs/beb = 4 bobs
   
   So, there are 4 bobs per bib.
Nitpick: A properly done substitution would've arrived at

   1 bib = 2 × (2 bobs)
without needing any of the "2 bebs × 2 bobs/beb" nonsense. It doesn't teach this task very well.

Re: OpenAI’s policies hinder reproducible research on language models

#172

Maybe OpenAI has the right to not reveal anything about their research and algorithms. But why don't we see similarly powerful truly open research backed by public, universities and companies? A truly open research will benefit lots of people and businesses.

Resources most likely. Training data, training a proxy that trains the real model, hardware, time, money. Managing such an open source project by itself would be terribly hard, considering the nature of model training, training data collection etc.

Re: OpenAI’s policies hinder reproducible research on language models

#173

I understand any individual's company anti-competitive measures. OpenAI looks at Google the same way Apple looked at IBM in the 80s. What I'm worried about is a lot of the talk about guarding models, public safety and misuse of models will end up leading every big company to pull public access of their APIs. We might look at 2022-2023 as a brief golden age when regular people could use stuff like GPT-4 before it was…

I only hope it doesn't go the way of that one paper which wanted to ban GPUs for sale to the public.

I'd say this is a real possibility though? Not necessarily for or against it, but you can't see this happening, or at least serious discussion of it?

Re: OpenAI’s policies hinder reproducible research on language models

#174
post #169

Earlier quoted context omitted.

> Since OpenAI didn't release the parameter count of GPT-4 That makes me ask what the open in OpenAI stands for?

Just like MTV doesn't mean Music TV anymore. As a joke I'd say, Open means "open your wallets"

Didn't know it was "Music TV", made me think about Skyrock... the biggest Rap channel in France, and essentially no Rock there.

Re: OpenAI’s policies hinder reproducible research on language models

#175

Earlier quoted context omitted.

This seems very difficult to solve incrementally. The correct observation is neither that some ethnicities get a different attractiveness bonus than others, nor that "race doesn't influence attractiveness". Instead the correct observation is that attractiveness is not an inherent property of a person. It exists only in the mind of the observer. I might find someone very attractive whom someone else does not find very…

> attractiveness is not an inherent property of a person This is like saying "value is not an inherent property of an object" - which is true in a philosophical sense, all value and beauty is a subjective, and depend on the opinions of people. But how would you then explain the existence of objects that have value to almost everyone in society (e.g. a car)? Similarly, how would you explain the existence of widely-rec…

The value of a thing to someone is also subjective. Ask two people (or even the same person twice in one day) how much they'd pay for a sandwich and you'll get different results. But "what's the value of a sandwich" has a very simple objective answer if you're at a sandwich shop. Maybe a slightly less objective answer if you're talking about the average price of a sandwich in all sandwich shops in the country, but it's still sensical to give a straight answer based on that metric.

No such objective answer can be found for attractiveness, though there isn't any fundamental reason why not; maybe if we had a culture of fetishizing appearance to the degree that we'd rank people and their attributes on the spot, we'd have more "objective" agreed upon measures available.

Re: OpenAI’s policies hinder reproducible research on language models

#178
If you came here after only reading the headline, you missed what the complaint is actually about:

It's not that GPT-4 is closed source. It's that access to `codex` model was pulled with only three days notice, and the model itself was not open-sourced. Since apparently a large number of researchers were writing papers which used that particular model, that means all of those research papers are now non-reproducible.

An obvious thing to do would be to either open-source older models (including the weights) when retiring them; or possibly transfer them to an institution who see their role specifically as serving as an archive / reference for this type of purpose. Open-sourcing older models shouldn't result in too much of a risk, either from an "AI Safety" perspective, or from a competitive perspective.

Re: OpenAI’s policies hinder reproducible research on language models

#179
post #166

Earlier quoted context omitted.

This misunderstanding may have something to do with how OpenAI was originally founded and the name: OpenAI.

Things change over time. Do you also complain that Apple doesn't actually sell any fruit?

Apple don't masquerade as a fruit seller. OpenAI started as a charity, took millions in donations, and have now abandoned their 'Open' principles.

Re: OpenAI’s policies hinder reproducible research on language models

#180
post #169

Earlier quoted context omitted.

> Since OpenAI didn't release the parameter count of GPT-4 That makes me ask what the open in OpenAI stands for?

Just like MTV doesn't mean Music TV anymore. As a joke I'd say, Open means "open your wallets"

Or TLC as the learning channel or History channel (assuming these still exist).

There are also lots of "Open Government" initiatives that end up being about making everything as opaque and confusing as possible. There were (are?) popular in the "big data" era, though funnily enough, if you watch "Yes Minister!" from ~40 years ago, there is a similar gag about "open government" in the first few episodes, so it's not new.

See of course Orwell, "we care about your privacy" banners, etc. People like to lie as blatantly as possible.

Post reply on HN