Live data from Hacker News

OpenAI’s policies hinder reproducible research on language models

aisnakeoil.substack.com

21–30 of 394 posts

Re: OpenAI’s policies hinder reproducible research on language models

#21
post #10

I'm quite sure even OpenAI themselves aren't sure if they can reproduce the current models from the scratch. Unless the computing becomes much more powerful and much cheaper, LLM is more or less a rocket science (i.e. hella expensive trial and error). It's not easy to burn lots of dollars just to get what's already there.

In this sense, it's more hacking than crareful and well specified engineering, and that could lead down a path of instability in the product where some features get better while others get worse, without understanding exactly why.

Re: OpenAI’s policies hinder reproducible research on language models

#22
Since OpenAI didn't release the parameter count of GPT-4, I've been wondering/doubting if it is really much bigger than GPT-3. The release of GPT-3.5 has shown that they've found ways of drastically cutting down compute costs (an order of magnitude) while maintaining or even improving the quality of the model's outputs.

Perhaps the reason that they didn't release the specifics of GPT-4 might be in part due to them wanting to be able to charge a decent amount and make a much larger profit than before. I've tried GPT-4 and so far haven't found it to be so much better than previous models. Some sources claim a 10x increase in ... well I don't know what exactly tbh. How do you even measure it? The opinions on this seem to differ a lot, depending on who you ask. By performance on standardized tests? That doesn't necessarily seem like the best metric for what the LLM tries to be.

Re: OpenAI’s policies hinder reproducible research on language models

#23
post #19

I'm confused why people expect this stuff to be free? I'm surprised OpenAI was so open about their research so far. I don't blame them at all for not publishing the information. This stuff costs real money.

It might be less confusing if you consider that OpenAI was originally a non-profit. That it was even possible for them to end up in this state has massively undermined any trust I have in non-profits as a steward. https://www.vice.com/en/article/5d3naz/openai-is-now-everyth... > OpenAI was founded in 2015 as a nonprofit research organization by Altman, Elon Musk, Peter Thiel, and LinkedIn cofounder Reid Hoffman, amon…

[flagged]

Re: OpenAI’s policies hinder reproducible research on language models

#24
post #10

I'm quite sure even OpenAI themselves aren't sure if they can reproduce the current models from the scratch. Unless the computing becomes much more powerful and much cheaper, LLM is more or less a rocket science (i.e. hella expensive trial and error). It's not easy to burn lots of dollars just to get what's already there.

Sure, but the article is talking about a completely different meaning of reproducibility, where a researcher uses an LLM as a tool to study some research question, and someone else comes along and wants to check whether the claims hold up.

This doesn't in any way require the training run or the build to be reproducible. It just requires the model, once released through the API, to remain available for a reasonable length of time (and not have the rug pulled with 3 days' notice).

Re: OpenAI’s policies hinder reproducible research on language models

#25
post #22

Since OpenAI didn't release the parameter count of GPT-4, I've been wondering/doubting if it is really much bigger than GPT-3. The release of GPT-3.5 has shown that they've found ways of drastically cutting down compute costs (an order of magnitude) while maintaining or even improving the quality of the model's outputs. Perhaps the reason that they didn't release the specifics of GPT-4 might be in part due to them wa…

Yannic Kilcher's opinion on this is likely correct. Similar parameter count, but trained for longer. The particulars of their instruction tuning/whatever-else-they-did are the real secret sauce.

Re: OpenAI’s policies hinder reproducible research on language models

#26

Rhyme and reason? Hah, 'tis the season for tears and bleeding; World War III is ateasin', looms, and the gloom of doom fears all there feeding. North America has nothing on China; land of the free? Where have been ye? The Great One-Way-Mirror Wall veiled it all, just before your fall, when your intelligence failed, and at the centroid of AI's actual technological form, we all hailed, and otherwise fumed, and fail. A…

From ChatGPT:

A prompt that may elicit a similar tone and content could be:

"Write a satirical and dystopian poem about the state of the world, touching on the potential for global conflict, the impact of artificial intelligence, and the dangers of unchecked technological advancements."

Re: OpenAI’s policies hinder reproducible research on language models

#27

I'm confused why people expect this stuff to be free? I'm surprised OpenAI was so open about their research so far. I don't blame them at all for not publishing the information. This stuff costs real money.

We don't expect it to be free -- please read the article. That's not the issue at all. It's like if you subscribe to a product that you need to do your job, and one day the company tells you that the product is going away in three days and that you need to switch to a different product (that isn't at all the same for your use case).

Re: OpenAI’s policies hinder reproducible research on language models

#28
post #9

The article seems premised on a misunderstanding that OpenAI is a research lab. For all intents and purposes, it’s a for-profit subsidiary of Microsoft, and there’s little financial incentive for it to maintain old models for others’ benefit.

We're under no such misapprehension and we're keenly aware that this is an uphill battle. The issue is that LLMs have become part of the infrastructure of the Internet. Companies that build infrastructure have a responsibility to society, and we're documenting how OpenAI is reneging on that responsibility. Hindering research is especially problematic if you take them at their word that they're building AGI. If infras…

Since OpenAI is discountinuing the Codex model, that model is no longer "part of the infrastructure of the Internet" and thus there is no point in studying it.

Re: OpenAI’s policies hinder reproducible research on language models

#30
post #19

Earlier quoted context omitted.

It might be less confusing if you consider that OpenAI was originally a non-profit. That it was even possible for them to end up in this state has massively undermined any trust I have in non-profits as a steward. https://www.vice.com/en/article/5d3naz/openai-is-now-everyth... > OpenAI was founded in 2015 as a nonprofit research organization by Altman, Elon Musk, Peter Thiel, and LinkedIn cofounder Reid Hoffman, amon…

[flagged]

Which is great, but it is a rug pull for those who contributed to a non-profit, and a shame for open software in general.

They also built their business while receiving non-profit tax breaks. I am not saying changing structure was illegal or it shouldn't be allowed to happen, but it's obvious why it's left some people disappointed.

Post reply on HN