Live data from Hacker News

OpenAI’s policies hinder reproducible research on language models

aisnakeoil.substack.com

271–280 of 394 posts

Re: OpenAI’s policies hinder reproducible research on language models

#271
post #197

Earlier quoted context omitted.

And those suggestions would be very in-line with the original purpose of OpenAI. A purpose they are now actively hindering in the name of profit.

I think what most of the people here are missing is how big, how paranoid, and how influential the "AI alignment" movement is. To you it looks like they're being overly careful and paranoid, perhaps as an excuse to set up a monopoly silo to extract money. But a lot of the people the OpenAI researchers work closely with -- people deep in the "AI alignment" community -- are telling them that they're being wantonly reck…

I'm still trying to figure out if I'm alone here but I feel like it's much harder to find a developer job currently (well, unless you work on AI... Perhaps it's time to bank my Stanford ML class certificate?) because GPT4 could potentially make everyone's existing employees twice as productive at the same cost, and (especially considering the extreme Fed rate hike in 1 year) who's going to take the risk of hiring someone new in this economic climate? The sheer number of new variables being thrown into the mix out there right now is complete chaos to any sort of prediction model

Re: OpenAI’s policies hinder reproducible research on language models

#272

Earlier quoted context omitted.

This seems very difficult to solve incrementally. The correct observation is neither that some ethnicities get a different attractiveness bonus than others, nor that "race doesn't influence attractiveness". Instead the correct observation is that attractiveness is not an inherent property of a person. It exists only in the mind of the observer. I might find someone very attractive whom someone else does not find very…

> attractiveness is not an inherent property of a person This is like saying "value is not an inherent property of an object" - which is true in a philosophical sense, all value and beauty is a subjective, and depend on the opinions of people. But how would you then explain the existence of objects that have value to almost everyone in society (e.g. a car)? Similarly, how would you explain the existence of widely-rec…

Surely attractiveness is a function of both the person being evaluated and the person doing the evaluation?

That is, a person's visual appearance has N aspects, and each person evaluates those N aspects differently. Attractiveness is then a kind of dot product between the two.

Seen this way, a person which is universally attractive is one with aspects u that is the solution to Au = 1, where A is a matrix of valuation vectors (one row per person), and 1 is a vector of ones.

Obviously very simplified but...

Re: OpenAI’s policies hinder reproducible research on language models

#274

Earlier quoted context omitted.

Super-human cognition? Hard to say. GPT-4 does raise the possibility of a machine writing smarter text than a human. What perplexes me is that since GPT is a predictor, it shouldn’t be able to write the smartest text - it should write the average text (since that has the largest frequency in the training set). Yet this does not seem to be the case. Is it inevitable that despite the quality of the data, better models…

> could the GPT-4 secret sauce be RLHF weighing intelligent answers higher? That part is one of the rare things that the technical report addresses. In Appendix B[0], they show that RLHF does not improve capabilities on human tasks. It does improve alignment. To me, this is an indication that they performed better scaling analysis and pretrained until it no longer improved. As the Chinchilla paper showed, GPT-3 was u…

> (As a side-note, it feels weird to me that they used a free academic archive to store their technical report, even though it cannot go through peer review or be accepted in any academic publication.)

I find this to be particularly egregious. Such an obvious false front is a red flag to me.

Re: OpenAI’s policies hinder reproducible research on language models

#275
post #242

Earlier quoted context omitted.

>> An obvious thing to do would be to either open-source older models (including the weights) when retiring them; or possibly transfer them to an institution who see their role specifically as serving as an archive Another obvious thing to do is do your research on non-commercial or open source things that can not be taken away from you. Sorry, I don't mean for the snark present in that statement. The frustration lie…

I thought I must be going crazy until I saw your comment. This sounds like a bad research practice that probably shouldn't be reproduced to begin with.

Research into systemically important infrastructure cannot be damned because that infrastructure isn't public. It's a cheap moralizing argument to say "pfff, this was predictable". Maybe so, but there isn't an alternative. Much like research on Twitter. Once these companies start to drift into providing what become broadscale social utilities and public services it doesn't matter that they're private. There are(/should be) obligations that come with that.

You can't handwave and say go do your research on some micro-niche open source project that's way behind the SOTA and has nowhere near the same reach. That's not what "best practice" means here.

Re: OpenAI’s policies hinder reproducible research on language models

#276

Earlier quoted context omitted.

GPT 3.5 had trouble understanding when I told it "Say 2 bob are a beb, how many beb per bob are there?" and it wrote a goddamn essay about shoes. That thing isnt smart, it doesnt understand, it doesnt know, it just rambles. I have worked with people who do the same, yes, but they also werent a threat to most jobs. I said it before, and I will say it again: If ChatGPT 3,4,5,... can take your job, maybe youre not reall…

Answer from GPT-4: "This question seems to be intentionally nonsensical or is using unfamiliar terminology. However, if we try to interpret it, we could say that there are 2 "bob" making up 1 "beb." In this case, there would be 0.5 "beb" per "bob." Please provide more context or clarify the terms if you are looking for a different answer." Answer from GPT-3.5 (subscription version, not free): "If 2 bob are a beb, the…

What do LLaMA-based models answer for this?

Re: OpenAI’s policies hinder reproducible research on language models

#277
post #197

Earlier quoted context omitted.

I think what most of the people here are missing is how big, how paranoid, and how influential the "AI alignment" movement is. To you it looks like they're being overly careful and paranoid, perhaps as an excuse to set up a monopoly silo to extract money. But a lot of the people the OpenAI researchers work closely with -- people deep in the "AI alignment" community -- are telling them that they're being wantonly reck…

Until the alignment movement begins to take seriously the idea that we already have misaligned artificial general intelligences I think they are best viewed as a convenient foil Paperclip maximizers exist, they're made not only of code but of people

Precisely! I'm much less concerned about super-intelligent AIs and much more concerned with shortsighted, greedy humans using pretty-good AIs (like those we have now) to squeeze out every ounce of profit from our already misaligned systems, at the expense of everyone else. Not to mention the political implications of being able to convincingly fake voices, photos, and videos.

In this sense, I'm pleased to see Open AI claim to be taking a more careful stance, but to be honest I think the genie is already out of the bottle.

Re: OpenAI’s policies hinder reproducible research on language models

#278
post #242

Earlier quoted context omitted.

I thought I must be going crazy until I saw your comment. This sounds like a bad research practice that probably shouldn't be reproduced to begin with.

Research into systemically important infrastructure cannot be damned because that infrastructure isn't public. It's a cheap moralizing argument to say "pfff, this was predictable". Maybe so, but there isn't an alternative. Much like research on Twitter. Once these companies start to drift into providing what become broadscale social utilities and public services it doesn't matter that they're private. There are(/shou…

Sure but OpenAI isn’t preventing research. It’s not their responsibility to provide reproducibility, at their expense, for any researchers looking at GPT, that job is the responsibility of the researchers, and the researchers still can work. It might be unfortunate from their perspective that there used to be a nice tool that makes their job easier, but the flip side here is that OpenAI didn’t say why they’re removing access to codex, and they probably have good reasons, not least of which is it costs them money that researchers aren’t subsidizing.

Re: OpenAI’s policies hinder reproducible research on language models

#279

IMO established companies (Meta, Google, etc) had their researchers publish papers as a competitive benefit or way to attract talent from academia (a researcher wouldn't want to stop publishing). Companies didn't see an issue with doing that because those papers were not "giving away" the core of the company, for example, Facebook's DeepFace paper from 2014 couldn't hurt its ad business. OpenAI on the other hand will…

Yep, and that's the difference between a big profitable company doing research as a side-hustle, and a company whose business IS the research.

One interesting and somewhat scary exception seems to be Microsoft; they seem to be converting a lot of their recent research projects into commercial value.

Re: OpenAI’s policies hinder reproducible research on language models

#280
post #197

Earlier quoted context omitted.

And those suggestions would be very in-line with the original purpose of OpenAI. A purpose they are now actively hindering in the name of profit.

I think what most of the people here are missing is how big, how paranoid, and how influential the "AI alignment" movement is. To you it looks like they're being overly careful and paranoid, perhaps as an excuse to set up a monopoly silo to extract money. But a lot of the people the OpenAI researchers work closely with -- people deep in the "AI alignment" community -- are telling them that they're being wantonly reck…

A powerful "Bootleggers and Baptists" pattern seems to have emerged in tech space.

In online media, and social media the power of major platforms became apparent at some point. Happenings in twitter or FB can determine politics, catalyze rebellions (eg Arab Spring), uprisings, even genocide.

At this point the pressure and desire to act responsibly becomes irresistible.

This "camp" finds common cause with "bootleggers" who want to lock down the platforms and markets for commercial reasons.

Post reply on HN