Live data from Hacker News

Research acceleration: The view inside OpenAI

openai.com

171–180 of 210 posts

Re: Research acceleration: The view inside OpenAI

#173
post #158
post #125

Earlier quoted context omitted.

> Personally I’d like to see them actually start benefiting humanity by doing all the things Sam has claimed they will like curing disease, cancer, global warming, etc. It makes more sense to leave curing disease & cancer to the experts, with tools (like AI) being developed by AI experts. Call me crazy, but I want separate organizations and experts for medical vs finance vs space vs climate vs AI research.

What the op was pointing out is that guys like Altman and Dario are repeatedly saying they’re going to cure xyz diseases and solve xyz huge global problems. Maybe their companies will eventually do these things, but haven’t yet. I don’t have an opinion either way, I think it’s too soon to tell if llms will be able to cure cancer or whatever. But at the very least it will be a good tool to help researchers do their jo…

> Maybe their companies will eventually do these things, but haven’t yet.

I think they are working with customers to improve the LLMs and tools for these use-cases. They almost certainly also hire experts to help filter out nonsense, pseudo-science and help curate trusted knowledge bases for training, but it will almost certainly be the customers who deliver the major results, and the AI companies will claim some of the credit. That said, patents for important medicine might help with the bottom line, so I could imagine partnerships and JVs.

> at the very least it will be a good tool to help researchers do their jobs.

Indeed.

Re: Research acceleration: The view inside OpenAI

#174
post #71

Earlier quoted context omitted.

Can you elaborate on this? Especially the tooling. I tried something similar and I remember it was still pretty dodgy in February.

my stack in a sentence: refine the docs/prompts/skills often, that's your biggest job, use both frontier labs models reviewing each other, don't solve individual problems only the systemic ones (set standards strategically, don't define tactics) If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build co…

Thanks!

> If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review.

This part jumped out at me. There's something to watch out for here.

I recently had a funny experience. I delegated a major feature to an agent.

It turned out that it had implemented it precisely backwards, in a way which was pointless and which made things worse.

But it had written countless tests for the feature and all the tests were green.

I realised in that moment that even formal verification would not have helped, because it would simply have written a mathematical proof of the correctness of the incorrect feature...

Re: Research acceleration: The view inside OpenAI

#175
post #56

Earlier quoted context omitted.

Yeah, I kept looking for the first place it was defined in the article and... nothing

Same. Defining acronyms should become a habit when writing.

Re-become a habit. It has long been standard good writing to always define an acronym on first use.

Re: Research acceleration: The view inside OpenAI

#176
post #37

The burning question I can't get any information nn is whether, if they determined an earlier misaligned generation may have transmitted misalignment to the current models, they would roll back to a safe checkpoint to rebuild from there. I suspect they would not unless forced to.

They would just publish new articles explaining how they are taking the issue seriously. Maybe take the model offline for a few days. They are irresponsible and unserious. Their own Astra system card says: > GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more ca…

That is shocking. I can't even bring myself to give them credit for making this heavy misalignment and dangerous lack of monitorability public. It is surprisingly honest of them to say it though.

Re: Research acceleration: The view inside OpenAI

#177

Earlier quoted context omitted.

TSMC doesn't know how to build EUV machines - they are stuck buying them from ASML like everyone else. Putting a model's weights in read-only memory close to the processor is certainly a way to increase token/sec generation speed, but of course does nothing to increase intelligence. Robots aren't going to help though - semiconductor manufacturing is semiconductor manufacturing regardless of whether you are etching GP…

Ah, sorry, it's ASML expertise that needs to be cloned and scaled up. I don't see how it changes things, though. Robots don't need that much intelligence. High-speed joint control, "hand-eye coordination", the higher level tasks can be delegated to external models. Distillation already works quite well for isolating the required functionality.

I'm pretty sure ASML, and their supply chain, do have plans to increase production, as do the chip fabs - they all see the demand, while at the same time being leary of boom and bust which is the historical reality of the chip business.

But, the production expansion rate of none of these companies is being limited by lack of trained personnel, and if it were it would surely be faster to hire/train more humans since robots are still very far from human dexterity, not to mention intelligence.

Robots and AI are tools of automation, a way to replace humans with machines, but not all the problems in the world are bottle-necked by lack of humans, or the cost of humans.

Re: Research acceleration: The view inside OpenAI

#178
post #174

Earlier quoted context omitted.

my stack in a sentence: refine the docs/prompts/skills often, that's your biggest job, use both frontier labs models reviewing each other, don't solve individual problems only the systemic ones (set standards strategically, don't define tactics) If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build co…

Thanks! > If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review. This part jumped out at me. There's something to watch out for here. I recently had a funny experience. I delegated a major feature to an agent. It turned out that it had implemented it…

Yes, you need some kind of other source of truth. I think the best way to get that is to do clean room development with a different agent, but ultimately if you give them the wrong idea they'll do the wrong thing.

The other thing I do, not as much as I should, but it's very powerful, is to generate spikes and deliberately throw them away to understand how to prompt better. Like I generated a swift version of the react native app I'm working on, and Alloy provers for the state transitions. None of it is production quality but getting great results that way is useful to scope future work.

Re: Research acceleration: The view inside OpenAI

#179

Earlier quoted context omitted.

Ah, sorry, it's ASML expertise that needs to be cloned and scaled up. I don't see how it changes things, though. Robots don't need that much intelligence. High-speed joint control, "hand-eye coordination", the higher level tasks can be delegated to external models. Distillation already works quite well for isolating the required functionality.

I'm pretty sure ASML, and their supply chain, do have plans to increase production, as do the chip fabs - they all see the demand, while at the same time being leary of boom and bust which is the historical reality of the chip business. But, the production expansion rate of none of these companies is being limited by lack of trained personnel, and if it were it would surely be faster to hire/train more humans since r…

Investing in training a person gets you one trained person. Investing in training an ML system gets you a cloneable ML system that can be scaled on demand much faster. ROI might change quickly.

Re: Research acceleration: The view inside OpenAI

#180
post #131

Earlier quoted context omitted.

Am I the only one who's a bit disappointed that we're spending trillions, destroying the ecosystem, drowning democracies and learning in slop, preparing a big financial crash, all of this to achieve an "average AI research intern"? A long time ago, I used to be a (AI-adjacent) research intern, and frankly, I wouldn't trust any non-trivial task to that younger me. Fortunately, by opposition to an already trained LLM o…

"Destroying the ecosystem" is just FUD. And if you don't find "average AI research intern" impressive, I'm not sure what to tell you. Have the goalposts moved so far that open ended problem solving at "average CS student fresh out of the uni" levels is suddenly trivial? Think of what AI was capable of in 2016. Or even 2022. Compare that to now. We had more AI progress in the last five years than I expected to happen…

The ecosystem absolutely is being destroyed. We're looking at anywhere from 3-5C warming by 2060, which is going to be devastating if not flat out apocalyptic. And wherever we're at in 2060 it's not like it's going to stop there, nor is it going to be comfortable until then. Things may start to crumble much sooner.

I don't believe AI and data centers have played that much of a role in this though, we could have powered those without burning billions of tons of coal and gas etc, and im sure there's already a significant fraction of green energy powering then depending on location. Anyway we would have been roughly in the same spot right now with or without AI and some new data centers. The media just loves spinning the narrative to make the hordes of sheep scream about anything other than the real issues.

Post reply on HN