Live data from Hacker News

GPT-5.6 Sol Ultra will be in Codex

twitter.com

221–230 of 433 posts

Re: GPT-5.6 Sol Ultra will be in Codex

#221

Earlier quoted context omitted.

https://archive.ph/NEwVz "However, these inference optimizations, which rival Anthropic refers to as “compute multipliers,” are a big focus for all the labs. Anthropic CEO Dario Amodei has been publicly talking about the concept since at least mid-2023, when he said on a podcast that the company limits “the number of people who are aware of a given compute multiplier” because it could give other AI labs a leg up if t…

Ok I’m not sure I follow your point here. Isn’t all that he’s saying that if they find some optimization techniques, that gives them an edge? And that makes sense? How is this suddenly evidence of him being a villain?

It has got to be one of the most insane takes I've read on HN, which to be fair has been trending towards "unhinged" when it comes to Anthropic and AI safety.

Compute multipliers are like a quant firm's trading algorithms. They're the crown jewels, the whole alpha of the lab. If you leak them, the lab dies.

Protecting them does not make Dario a villain, it's literally his job. It's also Sam's job, Denis's job, Mira's job, etc. Every lab guards these multipliers closely because they represent the entire worth of the lab.

Re: GPT-5.6 Sol Ultra will be in Codex

#222
post #73

Earlier quoted context omitted.

Humans use em dash as well. I hate that I have had to remove it from my writing style because people assume it’s AI generated. But I think that ship has sailed. I’ll have to do without now.

How do you type the em dash. I thought the point about the em dash "—" is, that it is longer than the normal minus "-". Humans normally have no way to produce it, cause there is no key on the keyboard.

OMG — I'm a robot.

Re: GPT-5.6 Sol Ultra will be in Codex

#223
post #214
post #193

Earlier quoted context omitted.

No need for the downvotes imho. Can’t speak for u/noduerme but in putting outsourcing (labor, LLM/AI) in the same basket I don’t see a category mistake, but a dry way of looking at business. What are the risks, what are the rewards, what future skills are we at risk of losing (the business logic part) if we go in direction XYZ. Something completely different (but with the same logic): do you outsource legal, hire you…

since when did HN have downvotes?

Its one of the features that is behind a minimum karma requirement.

https://github.com/minimaxir/hacker-news-undocumented/blob/m...

Re: GPT-5.6 Sol Ultra will be in Codex

#225

Earlier quoted context omitted.

Other than the delay time, I'm not sure I see the difference. You're removing your primary from the job of writing code, putting them into an editorial role, which removes responsibility and agency and actual hands-on understanding. The quality of the code is beside the point. More friction (language barriers, time zone difference) is actually better if you want to maintain institutional knowledge, because it require…

I reject the premise that using LLMs absolutely leads to loss of institutional knowledge. It is trivial for an LLM to generate a knowledge base of any kind in any language which can answer any question about your institution at any time. How is a bunch of fragmented humans with limited knowledge who can’t all communicate with each other better than that?

Do you seriously think that "institutional knowledge" can be maintained by simply writing down a few pages on Confluence or whatever system your company uses for internal documentation?

Re: GPT-5.6 Sol Ultra will be in Codex

#226

Hope this forces Anthropic to be less stingy with Fable.

It is not because they want to but because they literally don't have the capacity to.

What a suprise, to make better models, they just making them larger and larger, then they can barely run it, after a round of LARP-ing that they invented some dangerous LLM?

Re: GPT-5.6 Sol Ultra will be in Codex

#227

Earlier quoted context omitted.

I find this kind of cynicism fascinating tbh. On the one hand, it seems so relatable in some ways, because there is something uncomfortable about being seen as naive, in a way that being seen as cynical or negative doesn't seem to carry. I guess it's just self-protective, almost like some kind of perverse Pascal's wager: it's better to think everyone is horrible and be wrong than to think the opposite and be taken ad…

It's the Tragedy of the Commons. There will always be a vacuum of predatory bullshit that can be filled, and the victor is always the biggest sociopath. Rockefeller, Cecil Rhodes, Elon Musk, it's the same traceable pattern all the way back through history. It's not that everyone is like this, but that a few crafty marketeers are able to ruin it for everyone. Why should I treat Sam and Dario with special white gloves?…

The only reason Chinese labs release their research is because they are behind. They lose nothing by doing so because the alpha has already been discovered.

This is not hard to understand. Do you really think DeepSeek would publish their algorithms if they led the American companies? Lmao.

Re: GPT-5.6 Sol Ultra will be in Codex

#228

Earlier quoted context omitted.

The craziest for me is companies that sticking stochastic agents into automated business processes and expecting stable/reliable outcomes. Businesses want deterministic processes in the vast majority of cases.

I'm struggling with the assertion that these models cannot provide reasonably deterministic guarantees. I am using gpt to populate JSON objects conforming to a list of natural language constraints for purposes of generating fake customers. I am finding that gpt5+ never fucks up. Not even a little bit. I've ran this test hundreds of times with 20+ constraints and it's been perfect every time. Stable information yields…

We could probably debate this ad nauseam, so I'll just give you my most compelling arguments.

1) Writing code the "old fashioned" way (i.e., a Python program that does X, Y, Z) allows you to arrive at a battle tested solution that will not change over time. From a risk assessment perspective, the behavior is essentially immutable, allowing a business to guarantee consistent behavior over long periods of time.

2) Just because something hasn't happened to you do, does not mean that it will not happen. LLM are opaque. If you stay on the "happy path", you may see consistent behavior for long periods of time, but there's always potential for an edge case where something goes catastrophically wrong. This is without even opening the can of worms regarding prompt injection and intentional sabotage of a working system.

3) There are plenty of real world examples of an LLM spontaneously deleting data from a DB (or the entire DB) or otherwise going completely off the rails. These might seem hyperbolic, but it happened at our company (to a test DB, not production). The severity of errors that occur can be existential to a business' survival without the proper guard rails.

4) There's no concrete way to truly confirm understanding between an LLM and a human. It can tell you that it completely understands what you want, and then it can do exactly the opposite. Followed by, "my bad" (Claude's new favorite catch phrase). Code can be audited and even proven to be correct given the appropriate level of time and energy.

My best results have been gleaned in using LLM to produce deterministic systems. I recognize everyone has different use cases and needs, but this seems to be the best use of the technology in my experience.

Re: GPT-5.6 Sol Ultra will be in Codex

#229

there seems to be very big misunderstanding about what the "ultra" is, so let me explain it basing on the codex source code: it's similar to Claude code ultracode. there is no ultra effort level implemented on the backend. it's just alias in the codex to max effort setting and single line addition to prompt to use subagents proactively. that's all as far as we know pro models work differently. for once those are back…

ultracode in Claude Code kicks off a dynamic workflow.

Re: GPT-5.6 Sol Ultra will be in Codex

#230

Earlier quoted context omitted.

implode, how?

They're in a price war with the People's Republic of China running flat out with the full backing of a government that literally does not care if they ever see a financial return on the investment, they just want to drive the value of LLM training and inference to zero because we banked the market on it being arbitrarily high margin forever. China was like hold my beer. They have a staggering surplus of grid capacity…

> They're in a price war with the People's Republic of China running flat out with the full backing of a government that literally does not care

There is nothing to support this. You get cheap Deepseek tokens by foreign providers too.

It is the same thing with automakers. They complain about not being able to only make luxuary cars with high profits with BYD raining on their parade and blaming the Chinese gov.

Post reply on HN