Live data from Hacker News

Accelerating GPT-5.6 Sol Ultrafast

cerebras.ai

291–295 of 295 posts

Re: Accelerating GPT-5.6 Sol Ultrafast

#291

Now I can blast through my weekly 20x pro codex credit in like an hour, great! The amount of usage you receive on Codex these days is dismal compared to what it was a few months ago, FYI. And they charge more for going faster. As a Codex customer, I am not impressed with their shenanigans over the past few months and I have resolved to master the art of Pi Coding Harness creation and loving it. Thanks for all the fis…

Wait, you have noticed it too? I burned through weekly limit in 25 hours after last reset on Thursday, without changing what I was doing before that.

Interestingly enough, previous subscription to Plus gave me about 2-3 days of coding. So I switched to Pro Light now, gave me about the same amount for a week or so (G-d bless these quota resets of theirs!). Now, with the last reset I only gathered 25 hours before I hit my weekly limit. Now I am buying credits, I switched to Luna to execute very narrow sets of patches, and offload things to "free" Spark model, and it is blowing through tokens less actively, but noticeable quickly too. I am not sure if it is something with how tokens are counted, or how they are counted depending where you are in a subscription cycle. Could it be that one person's "limit" is not the same as another, trying to push you to buy token credits?

Re: Accelerating GPT-5.6 Sol Ultrafast

#292
post #252

Earlier quoted context omitted.

OpenAI and Anthropic are believed to be making billions per month in revenue selling tokens to enterprises. It’s not that surprising for enterprise SaaS they would have individual customers making up ~1% of their total sales. Although, perhaps a bit more surprising those few customers are posting about it here. Another interpretation would this is a counterfactual savings, like they previously paid $1M for y tokens,…

A second angle on the counterfactual savings would be Luna telling them not to pursue a potential session with the expected/extrapolated (from which sessions they did ignore Luna on, e.g. just to keep efficiency statistics current) sunk costs at time of getting shut down used to derive the quoted number.

A third angle is that it is not true and that's just some attempt at influencing an HN thread, which is happening quite regularly pro and against AI. We know that the pro-AI has a way bigger budget to burn with this kind of operations tho.

Re: Accelerating GPT-5.6 Sol Ultrafast

#293
post #178

Earlier quoted context omitted.

We’re reaching transvestigation levels of people trying to spot AI text everywhere they look

Hmm, this sounds like something a bot would say to prevent being caught.

15 years ago i was astounished by the intellectual deepness of this community. Now i'm aware this has always been a cult and their former cult leader Mr. Altman wants to destroy this capitalist society. He isn't even hiding motives. People just stopped listening carefully.

Re: Accelerating GPT-5.6 Sol Ultrafast

#294

Earlier quoted context omitted.

no, because closed sourced model pricing has no relationship to its size. Thats what im saying. the inference margins are crazy, but people think the fonrtiner models must be 10T params or something because theyre expensive

> closed sourced model pricing has no relationship to its size. That's too strong. Only in an actual monopoly for a product with no substitutes that has price inelastic demand can pricing fully disconnect from costs. Frontier model serving is only maybe a soft version of that, where costs and moat both contribute to pricing.

sure, but my point is that supply and demand are what determines pricing, not cost to serve. its like thinking that because something costs $1 to make, its not possible for a company to sell it for $10, or that becuase they are selling something for $10, it must cost them $9 to make.

Re: Accelerating GPT-5.6 Sol Ultrafast

#295

Earlier quoted context omitted.

I still personally think that a heavy lean into MoE will be better for that sort of thing. Our brains are subdivided into large parts but I'm sure (and I'm not a brain scientist) that those parts can be subdivided even further into systems that run at various frequencies and latencies depending on what they're used for. I was thinking about it the other day actually. How our brains evolved structure. I imagine it was…

like MoE with a billion "experts". That seems promising to me too, although, the thing I've always read is that you can't make the "experts" too narrow. Even if you had a "coding expert" it has to know a lot more than coding - if you tell it to make an online store it needs to parse your language, understand the internet , what a "store" is in this context, etc. I am not a primary source, probably not even a secondar…

Oh for sure, but I think generally multiple experts are selected in an MoE pass for a token, so presumably it'd select programming related ones as well as general knowledge/language.

Only problem would be the routing layer works on the previous token as far as I understand so it might need more informational depth than just "a token" to select experts, I suppose in the same way attention works.

I wish I had the GPUs to run those sorts of experiments ha ha.

Post reply on HN