GPT‑5.3‑Codex‑Spark
openai.com
GPT‑5.3‑Codex‑Spark
1–10 of 415 posts
Re: GPT‑5.3‑Codex‑Spark
#2(Yes I know they released /fast last week but I’m loving the constant oneupsmanship)
Re: GPT‑5.3‑Codex‑Spark
#3Re: GPT‑5.3‑Codex‑Spark
#4When they partnered with Cerebras, I kind of had a gut feeling that they wouldn't be able to use their technology for larger models because Cerebras doesn't have a track record of serving models larger than GLM.
It pains me that five days before my Codex subscription ends, I have to switch to Anthropic because despite getting less quota compared to Codex, at least I'll be able to use my quota _and_ stay in the flow.
But even Codex's slowness aside, it's just not as good of an "agentic" model as Opus: here's what drove me crazy: https://x.com/OrganicGPT/status/2021462447341830582?s=20. The Codex model (gpt-5.3-xhigh) has no idea about how to call agents smh
Re: GPT‑5.3‑Codex‑Spark
#5Re: GPT‑5.3‑Codex‑Spark
#6Re: GPT‑5.3‑Codex‑Spark
#7> more than 1000 tokens per second
Perhaps, no more?
(Not to mention, if you're waiting for one LLM, sometimes it makes sense to multi-table. I think Boris from Anthropic says he runs 5 CC instances in his terminal and another 5-10 in his browser on CC web.)
Re: GPT‑5.3‑Codex‑Spark
#8In my opinion, they solved the wrong problem. The main issue I have with Codex is that the best model is insanely slow, except at nights and weekends when Silicon Valley goes to bed. I don't want a faster, smaller model (already have that with GLM and MiniMax). I want a faster, better model (at least as fast as Opus). When they partnered with Cerebras, I kind of had a gut feeling that they wouldn't be able to use the…
> I don't want a faster, smaller model. I want a faster, better model
Will you pay 10x the price? They didn't solve the "wrong problem". They did what they could with the resources they have.
Re: GPT‑5.3‑Codex‑Spark
#9Off topic but how is it always this HN user sharing model releases within a couple of minutes of their announcement?
Re: GPT‑5.3‑Codex‑Spark
#10If 60% of the work is "edit this file with this content", or "refactor according to this abstraction" then low latency - high token inference seems like a needed improvement.
Recently someone made a Claude plugin to offload low-priority work to the Anthropic Batch API [1].
Also I expect both Nvidia and Google to deploy custom silicon for inference [2]
1: https://github.com/s2-streamstore/claude-batch-toolkit/blob/...
2: https://www.tomshardware.com/tech-industry/semiconductors/nv...