Earlier quoted context omitted.
They already are. I have been using Kimi k2. It is 90% as good as Sonnet and on Groq 3x faster and 1/5th the price.
What kind of GPU setup are you using for Kimi?
OpenAI prepares to launch GPT-5 in August
51–60 of 75 posts
Re: OpenAI prepares to launch GPT-5 in August
#52There’s so much work to be done developing coding related tools that integrate AI and traditional coding analysis and debugging tools. Also programming needs to be redesigned from the ground up as LLM first.
Yes, because non-deterministic systems are great softwares. I mean, who want's repeatable execution on the control program for their nuclear submarine or their hospital lighting controls. Why would anyone want a computer capable of actual math running on the President's nuclear "football" when we can have the outputs of hallucinating tools running there.
Re: OpenAI prepares to launch GPT-5 in August
#53Earlier quoted context omitted.
Not sure. Since Google started to include a Gemini response on top of their search results I stopped using chatgpt for search
Funny, the AI summary makes the experience shittier for me. The UI jumps and everything moves, I now have to wait until it loads. Massive UX mistake, you learn this the first week you make websites ...
Re: OpenAI prepares to launch GPT-5 in August
#54Earlier quoted context omitted.
I am still skeptical about the value of LLM as coding helper in 2025. I have not dedicated myself to an "AI first" workflow so maybe I am just doing it wrong. The most positive metaphor I have heard about why LLM coding assistance is so great is that it's like having a hard-working junior dev that does whatever you want and doesn't waste time reading HN. You still have to check the work, there will be some bad decisi…
> And the code you get out, in my experience at least, is pretty crap I think that belies the fundamental misunderstanding of how AI is changing the goalposts in coding Software engineering has operated under a fundamental assumption that code quality is important . But why do we value the "quality" of code? * It's easier for other developers (including your future self) to understand, and easier to document. * Easie…
Often AI critics say things like “quality is bad or “it made coding errors” or “it failed to understand a large code base”.
AI proponents and expert users understand these constraints and know that these are but actually that important.
Re: OpenAI prepares to launch GPT-5 in August
#55> Altman decided to let GPT-5 take a stab at a question he didn’t understand. “I put it in the model, this is GPT-5, and it answered it perfectly,” Altman said. If he didn't understand the question how could he know the model answered it perfectly ?
It takes a really special kind of self-delusion to recognize that you don't understand the question and also think you are qualified to evaluate the answer.
Re: OpenAI prepares to launch GPT-5 in August
#56They can call it whatever they want…not sure that has a great deal of meaning unless there’s a GPT-4/Claude 3.5 level step change.
Turning to their reasoning models, it’s also known and documented through SimpleQA and PersonQA that OpenAI o3 hallucinates more than o1, and o4-mini even more than o3. There’s an unmanaged issue where training on synthetic data improves benchmark results on STEM tasks but increases hallucination rates, especially troubling OpenAI models for some reason (my guess: they’re fine-tuned to take risks since it’s known to also increase likelihood of getting it right for hard tasks?)
Google has long known OpenAI struggles with hallucinations more than them according to an anonymous Googler that I saw commented on this. This has been verified by the aforementioned benchmarks. Anthropic also struggles less. But as far as I can tell, they’re all facing issues with synthetic data acting like a double edged sword.
So GPT-5 is going to be interesting. How well it exactly does will bear a lot of meaning for the kind of trouble OpenAI is in right now. Maybe OpenAI has found a novel approach in reducing hallucinations? I think that’s among their most crucial points right now. But other than this, no, I don’t expect a revolution, only an evolution. They might currently win benchmarks, but it will hardly be something that catapults them.
If GPT-5 underwhelms, it will bear a stronger signal than merely the one that GPT-5 underwhelms. Because then OpenAI has trouble with both non-reasoning and reasoning models, and we’re likely to be looking at the end of the road on the horizon for current GPT based LLM’s and one where the winner will probably ultimately be cheaper open weight models once they catch up.
Re: OpenAI prepares to launch GPT-5 in August
#57What's the point of this article besides free propaganda? It seems to me like every other AI shop except for OpenAI and possibly Anthropic only gets mentioned once they actually release something.
I was surprised at the number and calibre of orgs that came to me who would basically say anything I wanted for cash, this opened my eyes a lot and made me very suspicious of published media.
Re: OpenAI prepares to launch GPT-5 in August
#58Re: OpenAI prepares to launch GPT-5 in August
#59If it were any good I would assume there would be no need to hype it up. My theory is that LLMs will get commoditized within the next year. The edge that OpenAI had over the competition is arguably lost. If the trend continues we will be looking at inference like commodity prices, where the most efficient like cerebras and groq will be the only ones actually making money at the end.
Anyway, I'm sure gpt-5 will be AGI.
Re: OpenAI prepares to launch GPT-5 in August
#60Evolution or revolution? They’d better deliver, as Gemini has been hogging all the attention and open source models are fast catching up.
I use Gemini at work, and ChatGPT for everything else 'personal'. Also from my non-tech friends, I usually hear them talk about ChatGPT and not Gemini or other models. I think, even if Gemini outperforms ChatGPT, there's definitely a strong 'first mover advantage' at play. I suspect that being outperformed by Gemini etc won't diminish their market share significantly.