I don't know where I need to sign up to try it out. What is pricing? Is it API or subscription, what?
I had the exact same experience with Grok 4.5 as well.
81–90 of 228 posts
I don't know where I need to sign up to try it out. What is pricing? Is it API or subscription, what?
I had the exact same experience with Grok 4.5 as well.
Earlier quoted context omitted.
> but they can't read minds I don't know who needs to hear this, but neither can humans. You've implicitly assumed here that AI systems will always be worse at contextualizing and framing questions than the average engineer. I'm not sure that it's even true now let alone in the nebulous future. You haven't narrowed the fundamental myopia of the assumption here, just dressed it in slightly different clothing.
> You've implicitly assumed here that AI systems will always be worse at contextualizing and framing questions than the average engineer. How would they know what to ask or contextualize if they don't know what the user wants?
Plus all the other things that software engineers generally have not learned to a professional level even if they picked up the basics on the job by osmosis, because figuring out the customer's needs (and what they'll pay you for which may be different) is the job of a business analyst, a PM, or a UX researcher, and those are different skills and two of them may come with a Business Informatics degree rather than a CompSci one.
LLMs can be "eh, better than nothing" at many things, not just code.
Earlier quoted context omitted.
At least in China a lot of software developers are now struggling. I think for a lot of type of software we have now reached peak employment. Someone payed a few k just for a normal website.
> At least in China a lot of software developers are now struggling. Do you think that Chinese software industry is that relevant to the kind of software market talked about on HN? I.e. lots of enterprise b2b and infra companies. Chinese companies have always had a very low willingness to pay for software which kinda breaks the flywheel of B2B SaaS companies and companies to service those companies all the way down.
Are we still left with this mindset? Maybe once upon a time but it has definitely been changing.
There's plenty of B2B and enterprise SaaS companies in China serving the Chinese market. Maybe not as many, but no longer the very low of the past.
I also would not say enterprise were not willing to pay, even many years ago. It's the SME that refused to pay. Large CRM, ERPs etc have always existed.
How is every company able to show itself at the top of every benchmark?
To be fair, seems more correct to compare against similar strength models if your main edge is pricing.
Interesting how the prevalent opinion until yesterday seems to have been that OpenAI & Anthropic are irreversibly ahead and now with xAI and Meta at least delivered something that's competitive with useful models and cheap too. Granted, the narrative that the two leading labs are ahead still holds with Fable (and perhaps an upcoming GPT6), but it's not as over as common knowledge by the opinion leaders would have us…
Not the way you're implying?
The GLM 5.2 hype was blowing way before this. Neither xAI nor Meta have really made a difference in a different way - similar results / similar pricing (to GLM 5.2).
The pricing is insane: $1.25/$4.5 for 1M tokens, and $0.15 for cached input! https://dev.meta.ai/docs/getting-started/pricing-rate-limits
I really dont see how anyone's willing spend more than $1.50 per mm output. Let alone $15-50. Does anyone actually pay for usage based billing as a consumer?
Earlier quoted context omitted.
To expand on Chinese models: - DeepSeek - GLM (Z.ai) - Minimax - Kimi (Moonshot) - Hy3 (Tencent) - Qwen (Alibaba) (Each one of these with weights available to download and run locally)
GLM 5.2 is great, but is so rate limited now I no longer recommend it
Interesting that neither meta nor xai chose to do open source given that they are both clearly behind Google, OpenAI and anthropic - and a serious us open source offering would give them a clear foothold.
Their published benchmarks seem to indicate that it's pretty good at coding and multimodal, but VERY good at successful tool calls. What kind of use case would be best for that shape?
Their published benchmarks seem to indicate that it's pretty good at coding and multimodal, but VERY good at successful tool calls. What kind of use case would be best for that shape?
I wonder if we'll start to see that pattern with every new release. Tool use likely changes rapidly, so the newest, rather than most intelligent, model may always have an edge.