Earlier quoted context omitted.
Sonnet 5 is a huge token hog, though, it uses far more reasoning tokens than Opus models while being priced at $2/$10 with promo, and $3/$15 (usual Sonnet price) afterwards.
I'll probably get hate for it, but I was not impressed by Fable, I felt like it was just Opus with more tokens for thinking. I feel like the second I turned on Fable I drained my usage more quickly, despite them billing it as though it were Opus level of usage. The value is just not there for me. I wish they could make Haiku remain low-cost and drastically more capable to the point you could use only Haiku.
Grok 4.5
81–90 of 1001 posts
Re: Grok 4.5
#82Isn't this the same Twitter company that was supposed to go bankrupt a few years ago? Now it is somehow part of a Space company that has an AI division inside of it? I think we are going to be waiting a long time for Twitter / X to go bankrupt as it was (erroneously) predicted a long time ago.
In the transaction announcement (xAI buying twitter) twitter reported $12b in debt on acquisition, roughly the amount originally sourced ($13b), so it apparently made good on its debt covenants during the operating period. I have no idea if it received additional capitalization from Musk to do that or not.
That said, the deal was classic Musk - anybody who went on the equity ride with him in Twitter just KILLLED it; xAI was valued at $80bn and twitter at $33bn, so the owners there became 30% owners of xAI. xAI was acquired for $250bn at a SpaceX valuation of $1 trillion, or 20% of the resulting entity, so the twitter stock was 6% of spaceX at about $2 trillion, or $120bn on an equity purchase price basis of $30bn. and that $120bn in value is on really good daily trading volumes; lots of depth.
Re: Grok 4.5
#83terminal is nice but codex desktop app is very useful
Re: Grok 4.5
#84How popular is Grok compared to other companies models for SWE tasks? I almost never hear it talked about against OpenAI's or Anthropic's products
Because of the of the political stuff, they have a bad reputation I think and are taken less seriously (I feel this way). They have an opportunity imo to break free from that and just not do the gatekeeping / condescension that the other providers are starting, and become more mainstream.
Using Grok is therefore a supply chain risk and it's not nearly good enough to offset that risk.
Re: Grok 4.5
#85So basically since US stopped OpenAI and Anthropic for 4 weeks, it allowed all other AI Labs to almost catch up. GLM 5.2 caught up, Cognition RL'ed Kimi 2.7, Grok 4.5 is out, DeepSeek v4 GA is out in a few days... What is the moat? and why should we pay for the expensive tokens today instead of just waiting a few months/weeks and getting AI for significantly cheaper? I must say, I feel like companies spending Million…
Re: Grok 4.5
#86Of the 3 models I tried, Grok did the best at making an iOS app I wanted for personal use (a bike computer with specific qualities). (Claude just gave up and did an HTML/CSS implementation but I insisted on native SwiftUI+Metal.) Grok definitely fumbles sometimes, but I have been surprised what it CAN intuit versus me having to micromanage it. (I am not an iOS developer, so getting something specific that I needed in…
Re: Grok 4.5
#87Earlier quoted context omitted.
Low effort and uninformed comment. The team published a good followup on why this happened: the model pulled in people's own tweets as context to prompting so edge lords that wrote innocuous prompts got to see edge lord content.
Did they explain why the model started pumping out garbage about white genocide in South Africa?
Re: Grok 4.5
#88> Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments.
This is what the big money was for. Cursor is the first big player that had real-world data from real-world projects, before cc / codex were a thing.
> We used reinforcement learning on difficult problems in realistic environments spanning both software engineering and broader knowledge work. These environments teach the model to investigate problems, use tools, recover from mistakes, and verify results.
> Many of these problems had to be designed to be difficult enough that even frontier models fail at them. As models improve, existing tasks stop teaching them anything new, and problems that once required extensive reasoning become routine.
> We developed a distributed agent system to construct these environments at scale. Engineers specify a problem and how a solution is verified, and large groups of agents construct, test, and refine each environment.
This is where scale comes in. You use the previous gen model to prepare datasets for the next model iteration. The better the models, the better the data, the better the next models. (they also have a comparison with their composer2.5 training run, for people still thinking chinese models are "close to SotA"...)
Reports of xAIs demise (after giving a lot of compute to Anthropic) were slightly exaggerated, it seems.
> Grok 4.5 was trained across tens of thousands of NVIDIA GB300 GPUs
Re: Grok 4.5
#89Earlier quoted context omitted.
Competition. You don't want to lose your customers trying out the competitors updated and better product. Release on the same day and they won't be able to compare their new to your old.
But how do they know what day is that? Unless you have already something ready to be announced (and you just hold it until the very last moment, which doesn’t make sense, since you could just announce it asap)
Maybe a little corporate espionage.
Probably more keeping an eye on the behavior of the competition and predicting what they might do and adjusting your own schedules.
Re: Grok 4.5
#90Its remarkable how Anthropic is able to maintain their edge against all competition. Anyone have any idea what the secret sauce is that has Anthropic at the top of all leaderboards for the past few years?
I have never liked the various nerfs Anthropic has used to balance GPU (slowing down responses, quota variance, model optimizations etc) and it definitely has burned a lot of good-will.
But it has seemed that being able to look beyond the short term pitchforks has worked quite well.