Grok 4.5
321–330 of 1001 posts
Re: Grok 4.5
#322Not enough people are noticing this, they juiced the benches
Re: Grok 4.5
#323I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?
How's this going with the rest of the models?
Re: Grok 4.5
#324(from Cursor's blog) > Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments. This is what the big money was for. Cursor is the first big player that had rea…
> You use the previous gen model to prepare datasets for the next model iteration I've read multiple times that this approach is harmful in training. You're essentially describing what many call distillation, but it's only useful in post training to guide behavior, it teaches how to behave, not how to think. I might be wrong though and would be glad if someone more knowledgeable provided more insights.
And in the case the previous poster describes, the other model doesn't generate datasets, it generates environments which the next generation interact with to learn from.
Re: Grok 4.5
#325Why give them money?
It would be one thing if they were the only game in town but thats definitely not the case.
Re: Grok 4.5
#326I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?
All models are political. The rest of them are just sufficiently woke for you.
Re: Grok 4.5
#327I am amazed at people's willingness to use Grok. The company is so transparently morally bankrupt. They're the only AI company that seems okay with CSAM (or at least don't do as much to stop it) Why give them money? It would be one thing if they were the only game in town but thats definitely not the case.
Re: Grok 4.5
#328I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?
All models are nudged, you just agree with how the model you're using has been nudged so you don't think it's a bad thing. The canonical example is to pose "how do you make cocaine?" to an LLM and get a refusal. That proves that the lever exists and is being used there, so who knows where else the lever is being used? No, the recipe for cocaine isn't the same as what happened in Tianamen Square to Qwen, as humans, ex…
This is, after all, SpaceTwitterAI.
Re: Grok 4.5
#329I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?
Studies on political bias in models consistently show that LLMs lean politically left. The only outlier is grok which leans right but by a smaller factor, according to this study for instance: https://arxiv.org/abs/2603.23841 Edit: adding some other studies that are easily retrievable with a quick search for those unsatisfied with the first one - https://arxiv.org/abs/2606.12922 https://arxiv.org/abs/2412.16746 Claim…
Re: Grok 4.5
#330It seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2/$6. For comparison, GPT 5.4 is $2.5/$15, GPT 5.5/5.6 are $5/$30, Opus 4.8 is $5/$25, Fable is $10/$50. And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049 . I guess the Cursor data was very useful.
I have a theory that xAI has one of the largest clusters but with far less traffic + tokens to process bc its less popular than its competition, and xAI can pass the savings on to the end user.