I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?
Depends on the domain imo. I work on a design tool - I don't think their political narrative will affect my work.
Grok 4.5
311–320 of 1001 posts
Re: Grok 4.5
#312Earlier quoted context omitted.
You could be typing the same about Google or a number of the other labs right now. A diverse market full of choices keeps it from becoming the browser wars all over again.
Google is playing a different game. I don't really know what game they're playing, but they're not trying to beat Claude Code. They have coding capabilities and Antigravity, but I'd be surprised if it's much more than an afterthought. They're focusing on efficiency, models at the edge, human interaction, image and video, etc. in ways Anthropic, in particular, is not. Google wants its AI to be pervasive in everyone's…
Re: Grok 4.5
#313(from Cursor's blog) > Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments. This is what the big money was for. Cursor is the first big player that had rea…
I've read multiple times that this approach is harmful in training.
You're essentially describing what many call distillation, but it's only useful in post training to guide behavior, it teaches how to behave, not how to think.
I might be wrong though and would be glad if someone more knowledgeable provided more insights.
Re: Grok 4.5
#314Can someone breakdown to me how this makes any sort of economical sense? Spending billions and billions to have the 3rd best model while even the number 1 and 2 players already seem to struggle making a profit. What am I missing here? Not trying to go full Ed Zitron but this doesn’t make sense to me.
Re: Grok 4.5
#315Earlier quoted context omitted.
Grok Build sucks compare to composer 2.5. Just use compose 2.5 and you'll have basically unlimited usage on the 40$ plan.
Composer 2.5 is so underrated IMO. I built a really feature rich application, insanely complicated, close to 200k LOC since it came out and for the most part it ran like a champ. Only used CLaude a couple times to get it unstuck. 8 hours a day and I'm paying about 30 a month.
Re: Grok 4.5
#316I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?
Edit: adding some other studies that are easily retrievable with a quick search for those unsatisfied with the first one - https://arxiv.org/abs/2606.12922 https://arxiv.org/abs/2412.16746
Claiming reality has a left wing bias is certainly an opinion you're welcome to have to explain this, but the reality of the bias in models is well evidenced. It seems that practically Grok's right wing tweaks mostly just combat the already pre baked bias existing models have (generally).
Re: Grok 4.5
#317I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?
How's this going with the rest of the models?
Re: Grok 4.5
#318So basically since US stopped OpenAI and Anthropic for 4 weeks, it allowed all other AI Labs to almost catch up. GLM 5.2 caught up, Cognition RL'ed Kimi 2.7, Grok 4.5 is out, DeepSeek v4 GA is out in a few days... What is the moat? and why should we pay for the expensive tokens today instead of just waiting a few months/weeks and getting AI for significantly cheaper? I must say, I feel like companies spending Million…
Re: Grok 4.5
#319Of the 3 models I tried, Grok did the best at making an iOS app I wanted for personal use (a bike computer with specific qualities). (Claude just gave up and did an HTML/CSS implementation but I insisted on native SwiftUI+Metal.) Grok definitely fumbles sometimes, but I have been surprised what it CAN intuit versus me having to micromanage it. (I am not an iOS developer, so getting something specific that I needed in…
As someone also not happy with my bike computer (some truly horrific UI/UX decisions), could you share or explain what you made? I like your web server.
Sure. I'm not sure if I will actually publish this thing, but I can show you: https://x.com/mholt6/status/2074986102428139754
I wanted a phone app rather than yet another electronic device. Phones do not have great screens in bright sunlight, and they run hot, so it's not ideal for a bike computer in the first place. But I can't deny the convenience of the multipurpose tool that is my phone.
This app will have a few UI/UX modes. The default is the futuristic-looking HUD, but it has a low-power mode that's mostly monochrome on black, and an even lower-power "Cruise mode" that removes the map entirely and just shows you speed, approximate heading, and nav directions. Still very WIP and mostly for my own amusement!
Re: Grok 4.5
#320I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?
All models are nudged. With out the actual source used to build a model, we don't know what's in them and it would be foolish to assume that people don't have their thumb on the scale when it's know, publicly, that they shouldn't be trusted.