Live data from Hacker News

Grok 4.5

x.ai

511–520 of 1001 posts

Re: Grok 4.5

#511

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

This website tracks AI model political leanings:

https://trakkr.ai/bias

Grok differs from some of the other models (it's more libertarian, and more right wing), but all models have their biases - particularly ChatGPT, which sits to the economic left of 81% of US adults. See https://trakkr.ai/bias/findings

Re: Grok 4.5

#512
post #457

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

Has it occurred to you that _all_ model providers are actively trying to shape their models' replies to fit their preferred political narratives?

Yeah but their political narratives involving dehumanizing people, and hurting minorities - which is way worse. The attempt at false equivalence between traditional American propaganda and right wing American propaganda is disgusting.

Re: Grok 4.5

#513
post #92

First impressions: - Very fast, easily beats GPT 5.5/Opus 4.8/GLM 5.2 because of higher t/s (around 90?) and very high token efficiency - Very good price, no contest vs GPT and Opus which are very overpriced if you pay API costs, and probably cheaper than GLM 5.2 when you take into account the token efficiency. - Will take quite a while to get a feel for how smart it is, but it's definitely good, I'd say in the same…

Concur. Tried on a "this test suite is weaker than I'd like, too often depending on internal state rather than outcomes" problem via Cursor, asking it to "review and suggest solutions." It gave me a quality overview of the test approaches, strengths, weaknesses, and gaps then recommended a disciplined multi-prong approach based on a common, trusted testing library ( https://hypothesis.readthedocs.io/en/latest/ ). It…

My benchmark is ripping tailwind out of a few year old elixir Phoenix liveview app, and replacing it with component level scoped styles

It's a good and complex task, that requires touching the build system, most components, the stylesheets, and more. Opus 4.6 could barely do it. Sonnet 4 cannot (haven't tried 5 yet). MiniMax actually did fairly well

Grok aced it, rather quickly and cheaply, surprisingly

I run each through Oh my pi, with dexter providing the LSP for elixir

Re: Grok 4.5

#514
post #406

Every time I get excited about Grok’s performance on benchmarks and demo videos, I test it myself and end up disappointed. I'll give this one a try with a grain of salt and lowering my levels of expectations

I am trying to benchmark it now, but: - It doesn't seem available in EU (?) - Using a VPN seems to sort of fix it, but it's way slower than I expected, when everyone was praising it, it feels like the speed is slowly ramping up - Cost is $2/$6 for https://aibenchy.com/compare/x-ai-grok-4-5-medium/z-ai-glm-5...

Instead of a VPN, might I suggest running litellm proxy on a server in the US and connecting to that

Re: Grok 4.5

#515
> Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools.

This -- training on work done on hard, real-world tasks -- seems to be how most frontier models are making capability gains these days. In fact people make decent money doing that for data companies like Mercor. However it's also striking that Cursor managed to gather so much of such data.

Turns out Cursor will train on everything you do unless you opt-out, even if you're already paying for it with cash! Are that many people really not opting out?

This is why it seems like a significant concern to me: It's very clear that typical, run-of-the-mill coding has been completely commoditized, so the primary value remaining is either in novel use-cases and applications, or novel technical solutions to hard problems.

Presumably the value for novel use-cases could be captured by building a business around it via the usual moats (distribution, relationships, network effects, first mover advantage, etc.) so the code and techniques do not matter as much.

However novel technical solutions, which are already hard to monetize without building a whole damn business around it, could at least be capitalized on by simply being able to claim credit for it. I'd at least like the option of being "paid in exposure" if I'm not getting paid in cash. But having them "leaked" unwittingly via the training corpus to whosoever happens to prompt the model with the same problem removes even that option.

I know people have been calling out this risk forever, and I don't use any tool that I can't opt-out of training completely, but the scale at which this is happening -- on an ongoing basis, mind you, after training on the data of the whole world, and that too after paying for the product -- is surprising. I'm bullish on the technology but we really should be way more careful handing these AI companies even more of our intellectual crown jewels.

Re: Grok 4.5

#516

Earlier quoted context omitted.

"Almost" is doing a lot of work there; there is no alternative to Fable.

You can get fable-ish performance with gpt 5.5 watching over opus output. Although it fundementally cannot work as well because gpt 5.5 doesn't see the thinking process behind opus 4.8 unlike fable which presumeably self-steers and is natively trained for it. See more: https://omp.sh - turn on advisor and set advisor role to gpt 5.5 xhigh thinking.

Advisor is such a killer hidden feature

Re: Grok 4.5

#517
post #457

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

Has it occurred to you that _all_ model providers are actively trying to shape their models' replies to fit their preferred political narratives?

From a business perspective, a company that trains it's LLMs to having boring, mainstream, generally-inoffensive views is a big selling point over whatever the hell Elon is doing.

Re: Grok 4.5

#518
post #457

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

Has it occurred to you that _all_ model providers are actively trying to shape their models' replies to fit their preferred political narratives?

That's the claim, and it's a belief that's self-fulfilling prophecy, like saying all politicians are corrupt.

If you can convince everyone that everyone is corrupt, it hurts anyone who isn't corrupt. You hear people preferring those who have no shame about their corruption, based on the premise that those who aren't overtly corrupt must be more sinister and dangerous if they hide their corruption so well.

It's a race to the bottom.

Re: Grok 4.5

#519

Earlier quoted context omitted.

It's not the preferred political narrative of the model that I worry about. It's how brazen they are about altering their models to achieve it. It makes me wonder what else they're altering. I have trust issues with OpenAI and Anthropic as well, but with those companies, at least I know their motives are purely profit driven. I don't have that assurance with xAI.

Implying you prefer your manipulation to be subtle an not discussed

I don’t think that is necessarily a bad preference if this was an actual dichotomy. Not all types of manipulation is equal, and when you at least try to hide it shows at least some respect for the user.

That said, I don‘t believe this dichotomy is real. Personally I don‘t use AI, political manipulation is however only a relatively tiny part of my reasoning for opting out.

Re: Grok 4.5

#520
post #92

First impressions: - Very fast, easily beats GPT 5.5/Opus 4.8/GLM 5.2 because of higher t/s (around 90?) and very high token efficiency - Very good price, no contest vs GPT and Opus which are very overpriced if you pay API costs, and probably cheaper than GLM 5.2 when you take into account the token efficiency. - Will take quite a while to get a feel for how smart it is, but it's definitely good, I'd say in the same…

This is what I don't understand. Why would I use this "cheaper" model when it's still going to be more expensive than Codex on the $200 plan? Are they only targeting business users who pay per token?
Post reply on HN