Live data from Hacker News

Grok 4.3

docs.x.ai

331–340 of 608 posts

Re: Grok 4.3

#331

Grok 4.3 was completed ahead of its CEO’s lesson on this common safety resource: Asked if he knew anything about OpenAI's "safety card," Musk smiled and replied: "Safety card? Why would it be a card?" https://www.axios.com/2026/04/30/musk-openai-safety-grok Low relevancy in spite of cluster size and musical chair gas generators for time being: Later in his testimony, Musk was asked about a claim he made last summer t…

Elon has publicly stated that he cares a great deal about safety. He has stated that the only safe models are those which align greatest with truth, that which is in reality. In this, xAI has lived up, as it has proved to hallucinate least (or close to least) in benchmarks. If you read that, quote again, he is saying "how can you quantify safety in a card?"

Elon publicly states a lot of things, most of which aren't truthful.

Re: Grok 4.3

#333
post #57

So, we have: - claude for corps and gov - codex for devs - grok for what, roleplay, racism? Those are the two things I've ever heard grok associated with around me.

So interestingly, I know of at least one application in a charity that deals with trafficking where grok was happy to do one-shot classification tasks where all other models refused to cooperate. I think there's a surprising number of actually useful applications in this sort of grey area for a slightly-less guardrailed, near-frontier model (also the grok-fast models are cheap!).

I am software dev and i was doing a security check on my own application (work) I was running in localhost and gave it access to the code.

every single model refused to attempt to run any sort of test to check if it was a n issue other than grok.

Re: Grok 4.3

#334

Grok is my favorite model for chatting, and my favorite voice mode. It seems to be the only voice mode that isn't routing to a extremely cheap model (like Haiku), and has been the highest quality out of all the frontier ones. When you subscribe to SuperGrok you can also create a "council" of agents, each with their own system prompt and when you ask something, they will all get asked in parallel to come to a conclusi…

The Gemini app voice mode uses one of their more recent models (and not some gimped small one), and is very capable. The personality is also fine, much more natural than the Gemini web chat, with my only complaint being it's insistence on suggesting a "next step" which seems to he something that they all do. I'm not sure if the "next step" is just to drive cost up for you (but makes no sense for free version), or bec…

I think the "next step" instruction is more about engagement than cost, basically giving the user some options to continue the chat. I always have had success by ending the prompt with "only reply with nothing else but the answer to the query in a precise way". This usually always works better than telling it to not ask leading questions etc but a straight up expectation of the answer format you need is an instruction that most models can follow imo

Re: Grok 4.3

#335
post #28

Earlier quoted context omitted.

When I signed up, I accidently paid for a full year. So from time to time, I'll throw it something just to see what it produces compared to the other LLMs. And, even after all this time, it still feels like a really "dumb" model compared to the other frontier ones. But, worse, many of my system prompts make it go wacky and puke jibberish. However it was pretty cool for those couple months awhile back when it was unce…

Ah yes the psychosis reinforcement vertical. It's such a lucrative market for those schizophrenics and bipolars. Great way to get lots of engagement. Groks portfolio is so diverse

It's a great way to get funded by your CEO and get good performance reviews; xAI employees know how their bread is buttered.

Re: Grok 4.3

#336
post #291
post #119

Earlier quoted context omitted.

Credit where it's due, Grok is currently the only model that has near-realtime updates from/access to a waterhose of data, and is casually used by regular people all the time. I don't think there's a single thread on Xitter whete people don't delegate some question to grok. (There's a separate conversation of failure modes, and whether it's a good thing, and how much control Elon had when he doesn't like Grok's "woke…

All the major tools can websearch guy

It's not just about web search though -- there's another element too. I go to Grok to find things I have failed to find with web search.

I agree with GP -- if I want sourced commentary on current events, Grok is my go-to above the other models. For whatever reason, its search feels better and more up-to-date -- whereas the others feel more like filters of media, Grok feels more like filters of sources.

Could just be my perception though. YMMV

Re: Grok 4.3

#338

Earlier quoted context omitted.

Sadly, it's more likely that people will just start talking like bots

I've seen this expressed as a concern even from one of my colleagues. My retort was: "English is not my native language and LLMs taught me quite a few very useful formalisms that do land well for people and they change their attitude towards you to be more respectful afterwards. It also showed me how to frame and reframe certain arguments. I agree sounding like an LLM is kind of sad but I am getting a lot of educatio…

It's impressive that you've even managed to use an em-dash in spoken language. /s

Re: Grok 4.3

#339

So, we have: - claude for corps and gov - codex for devs - grok for what, roleplay, racism? Those are the two things I've ever heard grok associated with around me.

No point in even trying to have close to a sensible discussion on this topic here. Musk-related posts seem to consistently get brigaded by his acolytes or bots. That and many HN users seem completely comfortable separating morality for what little progress "only Musk" can offer humanity, a la Wernher von Braun.

[flagged]

Re: Grok 4.3

#340
post #312

Earlier quoted context omitted.

Lol I wonder when Anthropic discussed the idea of Claude Code internally, were there bozos saying "3rd parties will eventually deliver this so we shouldn't waste time one it."

The only good thing Claude Code did was bring coding harnesses to a wider audience. It is not a good harness.

What are good harnesses? I haven't yet been able to get good agent teaming approaches out of other harnesses yet, before that feature I mostly regarded the space as competitive, but until another harness can do as well with Claude models it seems like it's better for now?
Post reply on HN