Live data from Hacker News

Grok 4.6

x.ai

371–380 of 696 posts

Re: Grok 4.6

#371
post #172

Earlier quoted context omitted.

These system prompts are not the only safety layer that these models use. There's other more deterministic filters in place both on input and (streaming) output.

I think it is fair to argue that prompts are not a safety layer at all and can't be relied upon for much. "Make no mistakes"

it is a heuristic though, and can be measured as such.

my steel yield strength table is similarly not guaranteed to be correct for the piece of steel that I have in front of me.

Re: Grok 4.6

#372
post #352

Earlier quoted context omitted.

So, what do you do for a living? You've made a personal attack and seem to be under the impression you're morally superior. So, I'm curious as to what highly virtuous role you take on in your daily life. That said, I see your comment history is a lot of one sentence personal attacks against people. Not a lot of thoughtful debate. This makes hypocrisy out of your supposed concern for social good.

I work for myself and I don't enable the production of CSAM so yeah I'm quite content being morally superior on this issue

OP doesn't either. You haven't answered the question.

Re: Grok 4.6

#373
post #172

Earlier quoted context omitted.

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

These system prompts are not the only safety layer that these models use. There's other more deterministic filters in place both on input and (streaming) output.

Is there concrete evidece that those are xAI's default prompts anyway? They seem plausible enough but how would company outsiders know?

Re: Grok 4.6

#374

As polarizing as grok is, it was basically inevitable for it to start being a real competitor given how much investment SpaceX made into its own inference capabilities. Seems if you are okay with it, there's no reason to use anything but the highest effort levels of some other frontier models for the price. I think Grok provides healthy competition to the other labs, though I do think they bank on groks reputation ma…

Curious - what is the main issue you find polarizing with grok?

It took many many many turns for me to have the model even acknowledge that the fake elector scheme was actually a thing. It's very much primed to answer vaguely when it goes against the current political ideals of its owner.

Re: Grok 4.6

#375

I am noting Opus 5 is omitted. Interesting as I thought it benched better than Fable 5 in a few benchmarks.

The Opus 5 release was a perfect example of how useless these benchmarks are for a head to head model comparison. Anthropic published a post showing Opus 5 beating Fable in almost every eval but then added a disclaimer that it was still a tier below Fable in intelligence (and thus pricing). So then what did all the numbers represent exactly?

Re: Grok 4.6

#376

Earlier quoted context omitted.

"I beg to differ, it is my opinion that reality should be different to what you have observed"

in reality even the mention of a prohibition is enough to make the model reject that no matter what

That’s just not true. There are bypasses that happen all the time.

Re: Grok 4.6

#377
post #294

Earlier quoted context omitted.

I'm a bit confused by the Cursor relationship here, the acquisition hasn't closed yet, what are they doing with Composer?

The $60 billion cursor option that SpaceX bought was exercised on June 16th. The deal is closed.

No it was just announced then but it's still going through regulatory/antitrust

Re: Grok 4.6

#378
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

NLP guys were right all along :)

Re: Grok 4.6

#379
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.

Re: Grok 4.6

#380
post #379

Earlier quoted context omitted.

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.

What do you think a human brain is…
Post reply on HN