Live data from Hacker News

Grok 4.6

x.ai

551–560 of 696 posts

Re: Grok 4.6

#551
post #477

Earlier quoted context omitted.

That has been the case for a while now: https://en.wikipedia.org/wiki/Ashcroft_v._Free_Speech_Coalit...

I think the comment you replied to was referring to the fact that when Twitter was taken over the entire Trust and Safety team was done away with. This has allowed child sexual abuse material to flourish on the platform.

It was referring to the feature they added where you could give a picture of a child to an AI module and ask it to undress it and it would comply, and millions of people did just that.

Re: Grok 4.6

#552
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> * If it becomes explicitly clear during the conversation that the user is requesting sexual content of a minor, decline to engage.

Why would they write "explicitly clear"?

'Explicitly is an adverb meaning to do or say something in a clear, exact, and direct way'

Surely they want to stop all requests for that content, even requests in an unclear, inexact or in-direct way. I only ask as I expect a lot of effort went in to defining that the wording of that prompt and it immediately stood out to me.

Re: Grok 4.6

#553
I’ve found Grok a bit like the open models from Chinese labs, it seems good in benchmarks and falls apart in real world use. How is this new version?

Re: Grok 4.6

#554

Earlier quoted context omitted.

Hm? I'm saying that the AI firms used to have the philosophy of "ok this kitchen knife is dangerous but we'll catch the murderers" on older AI models. But now, the AI firms think that any average person could send a flying knife to attack a political figure they don't like, from the comfort of their home. Now give this to a billion people, and suddenly you have chaos. So to continue the analogy, now they're mandating…

The analogy tracks because the stupidity of doing those other things is directly analogous. It's like pointing out that slamming your fingers in the door and slamming your toes in the door both hurt. That's why you shouldn't be purposely doing either one. How is the new stuff any different than the longstanding fact that anyone can go anywhere and then commit an act of violence? The thing that prevents this isn't tha…

The difference is the asymmetry of the potential warfare we're talking about here.

Committing physical, in-person crimes anonymously has obviously always been possible: there are unsolved murders, thefts, and other crimes every day. But they require a great deal of personal risk to the criminal because the criminal has to physically put themselves into the act of committing the crime, along the path of getting to where the crime is, and has to face an opponent, if their crime is against another person.

Now, that can be sourced remotely, routed through anonymizing tools, VPNs, etc., and do a great deal to cover their tracks so that the "pretty good chance they go to jail" can be substantively minimized in a way we couldn't previously contemplate.

The idea that we should let the US based models be permissive because at least they'll be subject to subpoena power is fatuous: yes, strictly speaking, a user committing crimes on a permissive foreign model will be harder to catch, but non-sophisticated users who have never heard of hugging face may find that being blocked by the US model is enough for them to reconsider their behavior. A dedicated enough individual is going to commit the crime they're going to commit, but there are tons of situations where preventing trivial access to tools that can be used for malice can actually prevent malice from occurring.

Re: Grok 4.6

#555

Earlier quoted context omitted.

The alternative is Claude-style "safeguards" aka censorship, which: 1. doesn't eliminate the possibility of a jailbreak anyway 2. frequently has false positives, triggering on innocuous requests, which is just really annoying Not saying that we can't (or shouldn't) do better than Grok, but I really don't know what the best solution is here...

> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.

at some point the kitchen knife analogy stops being useful

Re: Grok 4.6

#556
post #326
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

I think they dumb down their public models to be only slightly better than the competition. And the real competition is China, so the current state of the Chinese models would define the baseline. I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible for Chinese models to close the gap that easily. It's also very unlikely that they s…

> With that difference in training speed, it should be impossible for Chinese models to close the gap that easily.

They have a big advantage in that they can directly distill from frontier models.

Re: Grok 4.6

#557
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

My theory is that it all boils down to better data and longer post-training period. Cursor got curated data from the trillions reactions of real world developers in real jobs. xAI bought is and used it for its post-training and got Grok 4.5 . Longer post-training on the powerful Colossus cluster helped it get Grok 4.6 , although both versions use the same model with the same number of parameters. Thus, both must use the same pre-trained model as a baseline. See also an article infers the training and release timeline of popular models featured a few days ago here on HN.

Chinese labs must follow similar trajectories plus their specific efficiency improvements. That also explains the jump from DeepSeek 4 performance in April and July releases. They both use the same pre-trained model as well.

Re: Grok 4.6

#558
post #47
post #43

Tangental, but has anyone else noticed grok's voice mode got stupid and terse ~2 weeks ago? I've absolutely loved grok's voice mode since it came out (incredibly useful for brainstorming on walks and helping conceptualise and get the verbiage for expressing ideas) but it seems so have lost about 40 IQ points recently, and if the question is multi-part, it often answers just one part with no elaboration or explanation…

Presumably because Musk has been training it to be more like him.

Worth every one of the downvotes from the mecha-hitler youth.

Re: Grok 4.6

#559
post #313

Earlier quoted context omitted.

> CSAM is by definition limited to real imageries and cannot be generated. Where in the definition does it imply this?

I thought that's just the legal definition in any sufficiently developed countries?

People have been prosecuted and convicted here in Sweden for Japanese hand drawn CSAM.

I think it comes down to different ideas of why the law exists. If you believe removing access to pornographic material for this category means people will have a harder time becoming pedophiles, then that's how the Swedish law makes sense. If you believe pedophilia is a tragic disease that we can't treat and that synthetic pornography can help these people lead somewhat dignified lives without hurting children, then the Swedish law is actively damaging. Ultimately I don't think we have a strong scientific basis for any of those two view points currently. I'm leaning towards the second, but weakly.

Re: Grok 4.6

#560
post #514
post #408

Earlier quoted context omitted.

Grok 4.6 vs Sol 5.6 vs Opus 5: https://aibenchy.com/compare/openai-gpt-5-6-sol-low/x-ai-gro...

So Grok 4.6 is incredibly expensive in that comparison? It's double the cost of Sol 5.6 Low and still nearly double of the cost of 5.6 Medium (which scores +6% over Grok).

Yeah, Sol models have higher in/out cost but are incredibly token efficient.

Also, in those tests Sol Low did better, but you can also compare the price vs Sol High, then it's getting a bit closer.

So Grok 4.6 is still not the best choice when paying API rates, but they are improving fast.

Also, the more important difference is that sol is a lot faster.

Post reply on HN