Live data from Hacker News

Grok 4.6

x.ai

671–680 of 696 posts

Re: Grok 4.6

#671
post #314
post #94

Earlier quoted context omitted.

(I work on Grok) This isn't allowed. CSAM / deepfakes are against our acceptable use policy.

Elon Musk literally went to court to protect the ability to make child porn with Grok: https://www.ag.state.mn.us/Office/Communications/2026/07/31_...

If it were literal then you'd be able to quote the exact part of the text that says that, but you can't because that is literally not in the page. Even Ellison, who surely would love to make that claim, is intelligent enough not to make that claim directly.

A more balanced view[1] of the proceedings is available via mainstream news:

> In the 38-page lawsuit, xAI — whose AI model chatbot and image generator Grok is available on the social media platform known as X, formerly Twitter, and elsewhere — said it does not contest the state’s interest in banning the distribution of AI-generated nude images of real people without their consent. But it said Minnesota’s law “extends far beyond that goal,” banning many constitutionally protected images and video and subjecting the company to a penalty of $500,000 per violation.

> The lawsuit argues there is no “safe harbor” provision for companies that make good-faith efforts to prevent such images from being created by users, and that it covers images that were consented to by the depicted person, or even created by that person.

> It also says the law’s definition of “intimate part” is overly broad, covering body parts that are routinely displayed in public.

Calm down.

[1] https://apnews.com/article/minnesota-artificial-intelligence...

Re: Grok 4.6

#672
post #519

Earlier quoted context omitted.

Others have said this too but LLMs are the best approximation of magic we have. We etch runes on stones, put electricity through them and then try to “convince” them to do our bidding. The answers vary wildly sometimes depending on minutiae. Prompts should be really called spells. It really feels more like “should I add the frog’s eye or leg into the cauldron” than engineering.

>LLMs are the best approximation of magic we have. I don't think the alchemists suddenly became scientists, or died off to make way. It was a gradual transition. They didn't quite work out how to transmute lead to gold, but the alchemists and their descendants did eventually discover - and create - substances that are worth more than gold by weight. Now we have created sand that can teach itself how to talk. We covet…

   They fear the sand in the East.
I gotta use this in a book, this sounds great.

Re: Grok 4.6

#673

Earlier quoted context omitted.

There is a herd of companies all running a race. The technology is known. They all have roughly the same resources. It’s not unexpected that they have similar cycle times for model development and that those models will be of roughly the same quality. Then layer in corporate PR demands and you see all these models landing within weeks, sometimes days, of each other to keep the model developer’s name associated with “…

It still feels a bit strange to me that no one comes up with "secret sauce" that gives a big edge for at least a few months.

I agree. Thus far, we seem to just be training an absurd number of parameters, and that’s all.

Re: Grok 4.6

#674

Earlier quoted context omitted.

Humans, however, are highly variable, which may produce really varied and interesting results if they work together. One instance of an LLM is the same as another instance, so while you may get more out of it by stacking more of them, I strongly suspect it falls victim to diminishing returns. 100 instances of the same LLM may converge on the same result as 10.

You can change the temperature if you like. Have a 1000 agents being 'normal' and 10 being chaotic.

High temperature makes the LLM pick more out-of-distribution tokens, but the choices its presented with are still the same or same-ish. I'm not convinced that the more random outputs don't end up averaging to roughly the same conclusion after enough passes.

Re: Grok 4.6

#675

Earlier quoted context omitted.

Humans, however, are highly variable, which may produce really varied and interesting results if they work together. One instance of an LLM is the same as another instance, so while you may get more out of it by stacking more of them, I strongly suspect it falls victim to diminishing returns. 100 instances of the same LLM may converge on the same result as 10.

> 100 instances of the same LLM may converge on the same result as 10. Not in the highly verifiable domains. There you can take it from say 80-90% maj@x to 99% pass@n. Math, some parts of programming and cybersec are examples of highly verifiable domains. (e.g. if you're searching for a linux LPE, that's expensive to search but easy/cheap to verify - just have a token in /root and have the model retrieve that token)

Verifiability makes it easier to understand how well the LLM works, but this doesn't counter my hypothesis. If X number of instances get 99.0% on an objective, verifiable metric, is there any guarantee that 10X will get 99.9%? The fact that we are reliant on new model releases to push capability in big ways, and that people running gigantic clusters of LLMs end up beaten by new models implies that the capabilities of a given model have a hard upper limit, and that it may not even take much to reach it.

Re: Grok 4.6

#676

Earlier quoted context omitted.

Humans, however, are highly variable, which may produce really varied and interesting results if they work together. One instance of an LLM is the same as another instance, so while you may get more out of it by stacking more of them, I strongly suspect it falls victim to diminishing returns. 100 instances of the same LLM may converge on the same result as 10.

I think using different AGENTS.md can give the same model different perspectives on the same problem. For example a model with a well-tuned AGENTS.md by an expert mathematician approaching the same problem as the same model with a well-tuned AGENTS.md by an expert biologist can grind on the same problem from different perpectives. It's worth a shot at least, as a microservices architect I have a bias that we aren't n…

Crucially, does it make capabilities infinitely scalable? My comment just said that models may have a hard cap, and maybe doing specific setups like yours can make reaching it easier, but making the 'team' 10x larger after that optimal point may bring few to no improvements.

Although I'm also not sure about just how much better models can really get with this technique. Ultimately you're still getting the same model with the same training data, which are the important parts. Asking it to pretend to be something feels like it would just put a color filter in front of the conclusion the model has already predicted, or maybe alter the path to the conclusion slightly or pick a less likely answer that it still could've provided normally.

Re: Grok 4.6

#677
post #638

Earlier quoted context omitted.

Laughable thinking they aren't

we have free speech, don't be so gullible. you think Chinese citizens can go out in the streets and say anything they want. pathetically ignorant.

[deleted]

Re: Grok 4.6

#678
post #638

Earlier quoted context omitted.

Laughable thinking they aren't

we have free speech, don't be so gullible. you think Chinese citizens can go out in the streets and say anything they want. pathetically ignorant.

We're not China no, but where's Snowden right now?

Free speech doesn't make the government not authoritarian. It's not binary either.

Re: Grok 4.6

#679

Earlier quoted context omitted.

So basically, nothing that actually affects working with it in August 2026. Got it. Facebook has a far longer (and worse) laundry list of offenses and I'm sure you still use it. Or Threads, or Instagram. > My organization has outright banned Grok That's too bad, as it's currently the only model that won't consistently flag honest good-actor security questions, in my experience. So I'd ask you who you work for, but I…

Your experience is not reflective of mine at all, or my colleagues’, so I would check out the better SOTA models out again. I use codex extensively for security-related work - much of which is _overtly_ offensive - without issue. Same for Claude, minus Fable, after going through their approval process. I also went through OpenAI’s, but theirs was just basic KYC and instant. GPT-5.6 in Codex has produced full chain RC…

> after going through their approval process

Are you really trying to compare my experience as a NON-certified security professional, with yours?

Mostly understood re: the rest. Musk knows what damage he caused, and he knows it likely because it hit him in the dollars.

Re: Grok 4.6

#680

Earlier quoted context omitted.

*comprehend *people Also, that's not what strawmanning is. I never denied that Grok didn't act bizarrely offensively over a fucking year and a half ago (so did other LLMs, btw... and so have many other experiments over the years, remember Microsoft's?), which is an eternity in this space. I know Musk is polarizing, but give me a fucking break. Don't assume malice when social incompetence serves as an exculpatory fact…

So i do care that Elon Musk is responsible for USAID shutdown. The richest man on the world shuts down human support so abruptly that he causes real humans to die. Elon Musk, as the richest person on the planet, bought himself a propaganda platform he controls and started to finger around in democracy. Its a lot more than 'just' CSAM.

It disheartens me that if I want to find the most objective lowdown on the USAID and governmental meddling thing, I can get it trivially from an AI, and people will immediately dismiss it as slop, even if it provides receipts for all of its assertions:

https://chatgpt.com/share/6a7e8991-399c-83ea-bc65-2cf9a536fa...

Why don't more people do this, or trust it? Do they not want to get upset when their sacred cows are toppled? Do they not realize that getting their cows toppled actually makes them more correct? I don't know, and I'm starting to not care.

My statement to you is: I don't care what you believe or care about. I only care if you have actually faced the evidence against your beliefs and care. If they still stand after that from a good-faith point of view, then you can have them. Otherwise, they are just "partisan", and I have zero interest in partisan politics anymore, because it's way too bullshit-infused at this point, driven by a media whose profit only comes from outrage, and mired in people who are obsessed with identity politics.

(A friend once told me, "you are the most persuaded by evidence and argument of anyone I know... and that's a sad thing for humanity," so there's that.)

That said... My personal take (since I find this particular point unaddressed by the larger conversation around this) is that a lot of intangible value was lost by USAID going away, such as American goodwill. This then turns into fodder for terrorists... Which is, of course, very nearsighted. And the bill for that may come due one day.

Post reply on HN