Live data from Hacker News

Grok 4.6

x.ai

271–280 of 696 posts

Re: Grok 4.6

#271

Earlier quoted context omitted.

wonder if the world's richest man with no moral compass may pay botnet ranchers to astroturf on his behalf? we may never know!

Or, you know, people are sick of comment sections getting turned into political slap fights and react poorly to such.

[dead]

Re: Grok 4.6

#272
post #45
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

Yeah, I’m not convinced that there are any models as smart as Fable. Opus 5 definitely isn’t for all it has great benchmark scores. Fable displays judgement in a way I haven’t seen from any other model.

Any models available to us that is...

Re: Grok 4.6

#273

[flagged]

> I’m seriously considering switching to Grok

What's holding you back? According to your post history you've been calling Grok "awesome" for months now: https://news.ycombinator.com/item?id=47988753

Is there any part of Anthropic's offerings that you're struggling to leave behind?

Re: Grok 4.6

#274
post #94

Earlier quoted context omitted.

(I work on Grok) This isn't allowed. CSAM / deepfakes are against our acceptable use policy.

Your company's owner was promoting the feature and joking about it, and called enforcement against it "fascism". CSAM generation kept up for weeks after the initial news articles, and as far as I can tell deepfake generation is still a feature. It's hard to take your AUP seriously here when you've seemingly done nothing technical to actually prevent the action.

CSAM is by definition limited to real imageries and cannot be generated. "Generative CSAM" is like "false true information".

The thing about criticisms that Grok generates "CSAM" images, as well as many similar claims using that acronym, are actually more likely to be intentional mislabeling intending to refer to anime images. Advocates groups with British links love to do it, supposedly to avoid having to name states and/or ethnicity associated with it. which is frustrating because this is how BS like in GP is allowed to exist.

As for deepfakes... 100% they allow it, with weak plausible suggestion feature to decline it. They know that nobody will allow it if given an option. Same deal as Middle Eastern bot spams on Twitter: taking actual measures is against whatever their goals.

Re: Grok 4.6

#275
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts If the prompt guidance is causing the model to be so paranoid about leaking the system prompt... how do we already have it ?

[flagged]

Re: Grok 4.6

#276
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

> benchmark hacking

I think this is the main one. The benchmarks from this are heavily cherry-picked, and they also widely publicised their performance for 4.5 while downplaying the fact the benchmarks were "accidentally" in their training set

Re: Grok 4.6

#277
post #159
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

There is a widespread belief that the nature of intelligence is scalar, like how a person can have 100x more wealth than another person. If this were true, then we’d probably see breakaway RSI from a single lab. But I think we’re discovering that intelligence is about universality, not magnitude. This is analogous to how building a universal Turing machine wasn’t merely a matter of building a calculator that could mu…

But in theory you can make an LLM A LOT faster than a human.

You can also run massive amount of LLMs in parallel.

There might be a limit to a normal LLM but not to theo everall system.

Re: Grok 4.6

#278
post #50

Fable-like intelligence, beats GPT-5.6-Sol on most benchmarks, cheaper than Kimi K3 on API and quite generous usage on Cursor subscription.

I'm thinking of switching to Grok on Cursor (purely for $$ reasons). But Opus >= 4.8 has been fantastic; it's hard to leave, even just to dabble with other models.

I've been using Grok instead of Opus the past few weeks.

It's a downgrade, but barely noticeable for me and totally inconsequential for the amount of work required to fix it and the corresponding $$$ saving.

Re: Grok 4.6

#279
post #184

Earlier quoted context omitted.

China is clearly the US' main adversary. I don't take it personally and I don't believe China is inherently evil or something, but you'd have to be an idiot to be a US citizen and believe that you can trust China more than your own government in any general sense. Just the same, if you're a Chinese citizen and you believe you can trust the US more than your own government, then you're also an idiot. It's not a matter…

thats exactly why a lot of people in europe or america trust china more. enemy governments have zero direct power over you and they dont really want to work together with your government. they cant hurt you, only the country you live in. and with the snowden leaks, epstein files, ICE raids, rising fascism in europe, chat control, genocidal wars in ukraine and palestine, there is no reason to support your country anym…

Ah yes, just as there’s famously no such thing as Russian hackers (for example) given effectively total impunity to scam, defraud, blackmail, etc any company, so long as it’s not located in Russia. No direct harm! Oh wait…

The thing about your own country, especially the more democratic it is, is that there are brakes in the system. A lot of the control mechanisms are indirect, and thus slow and occasionally prone to failure, but the people do have the ultimate say. What you’re doing is looking at failures of the braking system and concluding that brakes don’t even exist! Faulty logic in the extreme.

Re: Grok 4.6

#280

Earlier quoted context omitted.

That's exactly what Anthropic said was going to happen! Their big bet is that models are going to keep getting sharply better, not that they're going to quickly reach a plateau of quality that they can then defend.

They will get sharply better in tasks with verifiable domains... math and coding Gradually the labs will start engineering verifiable sandboxes for wider domains like videogames This strategy will hit a plateau in about 18 months and then we're back to diminishing returns and incremental progress along other dimensions (like accelerated inference using ASICs)

You only mention math, coding and videogames.

They already hire and pay people with research titles for creating and solving problems in their fields.

And a lot of labs say that RL can help everywere and has plenty of way to go.

Post reply on HN