Live data from Hacker News

Grok 4.6

x.ai

331–340 of 696 posts

Re: Grok 4.6

#331
post #312
post #306

Earlier quoted context omitted.

It was said at the time that xAI acquiring Cursor was very smart because it would give them access to years of agent coding traces from millions of users. $60B in SpaceX stock for Cursor was a bargain Data + compute + being competent and smart enough to ship. fwiw I don't think these are yet Fable level - the difference tends to get discovered in the long tail of tasks - but they're close enough, they're cheap, and t…

> $60B in SpaceX stock for Cursor was a bargain Not if you go by financial fundamentals. All of Space X only has around $18B in sales.

yes but it was a stock deal - so they bought it using spacex bucks

Re: Grok 4.6

#332
post #129

In terms of using experience, I found Grok 4.5 to be way more pleasant to use than GPT 5.6 Sol and Claude 4.8/5. It just gets to the point, and is super fast and concise, no yapping. That's how AI agents should be imo. None of the weird "Claude ipsum" jargon like "load-bearing" and "stale folklore" or GPT 5.6-isms like "focused regression" and "provenance".

What harness are you using?

I've been using the grok cli, and this is what I love about it.

Re: Grok 4.6

#333
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

not suspicious at all. They are all doing the same scaling of test time, training data so getting similar results.

anyone with access to capital can produce frotier model. hell you can just ask chatgpt how to create a fontier model. recipe is not a secret despite what these 'labs' pretend

Re: Grok 4.6

#334
post #291
post #94

Earlier quoted context omitted.

(I work on Grok) This isn't allowed. CSAM / deepfakes are against our acceptable use policy.

[flagged]

So, what do you do for a living?

You've made a personal attack and seem to be under the impression you're morally superior. So, I'm curious as to what highly virtuous role you take on in your daily life.

That said, I see your comment history is a lot of one sentence personal attacks against people. Not a lot of thoughtful debate.

This makes hypocrisy out of your supposed concern for social good.

Re: Grok 4.6

#335
post #172

Earlier quoted context omitted.

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

These system prompts are not the only safety layer that these models use. There's other more deterministic filters in place both on input and (streaming) output.

I think it is fair to argue that prompts are not a safety layer at all and can't be relied upon for much.

"Make no mistakes"

Re: Grok 4.6

#336
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

"you may find and fix vulnerabilities in local codebases only" This seems like a bad idea, what does local mean? Anything Grok can access locally? This seems like asking for trouble.

> This seems like a bad idea, what does local mean? Anything Grok can access locally? This seems like asking for trouble.

It means you put "i.swear.this.is.localhost [remote ip]" in your hosts file.

Re: Grok 4.6

#337
post #313

Earlier quoted context omitted.

CSAM is by definition limited to real imageries and cannot be generated. "Generative CSAM" is like "false true information". The thing about criticisms that Grok generates "CSAM" images, as well as many similar claims using that acronym, are actually more likely to be intentional mislabeling intending to refer to anime images. Advocates groups with British links love to do it, supposedly to avoid having to name state…

> CSAM is by definition limited to real imageries and cannot be generated. Where in the definition does it imply this?

I thought that's just the legal definition in any sufficiently developed countries?

Re: Grok 4.6

#338
post #159

Earlier quoted context omitted.

There is a widespread belief that the nature of intelligence is scalar, like how a person can have 100x more wealth than another person. If this were true, then we’d probably see breakaway RSI from a single lab. But I think we’re discovering that intelligence is about universality, not magnitude. This is analogous to how building a universal Turing machine wasn’t merely a matter of building a calculator that could mu…

You're take basically lines up with Francois Chollet: https://arxiv.org/abs/1911.01547 intelligence is more like polishing a ball smooth than growing the ball to infinity. For many tasks, it will be smooth enough.

At a certain point the roughness of the ball reaches a size threshold where the imperfections are smaller than the wavelength of light, and the surface takes on a glassy smoothness. Intelligence has similar milestones, almost like phase changes, I think, where capabilities are reached. Maybe it's like a superposition of many small step functions.

Re: Grok 4.6

#339
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

I'm sure Grok 4.6 is not Fable level. Benchmarks are almost useless. Having said that, Grok 4.6 (1.5T params) is without a doubt way smaller than Fable, maybe a Fable sized Grok would be Fable level?

Supposedly grok 4.7 is a 5T model. But that's in musk units, so im not sure.
Post reply on HN