Live data from Hacker News

Grok 4.6

x.ai

591–600 of 696 posts

Re: Grok 4.6

#591
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

"you may find and fix vulnerabilities in local codebases only" This seems like a bad idea, what does local mean? Anything Grok can access locally? This seems like asking for trouble.

Seems clear to me it means don't go trying to change things over the internet.

Isn't it pretty standard to consider "local" to mean not remote or external? Local storage means storage on the machine, not attached via network or plugged into an external port. Localhost is the ip for the computer in question, not a remote one.

Re: Grok 4.6

#592

Earlier quoted context omitted.

This is like saying the person you see in the mirror is categorically a human being because both of you produce similar reflections of light rays

The person I see in a mirror is a human being. The person I see in a mirror is me. What do you think a mirror is?

The image you see in the mirror is a reflection of a human. Or, more precisely, a 2-dimensional projection of the frontal outer surface of a human.

One half of one dimension less than a human. But sure looks convincing on the surface.

Re: Grok 4.6

#593
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

I predicted this exact event several months before Fable. ,not in a provable way, but the reasoning was related to a paper I read from here that I basically self-internalized as variability knowledge. Two very similar papers, one unfortunately named.

I also stated recently (in informal conversation), based on the performance posted, that said variability was only applied to specific fields of information.

So allow me to make a more provable prediction:

There will be another significant jump related to full field converage, followed by another and from there (we'll call this v3), it will then be capable of automating ASI.

Re: Grok 4.6

#594
post #593
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

I predicted this exact event several months before Fable. ,not in a provable way, but the reasoning was related to a paper I read from here that I basically self-internalized as variability knowledge. Two very similar papers, one unfortunately named. I also stated recently (in informal conversation), based on the performance posted, that said variability was only applied to specific fields of information. So allow me…

There's some spelling mistakes. To avoid an edit. I would say it won't take 2 years.

Re: Grok 4.6

#595
post #172

Earlier quoted context omitted.

These system prompts are not the only safety layer that these models use. There's other more deterministic filters in place both on input and (streaming) output.

Is there concrete evidece that those are xAI's default prompts anyway? They seem plausible enough but how would company outsiders know?

I've spent the last year working as an annotator/evaluator for DataAnnotation. All the frontier/flagship model providers use independent contractors for iterating on their LLMs. I'm not able to tell you which models I've worked on as a term of my NDA.

The system prompt seems plausible, but in my experience they are much much much much longer and more verbose.

Re: Grok 4.6

#596
post #369

Earlier quoted context omitted.

It’s equivalent to having client-side input validation. Yes it can easily be bypassed, but in the vast majority of cases where users aren’t malicious it gets the job done quickly and cheaply.

But isn't the entire point of that system prompt to stop the malicious users. The majority of users are not going to ask those requests anyway.

A locked door stops the lazy thieves, and the lazy thieves are the most common ones.

Re: Grok 4.6

#597

Earlier quoted context omitted.

Because the primary is right handed. I enjoy asking my grandkids why mirrors reflect left to right and not top to bottom.

They do "reflect top to bottom". Hold a written word in front of your eyes to read it. Now flip it to the mirror to read the reflection: Did you flip it horizontally? Then it reflected left-to-right. Did you flip it vertically? Then it reflected top-to-bottom.

I think "reflect top to bottom" is intended to mean "swap top and button". A mirror reflects left, right, top and bottom perfectly. It's front and back that it swaps.

Someone saying that a mirror swaps left and right is comparing it to a photograph, and only because we, as bipedal creatures, really prefer to orient images of other humans with heads up.

Re: Grok 4.6

#598
post #50

Fable-like intelligence, beats GPT-5.6-Sol on most benchmarks, cheaper than Kimi K3 on API and quite generous usage on Cursor subscription.

I'm thinking of switching to Grok on Cursor (purely for $$ reasons). But Opus >= 4.8 has been fantastic; it's hard to leave, even just to dabble with other models.

If it's purely about $$, what about the newer open models. DeepSeek and Kimi are roughly equivalent performance for a hell of a lot cheaper.

Re: Grok 4.6

#599
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

4) Elon has access to some Anthropic models because Dario is desperate for compute and bought some from a competitor

Re: Grok 4.6

#600
post #525
post #377

Earlier quoted context omitted.

No it was just announced then but it's still going through regulatory/antitrust

There is an approximately zero probability that someone donating hundreds of millions of dollars to Super PACs in support of the most vain and corrupt president in US history will be held up by anti-trust enforcement. There's not a lot of reason for them to keep arms length at this point.

While you’re probably correct on this one, I thought the same thing about them ending the electric car rebates
Post reply on HN