Earlier quoted context omitted.
Enforce it then
We are and will continue to.
Grok 4.6
321–330 of 696 posts
Re: Grok 4.6
#322Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…
"you may find and fix vulnerabilities in local codebases only" This seems like a bad idea, what does local mean? Anything Grok can access locally? This seems like asking for trouble.
Re: Grok 4.6
#323As polarizing as grok is, it was basically inevitable for it to start being a real competitor given how much investment SpaceX made into its own inference capabilities. Seems if you are okay with it, there's no reason to use anything but the highest effort levels of some other frontier models for the price. I think Grok provides healthy competition to the other labs, though I do think they bank on groks reputation ma…
Curious - what is the main issue you find polarizing with grok?
Grok was supposed to be the unbiased model, that is: regurgitate everything it has read. Obviously all data has bias, even all of the data at once, but the sales pitch was that you would get that unfiltered. At least in open source models, this has been shown to improve the competence of the model.
Then this happened: https://futurism.com/artificial-intelligence/grok-describes-...
So not only has bias been introduced, but they are happily biasing it for trivial reasons. So now the model needs to be competitive in exactly the same way that others are: on benchmarks (which are still not a solved problem).
But, I (and many others) disagree with how Elon has behaved politically and don't want to hand money over to him, so all of that is a hypothetical.
Re: Grok 4.6
#324>Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass. As a designer, I'm always hesitant to believe these statements until there's independent comparisons between the old & new model, as well as comparisons to human made flows. Design can be so…
Re: Grok 4.6
#325Earlier quoted context omitted.
So basically, nothing that actually affects working with it in August 2026. Got it. Facebook has a far longer (and worse) laundry list of offenses and I'm sure you still use it. Or Threads, or Instagram. > My organization has outright banned Grok That's too bad, as it's currently the only model that won't consistently flag honest good-actor security questions, in my experience. So I'd ask you who you work for, but I…
The chinese are also mostly fine with good-actor security questions. Maybe even to comfortable.
Re: Grok 4.6
#326Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…
I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible for Chinese models to close the gap that easily. It's also very unlikely that they sell the same public models to their private customers (military etc). We also know they talk about "unpublished internal models" for things like the last HuggingFace hacking incident. So it's not a bad theory.
Re: Grok 4.6
#327Earlier quoted context omitted.
So basically, nothing that actually affects working with it in August 2026. Got it. Facebook has a far longer (and worse) laundry list of offenses and I'm sure you still use it. Or Threads, or Instagram. > My organization has outright banned Grok That's too bad, as it's currently the only model that won't consistently flag honest good-actor security questions, in my experience. So I'd ask you who you work for, but I…
You are only strawmanning around. Apparently you are unable to coprehend that other peole have values.
*people
Also, that's not what strawmanning is. I never denied that Grok didn't act bizarrely offensively over a fucking year and a half ago (so did other LLMs, btw... and so have many other experiments over the years, remember Microsoft's?), which is an eternity in this space. I know Musk is polarizing, but give me a fucking break. Don't assume malice when social incompetence serves as an exculpatory factor.
Apparently, you are unable to comprehend that your opinion of things has been tainted away from the truth by an algorithm incentivized to outrage you. That what you call your "values" are, in fact, driven by someone else's greed for eyeball attention. Do you think civilizations that become anti-Western-values over time are more driven by facts and empiricism, or by catchy slogans that twist the truth and a media that uses cherry-picked examples which immediately trigger emotions?
Re: Grok 4.6
#328Re: Grok 4.6
#329Earlier quoted context omitted.
I use both Grok 4.5 and Opus 5. They’re both very good and Grok is faster and cheaper.
Opus 5 is terrible. I'd even say it's a step backwards from 4.8. I'm getting high error rates from it, and then it catches the error, and then it sometimes errors the error fix (!). Just today I had to switch another agent to Fable with the instruction, "Please clean up the mess that Opus 5 made, thanks" The other day, Sol called Opus 5's handoff (a skill I have that is basically a compaction, but just written to a f…
Anthropic does this all the time (ruins their models for users) while they screw around with system prompts. Oh but it's for your own good of course! They know what's best for us all, if we would just give them a monopoly.
I can't wait until OpenAI/Grok/Chinese models surpass them enough that their main character syndrome and smug doomerism no longer draws much media attention.
Re: Grok 4.6
#330Earlier quoted context omitted.
> It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually. What are suspicious of? If the timing is similar maybe just everyone already are of similar capabilities and got there at a similar time? > Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? It means Anthropic had no…
Well, Opus 5 and Fable are the only models I don’t constantly swear at and call stupid, which seems like a pretty good moat to me. My guess is all the commenters (you are the 4th person I’ve seen say this) saying ‘Anthropic has no moat’ haven’t actually used Fable or even Opus 5 yet. Sol is laughable by comparison, and Grok… lol.
I've used plenty of Opus and Fable. Still do.
> Sol is laughable by comparison
Not really, it depends. Sol is better and useful in some areas. Definitely not all.
Fable is gimped just by those "guardrails" that silently downgrades you to Opus 4.8. Not only do you pay extra for Fable but your caching can be easily messed up. It also doesn't just find all the bugs or is bug-free. Sol has spotted lots of Fable issues and vice versa. Fable also costs 2-100x as much.
> I don’t constantly swear at and call stupid
That's not a judge of anything. There are models that may be stupid and you can swear at it, but if they still get the job done for 1/10th the price... maybe that's all you're paying for.