[flagged]
Grok 4.6
231–240 of 696 posts
Re: Grok 4.6
#232[flagged]
Re: Grok 4.6
#233Earlier quoted context omitted.
> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…
Mmm, quite. > I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. My vote is "machine psychology".
Re: Grok 4.6
#234Earlier quoted context omitted.
I use both Grok 4.5 and Opus 5. They’re both very good and Grok is faster and cheaper.
Opus 5 is terrible. I'd even say it's a step backwards from 4.8. I'm getting high error rates from it, and then it catches the error, and then it sometimes errors the error fix (!). Just today I had to switch another agent to Fable with the instruction, "Please clean up the mess that Opus 5 made, thanks" The other day, Sol called Opus 5's handoff (a skill I have that is basically a compaction, but just written to a f…
Re: Grok 4.6
#235Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…
asked grok to give a compute estimate for each: - SpaceX / xAI: ~1.4 GW (owned Colossus clusters) - OpenAI: ~2–3 GW (mostly rented/cloud) - Anthropic: ~1.5–2.5 GW (multi-cloud + xAI lease)
chatgpt estimates a lower: - OpenAI: ~1.5M H100-eq ± ~0.8M - Anthropic: ~1.4M H100-eq ± ~0.7M - SpaceX/xAI: ~0.6M H100-eq ± ~0.3M
but it felt obligated to mention that "for single tightly interconnected NVIDIA training clusters, SpaceX/xAI has been unusually strong."
Re: Grok 4.6
#236>Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass. As a designer, I'm always hesitant to believe these statements until there's independent comparisons between the old & new model, as well as comparisons to human made flows. Design can be so…
(I work on Grok) We've been working on teaching the model how to reason about great visual design principles. Obviously this is hard and somewhat subjective, but through a combination of writing down these principles (e.g. how to think about systems, not just "use this italic serif font on marketing pages"), and then creating a lot of data to pairwise compare designs/outputs, we've made a notable improvement over G4.…
Re: Grok 4.6
#237Earlier quoted context omitted.
> Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts If the prompt guidance is causing the model to be so paranoid about leaking the system prompt... how do we already have it ?
System prompts are more like suggestions than hard constraints.
That's the joy and pain.
Re: Grok 4.6
#238Earlier quoted context omitted.
I'd start here: https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a... And here: https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here…
So basically, nothing that actually affects working with it in August 2026. Got it. Facebook has a far longer (and worse) laundry list of offenses and I'm sure you still use it. Or Threads, or Instagram. > My organization has outright banned Grok That's too bad, as it's currently the only model that won't consistently flag honest good-actor security questions, in my experience. So I'd ask you who you work for, but I…
Re: Grok 4.6
#239In terms of using experience, I found Grok 4.5 to be way more pleasant to use than GPT 5.6 Sol and Claude 4.8/5. It just gets to the point, and is super fast and concise, no yapping. That's how AI agents should be imo. None of the weird "Claude ipsum" jargon like "load-bearing" and "stale folklore" or GPT 5.6-isms like "focused regression" and "provenance".
Re: Grok 4.6
#240Earlier quoted context omitted.
Elon Musk crediting his engineers: 1. "Please put in bold letters my quote that what people experience in the cars is the result of a large number of extremely talented engineers working very hard. Please give me the least credit." https://cleantechnica.com/2020/08/15/tesla-autopilot-innovat... 2. "It is extremely important to emphasize that Tesla Autopilot is the work of 300 super talented engineers." https://cleant…
Your source for his intelligence is that his employees glaze him? Surely one of the smartest people in the world has written or published something groundbreaking, right? Surely his sole intellectual contribution isn't shitposting on Twitter?
> At least Gates was honest that he "surrounded himself with smart people"
By reading the parent of a comment you can follow the conversation without needing to ask multiple questions.