Earlier quoted context omitted.
[flagged]
This is a matter of politics; it's a matter of reputation. I'm fine with using AI tools offered by companies like OpenAI, Anthropic, and Google despite knowing that these companies are ran by billionaires who are much more aligned, politically, to Musk than they are with me. What I'm not fine with is handing over valuable data to a guy that has literally completely captured the US government and has shown a disdain f…
Grok 4.6
401–410 of 696 posts
Re: Grok 4.6
#402Does anyone know how the grok allowances compare to OpenAI / Anthropic for the monthly plans? I heard they're not generous, which means I never really bother testing Grok.
Re: Grok 4.6
#403Earlier quoted context omitted.
What do you think a human brain is…
This is like saying the person you see in the mirror is categorically a human being because both of you produce similar reflections of light rays
Re: Grok 4.6
#404I stopped bothering with Grok for anything when 4.5 dropped. It was so awful that I figured Elon had given up and was going to give alll his compute to Anthropic. I’m extremely sceptical anyways - Grok 4.5 was probably the worst model I ever seriously tried to use going back 3 years.
Fast, speaks normally. Was able to figure out many issues Claude couldn’t. I thought code readability was a worse than Claude but I could just tell it how I wanted stuff written anyway.
What do you use it for? I’m genuinely curious. I’m also using it in cursor
Re: Grok 4.6
#405I didn’t expect we get 4.6 so soon and the increased limits to try it out are neat!
Re: Grok 4.6
#406https://aibenchy.com/compare/qwen-qwen3-8-2-4t-a95b-low/x-ai...
Re: Grok 4.6
#407Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…
The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability. All you need to have Fable-level AI is to announce it, and have enough fans shift from insisting that model Y is the best now, way better than model X.
Couldn't be further from the truth. The models can be tested and statistically evaluated.
I ran a massive Fable max code review on my lone lisp codebase. Now that I have switched to OpenAI, I decided to run an equivalent review using Sol max and compare them. I'm keeping all data so I can thoroughly evaluate their performance in multiple areas such as correctness, rigor, performance, security, maintainability, consistency, among others.
Fable pass is 100% done and I'm around 70% done with the Sol pass. Preliminary results are already becoming clear: Sol is capable of reproducing around 70% to 90% of Fable's performance. Haven't tested open weight models but I'd wager they have the same performance as Sol if not lower.
It seems Fable is still king, I'm afraid. It's undeniable that OpenAI is providing huge value here: up to 90% Fable performance at multiple times the usage on a subscription than what Anthropic offers us is a phenomenal deal. However, if one desires the best model, to me it looks like Fable is still it.
Re: Grok 4.6
#408Does really well and ~2x cheaper than Qwen3.8 2.4T, they have same pricing but grok is around 2x more token efficient: https://aibenchy.com/compare/qwen-qwen3-8-2-4t-a95b-low/x-ai...
https://aibenchy.com/compare/openai-gpt-5-6-sol-low/x-ai-gro...
Re: Grok 4.6
#409Earlier quoted context omitted.
I think it is fair to argue that prompts are not a safety layer at all and can't be relied upon for much. "Make no mistakes"
Yes that's been obvious since the beginning. That's why you should always monitor your agents closely. Just like supervised self driving cars, you have to watch the road and do some hand holding. The tooling around isolation, logging, and real time security/anonomly detection for regular LLM laptop users is very immature right now. I expect that to change soon. The alternative is extremely locked down models which is…
But if it's so obvious, then why are we still relying on it in the system prompt. It's just wasting context at this point.
Re: Grok 4.6
#410Earlier quoted context omitted.
not suspicious at all. They are all doing the same scaling of test time, training data so getting similar results. anyone with access to capital can produce frotier model. hell you can just ask chatgpt how to create a fontier model. recipe is not a secret despite what these 'labs' pretend
Google is not able to currently produce a frontier model despite all the capital.