Live data from Hacker News

Grok 4.6

x.ai

401–410 of 696 posts

Re: Grok 4.6

#401

Earlier quoted context omitted.

[flagged]

This is a matter of politics; it's a matter of reputation. I'm fine with using AI tools offered by companies like OpenAI, Anthropic, and Google despite knowing that these companies are ran by billionaires who are much more aligned, politically, to Musk than they are with me. What I'm not fine with is handing over valuable data to a guy that has literally completely captured the US government and has shown a disdain f…

Thats a very naive position to have my friend. You AT LEAST know Musk position/stance. What about the others CEOs? you never hear them talking about politics, perhaps some of them are 10x worst in ideology and bad influence against your values and you will buy and give your data to them without knowing. Ultimately, you are using a judgment you objetivelly can NOT do, nor trully compare the other options, so you are complicating and perhaps and ultimatelly being the grain of sand of a sand storm that can do damage to probably the most important entrepreneur of our time, who achieved so much. I Rather give him resources to achieve something, because the guy does eventually deilvers.

Re: Grok 4.6

#402
post #389

Does anyone know how the grok allowances compare to OpenAI / Anthropic for the monthly plans? I heard they're not generous, which means I never really bother testing Grok.

Cursor is very generous atm, you get a ton of Grok usage and then your monthly subscription cost in api pricing for Claude and GTP.

Re: Grok 4.6

#403

Earlier quoted context omitted.

What do you think a human brain is…

This is like saying the person you see in the mirror is categorically a human being because both of you produce similar reflections of light rays

The person I see in a mirror is a human being. The person I see in a mirror is me. What do you think a mirror is?

Re: Grok 4.6

#404
post #292

I stopped bothering with Grok for anything when 4.5 dropped. It was so awful that I figured Elon had given up and was going to give alll his compute to Anthropic. I’m extremely sceptical anyways - Grok 4.5 was probably the worst model I ever seriously tried to use going back 3 years.

I feel like I’m im a different universe than you… 4.5 is one of my favorite models of all time.

Fast, speaks normally. Was able to figure out many issues Claude couldn’t. I thought code readability was a worse than Claude but I could just tell it how I wanted stuff written anyway.

What do you use it for? I’m genuinely curious. I’m also using it in cursor

Re: Grok 4.6

#405
Grok 4.5 was the first time I considered giving me $100/mo to xai (currently on just SuperGrok). It’s just a very pleasant model to work with: fast, to the point, intelligent. It’s also much better in UI compared to gpt. Not as good as Claude but close!

I didn’t expect we get 4.6 so soon and the increased limits to try it out are neat!

Re: Grok 4.6

#407
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability. All you need to have Fable-level AI is to announce it, and have enough fans shift from insisting that model Y is the best now, way better than model X.

> The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability.

Couldn't be further from the truth. The models can be tested and statistically evaluated.

I ran a massive Fable max code review on my lone lisp codebase. Now that I have switched to OpenAI, I decided to run an equivalent review using Sol max and compare them. I'm keeping all data so I can thoroughly evaluate their performance in multiple areas such as correctness, rigor, performance, security, maintainability, consistency, among others.

Fable pass is 100% done and I'm around 70% done with the Sol pass. Preliminary results are already becoming clear: Sol is capable of reproducing around 70% to 90% of Fable's performance. Haven't tested open weight models but I'd wager they have the same performance as Sol if not lower.

It seems Fable is still king, I'm afraid. It's undeniable that OpenAI is providing huge value here: up to 90% Fable performance at multiple times the usage on a subscription than what Anthropic offers us is a phenomenal deal. However, if one desires the best model, to me it looks like Fable is still it.

Re: Grok 4.6

#409
post #349

Earlier quoted context omitted.

I think it is fair to argue that prompts are not a safety layer at all and can't be relied upon for much. "Make no mistakes"

Yes that's been obvious since the beginning. That's why you should always monitor your agents closely. Just like supervised self driving cars, you have to watch the road and do some hand holding. The tooling around isolation, logging, and real time security/anonomly detection for regular LLM laptop users is very immature right now. I expect that to change soon. The alternative is extremely locked down models which is…

> Yes that's been obvious since the beginning

But if it's so obvious, then why are we still relying on it in the system prompt. It's just wasting context at this point.

Re: Grok 4.6

#410
post #385

Earlier quoted context omitted.

not suspicious at all. They are all doing the same scaling of test time, training data so getting similar results. anyone with access to capital can produce frotier model. hell you can just ask chatgpt how to create a fontier model. recipe is not a secret despite what these 'labs' pretend

Google is not able to currently produce a frontier model despite all the capital.

Google is providing more TPUs to SpaceX and Anthropic than to its very own DeepMind. Most of that capital investment is going to Cloud, not frontier model development.
Post reply on HN