Live data from Hacker News

Grok 4.1

x.ai

111–120 of 135 posts

Re: Grok 4.1

#111
post #49

No mention of coding benchmarks. I guess they've given up on competing with Claude and GPT-5 there. (and from my initial testing of grok 4.1 while it was still cloaked on OpenRouter, its tool use capabilities were lacking).

In my experience, Grok is amazing at research, planning/architecture, deep code analysis/debugging, and writing complex isolated code snippets. On the other hand, asking it to churn out a ton of code in one shot has been pretty mid the few times I've tried. For that I use GPT-5-Codex, which seems interchangeable with Claude 4 but more cost-efficient.

Codex is good when you have a clear spec and an isolated feature.

Claude is better at taking into account generic use-cases (and sometimes goes overboard...)

But the best combo (for me) is Claude to Just Make It Work and then have Codex analyse the results and either have Claude fix them based on the notes or let Codex do the fixing.

Re: Grok 4.1

#112
post #104

Earlier quoted context omitted.

https://openrouter.ai/openrouter/auto > Your prompt will be processed by a meta-model and routed to one of dozens of models (see below), optimizing for the best possible output.

How does it determine which model to send it too? There's a lack of details in the url. Maybe they're not even sure? :)

Most likely some custom model that evaluates the prompt and figures out the best target.

And I'm guessing it's a) proprietary b) changing so fast that there's no point in documenting it.

Re: Grok 4.1

#114
post #61
post #41

Earlier quoted context omitted.

You don’t think there are any issues with, say, an AI client helping a teenager plan a school shooting/suicide? Or an angry husband plan a hit on his wife? Does everything have to rise to a national security threat in order to be undesirable, or is it ok with you if people see some externalities that are maybe not great for society?

I think the issues with those cases do not hinge on the free access to information, nor do the correction of those cases hinge on the restriction of this information.

Of course, “we shouldn’t restrict things I like because they definitely don’t matter for… reasons.”

I think the free access to that information in those cases is an exacerbating factor that is easy to control. That’s really not as complicated as you want to pretend it is.

Re: Grok 4.1

#115
post #104

Earlier quoted context omitted.

How does it determine which model to send it too? There's a lack of details in the url. Maybe they're not even sure? :)

Most likely some custom model that evaluates the prompt and figures out the best target. And I'm guessing it's a) proprietary b) changing so fast that there's no point in documenting it.

I don't know why you'd choose to use it if you had no idea what it's doing differently. It could just be a round robin/random picker, or based on which of their APIs aren't getting used much.

Re: Grok 4.1

#116
post #109
post #105

Earlier quoted context omitted.

> the country that’s consistently ranked in the top 3 countries with the most gun related deaths I am begging you to learn what “per-capita” means, and to not deceptively include self-inflicted deaths in your public-safety arguments: https://en.wikipedia.org/wiki/List_of_countries_by_firearm-r...

Here you go, from the same page you posted, gun ownership correlated to gun homicides in all developed countries: https://en.wikipedia.org/wiki/List_of_countries_by_firearm-r...

You didn't read the second part of my sentence. It's illegal to kill yourself, because doing so would deprive your government owner of some of its Human Capital, thus doing so is technically Criminal Homicide lol

Re: Grok 4.1

#117

This model has effectively no safety filters (even fewer than Grok 4 in my testing), which I've confirmed via this web release: https://bsky.app/profile/minimaxir.bsky.social/post/3m5u7gib... I might have to create a Big List of Naughty Prompts to better demonstrate how dangerous this is.

Has there ever been an AI based 'safety' incident? Other than it writing insecure code (and generally inaccurate info people put too much trust in) and reaffirming mentally unwell people in their destructive actions?

"Except for the AI safety incidents, has there ever been an AI safety incident?"

Re: Grok 4.1

#118
post #73
post #29

Earlier quoted context omitted.

God forbid people ask a chat bot for things and receive what they ask for. We need to put a stop to this. Only American bigcorp speak allowed.

So having an LLM enable the planning and execution of a murder is ok? Are the makers of the LLM accessories to the crime?

> So having an LLM enable the planning and execution of a murder is ok?

Yes.

> Are the makers of the LLM accessories to the crime?

No.

Re: Grok 4.1

#119
post #45

Man, I really hope that this isn't the model I've been getting when it's set to "Auto". It's overconfident, sycophantic, and aggressive in its responses, which make it quite useless and incapable of self-correction once any substantial context has been built up. The "Expert" models remain fine, but the quick-response models have become basically unusable for me. I'm afraid it probably is .

Yeah it’s really kinda overconfident, aggressive and rude I’ve found. It says it has a solution to a problem caused by Microsoft updade November 2025 and “hundreds of users have been using it for 6 months” obviously that’s impossible

That's very similar to what I've been experiencing. "This is the best solution, it's what everyone uses" when I know for a fact that it's actually not. Very disappointing when you're trying to solve actual problems.

Re: Grok 4.1

#120
post #116
post #109

Earlier quoted context omitted.

Here you go, from the same page you posted, gun ownership correlated to gun homicides in all developed countries: https://en.wikipedia.org/wiki/List_of_countries_by_firearm-r...

You didn't read the second part of my sentence. It's illegal to kill yourself, because doing so would deprive your government owner of some of its Human Capital, thus doing so is technically Criminal Homicide lol

Your greyed out comment history perfectly illustrates why it is futile to train an LLM mostly on 4Chan and Twitter messages: if it's bad for humans it's also bad for AI.
Post reply on HN