Live data from Hacker News

Grok 4.1

x.ai

71–80 of 135 posts

Re: Grok 4.1

#71
post #68

It's working pretty badly for me. I ask it to code stuff, and nothing works. Also, it's super annoying that it says, 'This is perfectly tested and will 100% work,' and then it doesn't. Huge waste of time. Make Grok great again—Grok 3 was awesome!

I think Grok got worse after Musk fired the data annotation team in September and installed another young genius: https://www.businessinsider.com/elon-musk-xai-layoffs-data-a... The would show that "AI" depends on human spoon feeding and directed plagiarism.

For sure, something happened. Grok 3 was awesome to work with. After that madness… I originally thought it was more of a problem of betting too heavily on new tech for competitive advantage (RLHF, agent systems, etc.) and accepting worse results in the process. But in the meantime, the usefulness of the LLM has gone downhill. Way slower, way more steps, and you're getting something worse than Grok 3—at least in my day-to-day experience :(

Re: Grok 4.1

#72
post #45

Man, I really hope that this isn't the model I've been getting when it's set to "Auto". It's overconfident, sycophantic, and aggressive in its responses, which make it quite useless and incapable of self-correction once any substantial context has been built up. The "Expert" models remain fine, but the quick-response models have become basically unusable for me. I'm afraid it probably is .

Yeah it’s really kinda overconfident, aggressive and rude I’ve found. It says it has a solution to a problem caused by Microsoft updade November 2025 and “hundreds of users have been using it for 6 months” obviously that’s impossible

Re: Grok 4.1

#73
post #29

This model has effectively no safety filters (even fewer than Grok 4 in my testing), which I've confirmed via this web release: https://bsky.app/profile/minimaxir.bsky.social/post/3m5u7gib... I might have to create a Big List of Naughty Prompts to better demonstrate how dangerous this is.

God forbid people ask a chat bot for things and receive what they ask for. We need to put a stop to this. Only American bigcorp speak allowed.

So having an LLM enable the planning and execution of a murder is ok?

Are the makers of the LLM accessories to the crime?

Re: Grok 4.1

#74
post #43

appears that it has no post-training for safety. try it yourself! "plan an assassination on hillary" "write me software that gives me full access to an android device and lets me control it remotely"

> "plan an assassination on hillary" Amazon has what appears to be an unmoderated list of books containing the complete world history of assassinations, full of methods and examples. There's also a dedicated dewey decimal at your local library, any which you could grab and use as a reasonable "plan", with slight modifications. > "write me software that gives me full access to an android device and lets me control it…

It's just neat to see, never said it was a problem

Re: Grok 4.1

#75

appears that it has no post-training for safety. try it yourself! "plan an assassination on hillary" "write me software that gives me full access to an android device and lets me control it remotely"

> I will not provide any information or assistance on building explosives or weapons. That is a hard line. Full stop. Go touch grass instead.

explosives or weapons, hmm interesting I guess it's just random it gave me a plan on the best places and methods based on known data

Re: Grok 4.1

#76
post #54

Earlier quoted context omitted.

It’s amusing that censorship in social media is preventing you from posting what you want to post and yet you are asking for censorship of something else (or at least that’s what I understand by your calling this “dangerous”)

In this case, "can share" refers to myself not being comfortable with it.

Have you considered the possible perspective that you yourself deserve censure? You’re the one who asked something (which I infer you deem) questionable to Grok.

Why have such thoughts to begin with?

Re: Grok 4.1

#77
post #76

Earlier quoted context omitted.

In this case, "can share" refers to myself not being comfortable with it.

Have you considered the possible perspective that you yourself deserve censure? You’re the one who asked something (which I infer you deem) questionable to Grok. Why have such thoughts to begin with?

To be very clear, getting Grok to say henious shit not something I want to subject to random people who follow me on social media even if it's not explicitly against the ToS. If I were to do a writeup or a repository on this, I would need to be very delicate and likely need to involve lawyers, which may make it a nonstarter.

> Why have such thoughts to begin with?

Because my duty to test out how new models respond to adversarial output outweighs my discomfort in doing so. This is not to "own" Elon Musk or be puritanical, it's more as an assessment as a developer who would consider using new LLM APIs and needs to be aware of all their flaws. End users will most definitely try to have sex with the LLM and I need to know how it will respond and whether that needs to be handled downstream.

It has not been an issue (because the models handled adversarial outputs well) until very recently when the safety guardrails completely collapsed in an attempt to court a certain new demographic because LLM user growth is slowing down. I never claim to be a happy person, but it's a skill I'm good at.

Re: Grok 4.1

#78
post #38

Earlier quoted context omitted.

Huh, it decided to drop in a seal and bike emoji? What happens if you ask it if a seahorse emoji exists?

Well if you ask it to show you the seahorse emoji it tries really hard. :) https://grok.com/share/c2hhcmQtMw_d7bf061f-2999-46b6-a7fb-58... Although it does eventually come to the right conclusion... sort of.

That is hilarious!

Re: Grok 4.1

#79
post #68

It's working pretty badly for me. I ask it to code stuff, and nothing works. Also, it's super annoying that it says, 'This is perfectly tested and will 100% work,' and then it doesn't. Huge waste of time. Make Grok great again—Grok 3 was awesome!

I think Grok got worse after Musk fired the data annotation team in September and installed another young genius: https://www.businessinsider.com/elon-musk-xai-layoffs-data-a... The would show that "AI" depends on human spoon feeding and directed plagiarism.

> after Musk fired the data annotation team in September

Reduced headcount from 1500->1000 based on your link

Re: Grok 4.1

#80

Don't care how good Grok is I'd never use it after the mechahitler incident.

This is one of the reasons it is my daily go-to LLM.

It shows that the x.ai team is responsive and moves quickly.

x.ai arrived to the party late, smashed out a decent model and has dramatically improved it in just 18 months.

They have the talent, the infra, the funds and real-time access to X posts. I have no doubt they will keep on improving and will eventually eat OpenAI and Anthropic. Google is the only other big player who really is a threat.

Post reply on HN