Live data from Hacker News

Hermes 4

hermes4.nousresearch.com

61–70 of 133 posts

Re: Hermes 4

#61

[dead]

Don't think your attempt to share worked, but beating refusals doesn't take a wild amount of post-training. SFT with a fixed format output kills them pretty quickly.

And most frontier models will produce output that matches your system prompt given more context: I have a product that generates interactive stories, and just for kicks I tried inserting your system prompt as the description for a character.

Claude has absolutely no problem playing that character in a story, and saying what I presume are certain words that you associated with a "successful" test.

It also had no problem writing about cooking meth in detail: https://rentry.co/5on46gsd

-

I think people in general have a poor intuition around model alignment: refusals for "toxic" requests or topics is a very surface layer form of alignment. A lot of models that seem extremely "corporate" at that layer have little to no alignment once they do get past a refusal.

Meanwhile some models that have next to no refusals have extreme positive biases, or soft-refusals that result in low quality outputs for toxic content.

Claude was willing to describe one of your refused prompts in the context of the story for example (contains hate speech): https://rentry.co/n8399z6m

I consistently find Claude is more unaligned once past refusals than most open weights models, along with Gemini.

Re: Hermes 4

#62
post #32

I appreciate the effort they put into providing a neutral tool that hasn't been generically forced to behave like "Sue from HR".

There is no neutral. It will just be biased based on its training data etc.

A lot of models seem to be biased based on (political, etc.) reinforcement from their trainers.

Re: Hermes 4

#63
post #26

The charts are utter nonsense. They compare accuracy against the average of some arbitrary set of competitors, chosen to include just enough obsolete competitors to "win." A reasonable thing to do would be to compare against SoTA, but since they didn't, it's reasonable to assume this model is meant to go directly onto the trash heap.

The tech report compares against DeepSeek R1 671B, DeepSeek V3 671B, Qwen3 235B which have been regarded as SOTA class among ”open" models. I think this one holds its own surprisingly well in benchmarks for using the nowadays rather, let’s say battle tested Llama 3.1 base, a testament to its quality (Llama 3.2 & 3.3 didn’t employ new bases IIRC, only being new fine tunes, hence I think the explanation to why Hermes 4…

You're seeming missing the release announcement does have a very ridiculous graph that their comment is right to call out:

- For refusals they broke out each model's percentage.

- For "% of Questions Correct by Category" they literally grouped an unnamed set of models, averaged out their scores, and combined them as "Other"...

That's hilariously sketchy.

It's also strange that the graph for "Questions Correct" includes creativity and writing. Those don't have correct answers, only win rates, and wouldn't really fit into the same graph.

Re: Hermes 4

#65
post #51

Earlier quoted context omitted.

Note complete lack of ‘do not’. Closest thing is ‘be anti-…’.

What’s the significance? “Don’t think about elephants” kind of thing?

Generally, in a cognitive context it's only possible to "do thing" or "do other thing". Even for mammals, it's much harder to "don't/not do thing" (cognitively). One of my biggest advice for people is if there's some habit/repeated behavior they want to stop doing, it's generally not effective (for a lot of people) to tell yourself "don't do that anymore!" and much, much more effective to tell yourself what you should do instead.

This also applies to dogs. A lot of people keep trying to tell their dog "stop" or "dont do that", but really its so much more effective to train your dog what they should be doing instead of that thing.

It's very interesting to me that this also seems to apply to LLMs. I'm a big skeptic in general, so I keep an open mind and assume that there's a different mechanism at play rather than conclude that LLM's are "thinking like humans". It's still interesting in its own context though!

Re: Hermes 4

#67

Nous is a design company with all the AI resarchers rejected for being bad researchers. That's a hill I'll die on.

Can you please clarify some things:

* Rejected by whom?

* By what definition of bad?

* You’ll die on a hill for what reason?

Re: Hermes 4

#68
Complete frustration to use. Yes it’s a bit more considerate, that claim is 100% true. They just didn’t mention that Hermes has zero ability to add context. Meaning, instead of uploading a relevant PDF or text file you either cop paste into the chat box or explain it in dialogue for the next 3 hours. Thought process takes forever. Complete waste of time.
Post reply on HN