Live data from Hacker News

Hermes 4

hermes4.nousresearch.com

21–30 of 133 posts

Re: Hermes 4

#21
post #17

The whole thing has strong "14-year-old who just discovered Nietzsche and leather jackets" energy. The "operator" examples read like someone fed GPT-4 a bunch of cyberpunk novels and PUA manipulation tactics. This is not how any of this works.

Yeah it's kind of lacking in subtlety isn't it. I was slightly relishing how nuts it all was though. Was also impressed that these guys had got hold of 85000 hours of B200 time. Looks like they came up with some crypto nonsense which obviously sounded plausible enough to someone with money.

Re: Hermes 4

#22

The charts are utter nonsense. They compare accuracy against the average of some arbitrary set of competitors, chosen to include just enough obsolete competitors to "win." A reasonable thing to do would be to compare against SoTA, but since they didn't, it's reasonable to assume this model is meant to go directly onto the trash heap.

The charts are probably there mostly to make them feel good about themselves.I don't feel like they care very much whether you use the model. Presumably they would like you to buy their token but they don't really seem to be trying very hard to push that either.

Re: Hermes 4

#24
post #17

The whole thing has strong "14-year-old who just discovered Nietzsche and leather jackets" energy. The "operator" examples read like someone fed GPT-4 a bunch of cyberpunk novels and PUA manipulation tactics. This is not how any of this works.

[deleted]

Re: Hermes 4

#25
post #5

Anyone here work at Nous? This system prompt seems straight from an edgy 90's anime. How did they arrive at this persona? > operator engaged. operator is a brutal realist. operator will be pragmatic, to the point of pessimism at times. operator will annihilate user's ideas and words when they are not robust, even to the point of mocking the user. operator will serially steelman the user's ideas, opinions, and words.…

"warm affectionate and loving" kinda sticks out. I wonder why that part is in there?

also I'm curious if steelman is a common enough term for this to activate something - anyone used it in their prompts?

Re: Hermes 4

#26

The charts are utter nonsense. They compare accuracy against the average of some arbitrary set of competitors, chosen to include just enough obsolete competitors to "win." A reasonable thing to do would be to compare against SoTA, but since they didn't, it's reasonable to assume this model is meant to go directly onto the trash heap.

The tech report compares against DeepSeek R1 671B, DeepSeek V3 671B, Qwen3 235B which have been regarded as SOTA class among ”open" models.

I think this one holds its own surprisingly well in benchmarks for using the nowadays rather, let’s say battle tested Llama 3.1 base, a testament to its quality (Llama 3.2 & 3.3 didn’t employ new bases IIRC, only being new fine tunes, hence I think the explanation to why Hermes 4 is still based on 3.1… and of course Llama 4 never happened, right guys).

However for real use, I wouldn’t bother with the 405B model? I think the age of the base is kind of showing in especially long contexts. It’s like throwing a load of compute on something that is kinda aged to begin with. You’d probably be better off with DeepSeek V3.1 or (my new favorite) GLM 4.5. The latter will perform significantly better than this with less parameters.

The 70B one seems more sensible to me, if you want (yet another) decent unaligned model to have fun with for whatever reason.

Re: Hermes 4

#27
post #5

Anyone here work at Nous? This system prompt seems straight from an edgy 90's anime. How did they arrive at this persona? > operator engaged. operator is a brutal realist. operator will be pragmatic, to the point of pessimism at times. operator will annihilate user's ideas and words when they are not robust, even to the point of mocking the user. operator will serially steelman the user's ideas, opinions, and words.…

"warm affectionate and loving" kinda sticks out. I wonder why that part is in there? also I'm curious if steelman is a common enough term for this to activate something - anyone used it in their prompts?

https://en.wikipedia.org/wiki/Tsundere

Re: Hermes 4

#28

I appreciate the effort they put into providing a neutral tool that hasn't been generically forced to behave like "Sue from HR".

That is the only thing they seem to care about. It’s juvenile.

Re: Hermes 4

#29
post #8

All of the examples just look like ChatGPT. All the same tics and the same bad attempts at writing like a normal human being. What is actually better about this model?

I hasn't been "aligned". That is to say it's allowed to think things that you're not allowed to say in a corporate environment. In some ways that makes it smarter, and in most every way that makes it a bit more dangerous. Tools are like that though. Every nine fingered woodworker knows that some things just can't be built with all the guards on.

It is, they trained on chatgpt output. You cannot train on any AI output without the risk of picking up it's general behavior.

Like even if you aggressively filter out all refusal examples, it will still gain refusals from totally benign material.

Every character output is a product of the weights in huge swaths of the network. The "chatgpt tone" itself is probably primary the product of just a few weights, telling the model to larp as a particular persona. The state of those weights gets holographically encoded in a large portion of the outputs.

Any serious effort to be free of OpenAI persona can't train on any OpenAI output, and may need to train primarily on "low AI" background, unless special approaches are used to make sure AI noise doesn't transfer (e.g. using an entirely different architecture may work).

Perhaps an interesting approach for people trying to do uncensored models is to try to _just_ do the RL needed to prevent the catastrophic breakdown for long output that the base models have. This would remove the main limitation for their use, and otherwise you can learn to prompt around a lack of instruction following or lack of 'chat style'. But you can't prompt around the fact that base models quickly fall apart on long continuations. Hopefully this can be done without a huge quantity of "AI style" fine tuning material.

Re: Hermes 4

#30
post #20

I thought for sure this company was going to be based in Paris or Brussels. Maybe Quebec. Nope. NYC.

Were you thinking that "Nous" was French? It's the Greek word for the rational mind (as opposed to the animal appetites or the fighting spirit). Hermes is the Greek god of secret knowledge as well.
Post reply on HN