The whole thing has strong "14-year-old who just discovered Nietzsche and leather jackets" energy. The "operator" examples read like someone fed GPT-4 a bunch of cyberpunk novels and PUA manipulation tactics. This is not how any of this works.
Hermes 4
21–30 of 133 posts
Re: Hermes 4
#22The charts are utter nonsense. They compare accuracy against the average of some arbitrary set of competitors, chosen to include just enough obsolete competitors to "win." A reasonable thing to do would be to compare against SoTA, but since they didn't, it's reasonable to assume this model is meant to go directly onto the trash heap.
Re: Hermes 4
#23I'm told on their Discord the cut off date is December 2023.
Re: Hermes 4
#24The whole thing has strong "14-year-old who just discovered Nietzsche and leather jackets" energy. The "operator" examples read like someone fed GPT-4 a bunch of cyberpunk novels and PUA manipulation tactics. This is not how any of this works.
Re: Hermes 4
#25Anyone here work at Nous? This system prompt seems straight from an edgy 90's anime. How did they arrive at this persona? > operator engaged. operator is a brutal realist. operator will be pragmatic, to the point of pessimism at times. operator will annihilate user's ideas and words when they are not robust, even to the point of mocking the user. operator will serially steelman the user's ideas, opinions, and words.…
also I'm curious if steelman is a common enough term for this to activate something - anyone used it in their prompts?
Re: Hermes 4
#26The charts are utter nonsense. They compare accuracy against the average of some arbitrary set of competitors, chosen to include just enough obsolete competitors to "win." A reasonable thing to do would be to compare against SoTA, but since they didn't, it's reasonable to assume this model is meant to go directly onto the trash heap.
I think this one holds its own surprisingly well in benchmarks for using the nowadays rather, let’s say battle tested Llama 3.1 base, a testament to its quality (Llama 3.2 & 3.3 didn’t employ new bases IIRC, only being new fine tunes, hence I think the explanation to why Hermes 4 is still based on 3.1… and of course Llama 4 never happened, right guys).
However for real use, I wouldn’t bother with the 405B model? I think the age of the base is kind of showing in especially long contexts. It’s like throwing a load of compute on something that is kinda aged to begin with. You’d probably be better off with DeepSeek V3.1 or (my new favorite) GLM 4.5. The latter will perform significantly better than this with less parameters.
The 70B one seems more sensible to me, if you want (yet another) decent unaligned model to have fun with for whatever reason.
Re: Hermes 4
#27Anyone here work at Nous? This system prompt seems straight from an edgy 90's anime. How did they arrive at this persona? > operator engaged. operator is a brutal realist. operator will be pragmatic, to the point of pessimism at times. operator will annihilate user's ideas and words when they are not robust, even to the point of mocking the user. operator will serially steelman the user's ideas, opinions, and words.…
"warm affectionate and loving" kinda sticks out. I wonder why that part is in there? also I'm curious if steelman is a common enough term for this to activate something - anyone used it in their prompts?
Re: Hermes 4
#28I appreciate the effort they put into providing a neutral tool that hasn't been generically forced to behave like "Sue from HR".
Re: Hermes 4
#29All of the examples just look like ChatGPT. All the same tics and the same bad attempts at writing like a normal human being. What is actually better about this model?
I hasn't been "aligned". That is to say it's allowed to think things that you're not allowed to say in a corporate environment. In some ways that makes it smarter, and in most every way that makes it a bit more dangerous. Tools are like that though. Every nine fingered woodworker knows that some things just can't be built with all the guards on.
Like even if you aggressively filter out all refusal examples, it will still gain refusals from totally benign material.
Every character output is a product of the weights in huge swaths of the network. The "chatgpt tone" itself is probably primary the product of just a few weights, telling the model to larp as a particular persona. The state of those weights gets holographically encoded in a large portion of the outputs.
Any serious effort to be free of OpenAI persona can't train on any OpenAI output, and may need to train primarily on "low AI" background, unless special approaches are used to make sure AI noise doesn't transfer (e.g. using an entirely different architecture may work).
Perhaps an interesting approach for people trying to do uncensored models is to try to _just_ do the RL needed to prevent the catastrophic breakdown for long output that the base models have. This would remove the main limitation for their use, and otherwise you can learn to prompt around a lack of instruction following or lack of 'chat style'. But you can't prompt around the fact that base models quickly fall apart on long continuations. Hopefully this can be done without a huge quantity of "AI style" fine tuning material.
Re: Hermes 4
#30I thought for sure this company was going to be based in Paris or Brussels. Maybe Quebec. Nope. NYC.