Live data from Hacker News

Hy4 preview

tencent.com

111–120 of 252 posts

Re: Hy4 preview

#111

Earlier quoted context omitted.

I can't help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it's subtle. Is there a chance the agent does less reasoning because of it?

some people made a 'caveman' speak qwen as a joke https://huggingface.co/ProCreations/grug-27b

It's not exactly a joke, it does reduce the amount of tokens. However, it does not improve performance (fine tunes are finnecky things, hard to get one right).

Re: Hy4 preview

#113

Earlier quoted context omitted.

some people made a 'caveman' speak qwen as a joke https://huggingface.co/ProCreations/grug-27b

It's not exactly a joke, it does reduce the amount of tokens. However, it does not improve performance (fine tunes are finnecky things, hard to get one right).

Personally the only 'enthusiast' modified qwen 3.6 27b or 3.6 35b-a3b I've found useful are the ones that have been run through heretic and adversarial data sets for innocent/dangerous prompts, to produce uncensored LLMs. They have some niche non-coding uses for things that a commercial LLM will never talk about.

https://github.com/p-e-w/heretic

Re: Hy4 preview

#114

> Notably, Hy4 preview also contributed to its own development process, participating for the first time in the automated optimization of training methods, data strategies, evaluation frameworks, and low-level operators. The model proposed approaches, ran experiments, and iterated based on the results, with the resulting code, logs, and feedback feeding into subsequent rounds of exploration. This established an early…

Don’t need a “better” hacker if you have ten thousand AIs all trying literally every single possible thing to exploit a system with. The main issue is that this will eventually bring down the exploitation cost enough to target very minor targets who weren’t worth it before.

Re: Hy4 preview

#115

Earlier quoted context omitted.

When GPT-5.6-sol's reasoning traces were leaked, they also used "caveman speak". Definitely a token efficiency optimization

I can't help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it's subtle. Is there a chance the agent does less reasoning because of it?

Training a variant to reason in early modern english in the style of the tudor elites might be an amusing way to test for that.

Re: Hy4 preview

#116
post #59

Earlier quoted context omitted.

Is the broken English an optimization or a byproduct of the model being developed in China?

Optimization. Why use many word when few word do trick?

Optimization on a idiosyncrasy. The same thing that makes Claude repeat "That was the most important thing you said in this whole conversation" is what makes grug speak optimize on token usage.

Real humans get non-primary information from word variation. It's reasonable to hypothesize that it has a role in thinking things, because it endures. Our languages need to breathe over time, and flourishing might be one of the aspects that allows that breathing space.

Re: Hy4 preview

#117
post #52

Earlier quoted context omitted.

This is not true, there isn't even a way to see a cache hit % model for a specific model, that wouldn't make any sense. You are confusing what I'm saying with cache cost, that has nothing to do with effective cache hit %. I'm talking about when you click on a specific provider for a specific model, you can scroll down on the view and see their cache hit % for that model [0]. These cache Hit % are accurate, I've done…

I think you're arguing the same general point that the person you're responding to is. But you're saying he's not understanding - he understands that they report a cache hit % but you can't look at that public metric with any level of accuracy _because_ most people aren't pinning their providers and they _are_ getting juggled around which is bringing that metric down. That's not to say that specific providers might h…

I know what they're saying. Why would openrouter calculate it thay way lol. They obviously dont. Think for a sec, they arent idiots.

Re: Hy4 preview

#118
post #66
post #46

Earlier quoted context omitted.

Cache hit % on openrouter is not a good metric, it's mainly driven by openrouter's own provider juggling than the providers themselves

OpenRouter randomizes which provider gets your request by default right? I think you have to pass a specific provider in the request to prevent that. (Or set up a preset or something.) This behavior makes it so you don't benefit much from the caching, unless you pin it to a single provider.

> This behavior makes it so you don't benefit much from the caching

I don't believe this is correct? AFAIK once it routes you to a provider for a given conversation that choice is sticky unless you hit technical difficulties. (It's more complicated than that, they recently added named routing strategies that you can append to the model name.)

IMO the relevant metric is cache TTL which isn't typically published AFAIK.

Re: Hy4 preview

#119

Earlier quoted context omitted.

It's not exactly a joke, it does reduce the amount of tokens. However, it does not improve performance (fine tunes are finnecky things, hard to get one right).

Personally the only 'enthusiast' modified qwen 3.6 27b or 3.6 35b-a3b I've found useful are the ones that have been run through heretic and adversarial data sets for innocent/dangerous prompts, to produce uncensored LLMs. They have some niche non-coding uses for things that a commercial LLM will never talk about. https://github.com/p-e-w/heretic

I think those are mostly vapor that runs on the small culture of "models should not be censored" thing. But from my experience, they unlock nothing meaningful.

Fine-tuning is great for really small models on specific applications, but it's not something that can essentially improve a more generic model.

That said, there seems to be a fine line in quantization+finetuning that could recover performance. It's just hard to get a hold of it (I feel it in some models, but it's hard to say yet; lots of small labs working on this RN).

Re: Hy4 preview

#120
post #42

Earlier quoted context omitted.

That sounds like an interesting challenge. Have you seriously considered solving it? Because in about 10 seconds I came up with a process that should work, provided enough compute power. Simply model the traditional film making process by starting with a script, character stories. Design your world, then design the storyboard, and all the scenes. Create a list of all the visual elements that need to be replicated bet…

I'm getting a Poe's Law feeling. I'm genuinely unsure about whether this post is a stone cold parody or not. I think I need to turn off the internet and go to bed. EDIT: Your username doesn't help, either.

[deleted]
Post reply on HN