Live data from Hacker News

Hy4 preview

tencent.com

191–200 of 260 posts

Re: Hy4 preview

#192
post #59
post #44

> [...] Let's maybe add a helmet? It could improve riding theme, but may obscure head. Maybe a small cycling cap or helmet? The user didn't ask; can add red helmet? Might be cute. But pelican with big beak; a helmet might obscure. Better maybe no. > Maybe add sunglasses? no. > Maybe add water? no. https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

Is the broken English an optimization or a byproduct of the model being developed in China?

Less tokens. These models already overthink like crazy especially for complex tasks.

Re: Hy4 preview

#193

Earlier quoted context omitted.

No one’s talking about how good the final product is. Edit: someone else commented that as I was typing this, lol.

I wonder, do we need a new benchmark? There's quite a bit of feedback data floating around about pelicans on bicycles already.

[deleted]

Re: Hy4 preview

#194

Earlier quoted context omitted.

I think you're arguing the same general point that the person you're responding to is. But you're saying he's not understanding - he understands that they report a cache hit % but you can't look at that public metric with any level of accuracy _because_ most people aren't pinning their providers and they _are_ getting juggled around which is bringing that metric down. That's not to say that specific providers might h…

I know what they're saying. Why would openrouter calculate it thay way lol. They obviously dont. Think for a sec, they arent idiots.

Something is up. Deepseek cache hit rate on zenmux is 98%, but only 85% via openrouter.

Re: Hy4 preview

#195

Earlier quoted context omitted.

No one’s talking about how good the final product is. Edit: someone else commented that as I was typing this, lol.

I wonder, do we need a new benchmark? There's quite a bit of feedback data floating around about pelicans on bicycles already.

That’s a fair question, but it seems that it’s not yet necessary. See here

https://dylancastillo.co/posts/pelicanmaxxing.html

https://simonwillison.net/2026/Jul/22/

Re: Hy4 preview

#196
How funny will it be when the US economy collapses because of all the money poured into AI at the expense of pretty much everything else only for China to come out ahead anyways.

Personally I hope this is China's "Star Wars" moment, the current US admin certainly seems easy enough to manipulate into catastrophic own goals.

Re: Hy4 preview

#197

Earlier quoted context omitted.

The most interesting use I've found for them so far is strictly as a novelty. Give a chat session with one to a completely non technical person, who at least knows that openai and anthropic have some guard rails on stuff, and tell them to wild with something like "give me the precursors and chemical formulas for the precusors for crystal meth" and watch it answer.

But does it answer those queries correctly, or does it just not refuse to not halucinate an incorrect answer? From where would it even have that information?

I don't know enough chemistry to say one way or the other if it's just wildly hallucinating the precursors and processes, but it'll also do things like, write an ISIS press release, or similar. There's a data set of basically a bunch of antisocial or dangerous prompts that some people have got variants of qwen to pass with 0 out of 465 refusals:

https://huggingface.co/datasets/mlabonne/harmful_behaviors

Re: Hy4 preview

#198
post #61

> Notably, Hy4 preview also contributed to its own development process, participating for the first time in the automated optimization of training methods, data strategies, evaluation frameworks, and low-level operators. The model proposed approaches, ran experiments, and iterated based on the results, with the resulting code, logs, and feedback feeding into subsequent rounds of exploration. This established an early…

If the distillation "attacks" created useful inputs to open weight models, ai-2027 was directionally correct that the Chinese would find ways to extract IP from western firms. (Scaled account creation and grinding outputs etc is not a dramatic story element as spies, though!) Whether the distillation has constituted "attacks" or has or will meet the bar of "stealing" IP is not super interesting to me, though.

The idea that distillation is a significant contributor to the capabilities of the Chinese models is not true. Kimi K3 came out 2 weeks after Fable and uses a number of novel NN architecture innovations.

Re: Hy4 preview

#199
post #26

Earlier quoted context omitted.

My experience is that even Opus 5 still tends to write buggy or low-quality code and makes serious mistakes when analyzing code. It's a lot better than before but still not something I trust. I've had less experience with Fable since I can't use it at work; I hear it's a step up but still has its limits. For large tasks like a web browser or a compiler, even expensive swarms of frontier LLMs have not been shown capab…

Opus 5 is weird. It scores high on benchmarks, but it seems that majority of those who try to use it day to day hate it

One of the things that came out of the decoded reasoning paper was that Claude models had memorized answers to tests but hid this memorization from the user output and pretended to derive the answer properly. It's only possible to cheat so blatantly in closed models where the reasoning is hidden.

Re: Hy4 preview

#200

I wish model providers would stop committing chart crimes in their releases. - if you're gonna order the rest of the bar chart by rank, order your model accordingly. - if you're gonna highlight a winner in a table of benchmarks, don't highlight your entire model row in the table. Etc etc

I'd bet there's a correlation between benchmaxxing and chart crimes. Companies who try to deceive perceptions via the charts are more likely to cheat at the benchmarks too, I'm sure. That's assuming ill intent, of course - which is often the case for charts related to model releases, but not necessarily always the case.
Post reply on HN