Live data from Hacker News

Hy4 preview

tencent.com

221–230 of 254 posts

Re: Hy4 preview

#221

Earlier quoted context omitted.

Well, it sure as hell isn't capitalism.

Lmao that’s exactly capitalism

Oh.

Administration corrupted up to it's very core? Check.

Nihilism of anyone not part of the proper color, gender, whatever agenda? Check.

Unlawful surveillance? Check.

Sending totally innocent citizens to prison with many of them dying mysteriously? Check.

Killing innocent people in the streets simply because they dare protest peacefully? Check.

Welcome to North Korea!

Oups. Confused.

Welcome to the GREAT US of A! Where True Capitalism is practiced.

Re: Hy4 preview

#222

Earlier quoted context omitted.

It's my understanding that the llm is not literally thinking those words, they are just the conversion of the matrix multiplication results (numbers) into the tokens. So the matrix is "multiplying" concepts and directions to come up with the final answer - which produces a somewhat readable reasoning trace. As far as the llm is concerned the reasoning trace could be random (to us) symbols. In fact, the reasoning trac…

This is different from the final output? I thought all traversals of the chain of language are through these mathematical means. You could tokenize the word "good" to resolve to only represent 'the opposite of "bad"' and to exclude "as opposed to evil", demanding that this second meaning will be reserved only for the new token "double-minus evil". You've therefore forced the words to be more univalent with no overlap…

I'm seeing it more like: take the concept of "good" put it in a scale of -10 (pure evil) to +10 (pure good). These concepts and the inbetweens have been ingrained into the model, the model can multiply its weights in any combination to express any level and any in between, even -2.16541 etc. So, suppose you give it a short story and ask the model to reason about it and how a character displayed good vs evil behaviour towards the story: Internally the model is making calculation that are very nuanced and extremely precise. This calculations are not the reasoning trace. The reasoning trace itself does not influence the calculations. You may read: "John starts bad and slowly becomes good" when inside the calculations are John goes from -5.245 to -4.24 to -5.221 again, and then 2.1. What matters for nuance is the inner calculations across the many matrix layers. What you see is like an independent program that looks at "-5.245 to -4.24 to -5.221 again, and then 2.1." consults the tokenizer and outputs: "John starts bad and slowly becomes good" or even "John first bad, then good". When in reality, inside, the model as been processing something more akin to "John starts the story as a despicable person, with a redemption arc that builds slowly, he can't yet be considered a good person, certainly not the kind of good person you'd leave your dog with, but he's certainly not as bad as before" the whole time. Now, what if when you continue the conversation, what does the model receive as context? It's original nuanced sentiment, or the brute reasoning trace? That I don't know. It might be that when the reasoning trace is converted from tokens back into numbers it loses all nuance, or it might be that the trace (the words you see) are not the only thing that is being saved and is not the only thing being fed back as context

Re: Hy4 preview

#223
post #44

> [...] Let's maybe add a helmet? It could improve riding theme, but may obscure head. Maybe a small cycling cap or helmet? The user didn't ask; can add red helmet? Might be cute. But pelican with big beak; a helmet might obscure. Better maybe no. > Maybe add sunglasses? no. > Maybe add water? no. https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

Halfway through it states “Let's mentally compose SVG.” Is this a common thing? I’ve never seen it before, the “mentally compose”

It's common for models to produce a draft of the SVG part way through their reasoning. Here's Qwen3.8-Flash-Next doing that, for example: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

The open weight models let you see the full reasoning trace. Models from OpenAI, Anthropic, and Gemini tend to obscure or summarize the reasoning traces so you can't see exactly what they're doing. Here's Gemini 3.7 Flash which looks like it's doing something similar: https://tools.simonwillison.net/markdown-svg-renderer.html#u...

One of the step summaries includes this:

> I'm now detailing the pelican's anatomy within the SVG. I've sketched the main body outline, including coordinates for the tail, chest, neck, head, and massive beak with a pouch. I'm focusing on the position of the eyes and considering the positioning of the wings, with the foreground wing on the handlebar for a confident look.

Re: Hy4 preview

#224
post #17

Is anyone here working on a problem for which current generation LLMs are inadequate, but that could possibly be solved by the next release of a first tier LLM? Or is it like bicycles? Unless your problem is named Tadej, you don't need a $13,000 bike.

I think it's more like steamshovels. 6 months ago they were enabling people who had never broken real earth to find out why osha has so many rules about shoring up walls. You could pull off something complex or delicate but it took a procedure with too many steps too much time to get there, good outcomes were pleasant surprises and required careful target selection. Now it's more like paying good money for a professional. They show up, measure twice, cut once, you're walking around your shiny new hole wondering what took the last guy so long.

The frayed edges on what I have slopped together as unreasonably ambitious, ludicrous projects with fucktons of tokens from models 6-9 months ago mostly look like situations where a capable-enough-to-be-dangerous developer tries to muscle through problems that explode in width & depth but keep digging (so, a tier below stopping early to do more design, two below recognizing the need for more planning from the outset). The primitives are there, most major things work well enough, but the remaining functionality and performance is inaccessible. At a cost of multiples of >1/8th of a $200/mo subscription.

Right now I can put $10 into DSv4 Pro/Flash or Qwen 3.8 Max/Flash, hand it a project and all of its unfinished forks in a state I barely remember, tell it that I want the things these forks have been working towards, and 8 hours later it has ie an working, tested, benchmarked multicore car physics simulation fabric with all of the forks evaluated, the gains merged in, the remaining work documented. It only needed a few hundred more lines of code but Codex 5.5 was never going to see that.

9 months ago I was saying developers are not being ambitious enough with these things, that's only more true now. They lend themselves to digging far deeper than they should: 200kloc god files, dozens of forks. Let it happen, you don't need to read it, it's for them, later. The only time you make them clean up is when it has a severe adverse effect on how long builds/lints/tests/benchmarks take. Spend a whole week having it do nothing but dig up published papers in relevant fields with cutting edge techniques and translating them into feature specs. Pick whatever state of the art is and try to crush it, throw everything at it, leave it looping on vague but wildly ambitious goals. When it modularizes and refactors it all down you might 'do a breakthrough', or maybe it happens in an hour over christmas break when you're trying the next one.

Re: Hy4 preview

#225
post #33
post #17

Is anyone here working on a problem for which current generation LLMs are inadequate, but that could possibly be solved by the next release of a first tier LLM? Or is it like bicycles? Unless your problem is named Tadej, you don't need a $13,000 bike.

I asked a current generation LLM to make me $1k a week and it hasn't so far.

If you're smart enough to try this you're worth more than $52,000 a year. Think about how foolish everyone else will feel when they didn't test the new release of Totally Working Golden Goose For Real This Time

Re: Hy4 preview

#226
post #205
post #125

Earlier quoted context omitted.

Yeah, it has been clear for a long time that there is reasoning and mental modeling going on here. The other option is that you do understand those words the same way, and the people making these (now nonsensical) anti-AI claims simply aren’t talking about the same programs/models we are. Their idea of SOTA is when chatgpt.com launched. If you took a point sample pre-Opus, and didn’t write a good prompt, of course yo…

it has been clear for a long time that there is reasoning and mental modeling going on here There is not. No one from these products is even claiming that's the case and they're so desperate to make the next big claim to re-ignite investment they'd be shouting it from every rooftop. It's just breaking out all the reasonable probabilities around what it's been tasked with and structuring them in a way that is designed…

What, as precisely as you can say, is the difference between an illusion of reasoning and reasoning?

(I am not claiming that there is none. But I personally would define "reasoning" in terms of its structure and its results, and it looks to me as if the best LLMs' "illusion of reasoning" has enough similarities in structure and results to much human reasoning that I don't see why we shouldn't also call it reasoning; if your opinion differs then I'm curious about where the disagreements lie. E.g., do we have different beliefs about what sort of thing LLMs' schmeasoning is able to accomplish, or does your notion of "reasoning" specifically require that it be done by humans, or what?)

Re: Hy4 preview

#227

How funny will it be when the US economy collapses because of all the money poured into AI at the expense of pretty much everything else only for China to come out ahead anyways. Personally I hope this is China's "Star Wars" moment, the current US admin certainly seems easy enough to manipulate into catastrophic own goals.

China basically has been running a "do nothing, win anyway" campaign since the Trump era.

More and more people view the US as a bad ally, and China is stepping in to help in many places.

People should get out more, go visit South America, Dominican Republic, etc. They have cheaper AC, cheaper cars, cheaper appliances, all because they can import from China without tariffs. This is not to say they have worse quality, I would argue the opposite, much of what they import is at least as good or better than what you get in the USA.

It will be no surprise to me when we end up forced to use equal or worse American AI for a premium in price, while the rest of the world moves on.

Re: Hy4 preview

#228
post #216
post #74

Earlier quoted context omitted.

It's very likely tencent games those stats, buying their own tokens.

Openrouter tracks what apps are using the model and the top ones for hy4 are all different coding harnesses. I guess it could be fake but seems more likely people are just trying it out. Hy3 was a very strong and underrated model.

The speed at which hy4 usage increased on openrouter, especially considering its not a cheap model, doesn't seem organic to me.

Its already serving as much tokens/day as the incredibly cheap and good GLM 5.3 flash, which had a crazy marketing campaign as ox alpha?

Also those top 5 apps are just 1.58B tokens out of 1.54T tokens from yesterday. Negligible.

Re: Hy4 preview

#229

How funny will it be when the US economy collapses because of all the money poured into AI at the expense of pretty much everything else only for China to come out ahead anyways. Personally I hope this is China's "Star Wars" moment, the current US admin certainly seems easy enough to manipulate into catastrophic own goals.

China basically has been running a "do nothing, win anyway" campaign since the Trump era. More and more people view the US as a bad ally, and China is stepping in to help in many places. People should get out more, go visit South America, Dominican Republic, etc. They have cheaper AC, cheaper cars, cheaper appliances, all because they can import from China without tariffs. This is not to say they have worse quality,…

Speaking as someone living in Asia who visits the US every year or so, you are falling more and more behind and I don’t see a way for you to catch up without a drastic societal reformation.

The EV market alone should scare the average citizen, but it seems like a non issue every time I bring it up.

I think the US economy has painted itself into a corner and this all in bet on AI is its last real play before the house of cards collapses.

Re: Hy4 preview

#230
post #158
post #61

Earlier quoted context omitted.

If the distillation "attacks" created useful inputs to open weight models, ai-2027 was directionally correct that the Chinese would find ways to extract IP from western firms. (Scaled account creation and grinding outputs etc is not a dramatic story element as spies, though!) Whether the distillation has constituted "attacks" or has or will meet the bar of "stealing" IP is not super interesting to me, though.

the chutzpah of calling it an `attack` or `stealing` is super interesting tho.

It's actually turbo boring and predictable. Capitalism has long since standardized on out and out lies to influence public perception and government action.
Post reply on HN