Live data from Hacker News

Hy4 preview

tencent.com

181–190 of 259 posts

Re: Hy4 preview

#181

Earlier quoted context omitted.

The results are boring. Not because the content is boring, but because you can so easily remix the results. Human curation is what creates value with these, not dumping and consuming. A personal perspective of a human being ups the respect, where the exact same sentences generated by an LLM carry no such value.

Tolkien did the human creative work - I just want a movie adaptation that’s as honest to the original text as possible. Think translating the text into video.

They are different mediums. While the underlying story might be Tolkien's, every frame is an artistic choice and while LLMs can make a choice is many situations, they are unable to 1) keep it coherent b) make it meaningful because art, to me and most, in an outcome of human experiences and thought, which by definition an llm cannot do.

Re: Hy4 preview

#182

Earlier quoted context omitted.

I can't help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it's subtle. Is there a chance the agent does less reasoning because of it?

Just so we are clear, no "caveman" spoke English. "Caveman speak" is just shortening the vocabulary of english, not a "caveman language". Given this, your concerns for "stereotypical caveman manner" makes very little sense since what caveman are you talking about?

The concern is not that the model was trained on actual caveman artifacts, rather on modern media representations of the stereotypical caveman (that never actually existed).

Re: Hy4 preview

#183

I wish model providers would stop committing chart crimes in their releases. - if you're gonna order the rest of the bar chart by rank, order your model accordingly. - if you're gonna highlight a winner in a table of benchmarks, don't highlight your entire model row in the table. Etc etc

I wonder if this is being reinforced via LLM because they see every other modeler doing the same thing.

Re: Hy4 preview

#184

Earlier quoted context omitted.

You can open source dataset without all the details how it was assembled. Models are lossy compressed datasets you can pick up and amend (fine tune / continue training / alter) according to license they were released under. Hy4 is released under OSI approved Apache License 2.0.

Parent poster is technically right - open “source” implies the source used to make something is open. The model source is training data and code, not just weights. But the reality is, the weights are a useful artifact that you can use to create derivative works. So, dismissing it as a photoshop binary is as technically wrong as calling it open source.

IIRC Nvidia claims to release enough data that it should be possible to fully reproduce Nemotron, so even if it's not as good as the current best models, GPT 5.1 or Opus 4.1.was still useful right? I guess it depends on what you wanted to do with them.

Re: Hy4 preview

#185
post #89

Earlier quoted context omitted.

Optimization. Why use many word when few word do trick?

What I find funny about "why use many word when few word do trick?" is that it's only slightly shorter than the regular "why use many words when few words do the trick?"

The latter is not a complete alternative, it is ambiguously conflating vocabulary scale with word count, and also, it is not as funny

Re: Hy4 preview

#188
post #26
post #17

Is anyone here working on a problem for which current generation LLMs are inadequate, but that could possibly be solved by the next release of a first tier LLM? Or is it like bicycles? Unless your problem is named Tadej, you don't need a $13,000 bike.

My experience is that even Opus 5 still tends to write buggy or low-quality code and makes serious mistakes when analyzing code. It's a lot better than before but still not something I trust. I've had less experience with Fable since I can't use it at work; I hear it's a step up but still has its limits. For large tasks like a web browser or a compiler, even expensive swarms of frontier LLMs have not been shown capab…

Opus 5 is weird. It scores high on benchmarks, but it seems that majority of those who try to use it day to day hate it

Re: Hy4 preview

#189
post #44

> [...] Let's maybe add a helmet? It could improve riding theme, but may obscure head. Maybe a small cycling cap or helmet? The user didn't ask; can add red helmet? Might be cute. But pelican with big beak; a helmet might obscure. Better maybe no. > Maybe add sunglasses? no. > Maybe add water? no. https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

If someone can look at that reasoning trace and see a stochastic parrot next word prediction machine, we don't understand those words in the same way.

Why can't a next token prediction machine not predict a train of reasoning?

Re: Hy4 preview

#190
post #151
post #44

> [...] Let's maybe add a helmet? It could improve riding theme, but may obscure head. Maybe a small cycling cap or helmet? The user didn't ask; can add red helmet? Might be cute. But pelican with big beak; a helmet might obscure. Better maybe no. > Maybe add sunglasses? no. > Maybe add water? no. https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

Did he seriously automate away one of the best quirks of his blog posts, i.e. evaluating new models with a touch of fun? I read AI slop all day, thanks.

That's just part of the output along the SVG/image.
Post reply on HN