Live data from Hacker News

Hy4 preview

tencent.com

201–210 of 266 posts

Re: Hy4 preview

#201
post #77

Earlier quoted context omitted.

You can open source dataset without all the details how it was assembled. Models are lossy compressed datasets you can pick up and amend (fine tune / continue training / alter) according to license they were released under. Hy4 is released under OSI approved Apache License 2.0.

So windows is open source because the binaries are a lossy compression of the original source?

Before we worry about source code, Microsoft doesn't grant me rights to modify/redistribute/sell copy of Windows I have.

Re: Hy4 preview

#202
post #146

Earlier quoted context omitted.

For me, an average long session results in about 200-300M cached input, 4-800K input, 2-400K output. Mostly the lower bound. Output depends on how much the model thinks. There are two problems here: - cache hit pricing (both Muse Spark 1.2 Contributor and MiMo 2.5 are around the $0.002-3/M mark) - cache persistence time Muse Spark drops the cache in less than 5m. MiMo keeps it around for at least an hour based on my…

even with the 5m cache, Muse Spark Contribs is still best bang for your buck for the intelligence you get. it is basically the old dsv4-flash prices, but even more smart.

I have used all three extensively. DS4 Flash is quite smart. I would rank MS 1.2 below it. MiMo is the dumbest of them all but good enough for basic stuff.

MiMo wins handsomely if you want to think about your code for minutes at a time as you write. I use it to make changes as I think. I know it will screw up some stuff. I then switch to MS/DS4 once every few hours and have it do a code review and fix the broken stuff. So much cheaper than getting MS to do it on its own.

Re: Hy4 preview

#203
post #89

Earlier quoted context omitted.

Optimization. Why use many word when few word do trick?

What I find funny about "why use many word when few word do trick?" is that it's only slightly shorter than the regular "why use many words when few words do the trick?"

less word better(, many words bad)

Re: Hy4 preview

#204
post #26

Earlier quoted context omitted.

My experience is that even Opus 5 still tends to write buggy or low-quality code and makes serious mistakes when analyzing code. It's a lot better than before but still not something I trust. I've had less experience with Fable since I can't use it at work; I hear it's a step up but still has its limits. For large tasks like a web browser or a compiler, even expensive swarms of frontier LLMs have not been shown capab…

Opus 5 is weird. It scores high on benchmarks, but it seems that majority of those who try to use it day to day hate it

If it were a Chinese model everyone would be screaming benchmaxxed.

Seriously something feels really off about Opus 5. I hope they correct it before 4.6 is removed.

Re: Hy4 preview

#205
post #125

Earlier quoted context omitted.

If someone can look at that reasoning trace and see a stochastic parrot next word prediction machine, we don't understand those words in the same way.

Yeah, it has been clear for a long time that there is reasoning and mental modeling going on here. The other option is that you do understand those words the same way, and the people making these (now nonsensical) anti-AI claims simply aren’t talking about the same programs/models we are. Their idea of SOTA is when chatgpt.com launched. If you took a point sample pre-Opus, and didn’t write a good prompt, of course yo…

it has been clear for a long time that there is reasoning and mental modeling going on here

There is not. No one from these products is even claiming that's the case and they're so desperate to make the next big claim to re-ignite investment they'd be shouting it from every rooftop.

It's just breaking out all the reasonable probabilities around what it's been tasked with and structuring them in a way that is designed to actively look human, and then feed it back to itself. Fundamentally that's the easiest way to iterate new features when the underlying architecture of LLMs is largely "fixed" right now. The fact it is output in a way that appears to reason through each is just a technical decision that creates an illusion of reasoning.

Re: Hy4 preview

#206

How funny will it be when the US economy collapses because of all the money poured into AI at the expense of pretty much everything else only for China to come out ahead anyways. Personally I hope this is China's "Star Wars" moment, the current US admin certainly seems easy enough to manipulate into catastrophic own goals.

it's not Star Wars. china doesn't have to start any war, just let us finish (or get finished in) the wars...

Re: Hy4 preview

#207

How funny will it be when the US economy collapses because of all the money poured into AI at the expense of pretty much everything else only for China to come out ahead anyways. Personally I hope this is China's "Star Wars" moment, the current US admin certainly seems easy enough to manipulate into catastrophic own goals.

it's not Star Wars. china doesn't have to start any war, just let us finish (or get finished in) the wars...

I think you misunderstood what I meant by "Star Wars".

https://en.wikipedia.org/wiki/Strategic_Defense_Initiative

Re: Hy4 preview

#208
post #162
post #159

Earlier quoted context omitted.

This is a remarkable coherent and clear reasoning trace. Maybe you should start also comparing reasoning traces when you do your pelican benchmark.

That would be very interesting but only the open models allow you to see the reasoning trace

So, open models will be better on this benchmark, which is deserved

Re: Hy4 preview

#209

Earlier quoted context omitted.

I can't help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it's subtle. Is there a chance the agent does less reasoning because of it?

Caveman invented fire, the wheel, domesticated wild plants and animals, organised society, survived the Toba catastrophe, cooked food, and was having sex ages before you and me. Don't write him off as stupid.

[dead]

Re: Hy4 preview

#210

Earlier quoted context omitted.

The results are boring. Not because the content is boring, but because you can so easily remix the results. Human curation is what creates value with these, not dumping and consuming. A personal perspective of a human being ups the respect, where the exact same sentences generated by an LLM carry no such value.

Tolkien did the human creative work - I just want a movie adaptation that’s as honest to the original text as possible. Think translating the text into video.

Do you really want a 20 minute pause in the action, every time a new room or space is encountered, because Tolkien goes on for 2-5 pages describing every new environment like that.
Post reply on HN