Earlier quoted context omitted.
You can open source dataset without all the details how it was assembled. Models are lossy compressed datasets you can pick up and amend (fine tune / continue training / alter) according to license they were released under. Hy4 is released under OSI approved Apache License 2.0.
So windows is open source because the binaries are a lossy compression of the original source?
Hy4 preview
201–210 of 261 posts
Re: Hy4 preview
#202Earlier quoted context omitted.
For me, an average long session results in about 200-300M cached input, 4-800K input, 2-400K output. Mostly the lower bound. Output depends on how much the model thinks. There are two problems here: - cache hit pricing (both Muse Spark 1.2 Contributor and MiMo 2.5 are around the $0.002-3/M mark) - cache persistence time Muse Spark drops the cache in less than 5m. MiMo keeps it around for at least an hour based on my…
even with the 5m cache, Muse Spark Contribs is still best bang for your buck for the intelligence you get. it is basically the old dsv4-flash prices, but even more smart.
MiMo wins handsomely if you want to think about your code for minutes at a time as you write. I use it to make changes as I think. I know it will screw up some stuff. I then switch to MS/DS4 once every few hours and have it do a code review and fix the broken stuff. So much cheaper than getting MS to do it on its own.
Re: Hy4 preview
#203Earlier quoted context omitted.
Optimization. Why use many word when few word do trick?
What I find funny about "why use many word when few word do trick?" is that it's only slightly shorter than the regular "why use many words when few words do the trick?"
Re: Hy4 preview
#204Earlier quoted context omitted.
My experience is that even Opus 5 still tends to write buggy or low-quality code and makes serious mistakes when analyzing code. It's a lot better than before but still not something I trust. I've had less experience with Fable since I can't use it at work; I hear it's a step up but still has its limits. For large tasks like a web browser or a compiler, even expensive swarms of frontier LLMs have not been shown capab…
Opus 5 is weird. It scores high on benchmarks, but it seems that majority of those who try to use it day to day hate it
Seriously something feels really off about Opus 5. I hope they correct it before 4.6 is removed.
Re: Hy4 preview
#205Earlier quoted context omitted.
If someone can look at that reasoning trace and see a stochastic parrot next word prediction machine, we don't understand those words in the same way.
Yeah, it has been clear for a long time that there is reasoning and mental modeling going on here. The other option is that you do understand those words the same way, and the people making these (now nonsensical) anti-AI claims simply aren’t talking about the same programs/models we are. Their idea of SOTA is when chatgpt.com launched. If you took a point sample pre-Opus, and didn’t write a good prompt, of course yo…
There is not. No one from these products is even claiming that's the case and they're so desperate to make the next big claim to re-ignite investment they'd be shouting it from every rooftop.
It's just breaking out all the reasonable probabilities around what it's been tasked with and structuring them in a way that is designed to actively look human, and then feed it back to itself. Fundamentally that's the easiest way to iterate new features when the underlying architecture of LLMs is largely "fixed" right now. The fact it is output in a way that appears to reason through each is just a technical decision that creates an illusion of reasoning.
Re: Hy4 preview
#206How funny will it be when the US economy collapses because of all the money poured into AI at the expense of pretty much everything else only for China to come out ahead anyways. Personally I hope this is China's "Star Wars" moment, the current US admin certainly seems easy enough to manipulate into catastrophic own goals.
Re: Hy4 preview
#207How funny will it be when the US economy collapses because of all the money poured into AI at the expense of pretty much everything else only for China to come out ahead anyways. Personally I hope this is China's "Star Wars" moment, the current US admin certainly seems easy enough to manipulate into catastrophic own goals.
it's not Star Wars. china doesn't have to start any war, just let us finish (or get finished in) the wars...
Re: Hy4 preview
#208Earlier quoted context omitted.
This is a remarkable coherent and clear reasoning trace. Maybe you should start also comparing reasoning traces when you do your pelican benchmark.
That would be very interesting but only the open models allow you to see the reasoning trace
Re: Hy4 preview
#209Earlier quoted context omitted.
I can't help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it's subtle. Is there a chance the agent does less reasoning because of it?
Caveman invented fire, the wheel, domesticated wild plants and animals, organised society, survived the Toba catastrophe, cooked food, and was having sex ages before you and me. Don't write him off as stupid.
Re: Hy4 preview
#210Earlier quoted context omitted.
The results are boring. Not because the content is boring, but because you can so easily remix the results. Human curation is what creates value with these, not dumping and consuming. A personal perspective of a human being ups the respect, where the exact same sentences generated by an LLM carry no such value.
Tolkien did the human creative work - I just want a movie adaptation that’s as honest to the original text as possible. Think translating the text into video.