Viewing profile — jug
jug
HN member- Joined
- Mon, Nov 17, 2014, 10:53 AM UTC
- HN karma
- 4,502
- Public activity
- 1,449 items
- HN profile
- View on Hacker News ↗
About jug
No profile information was provided.
Recent public activity
-
comment
Comment #49139824
If that's marketed well it feels like it should cause a system shock like R1 did. It would also be interesting to see the reaction with code models becoming so good already i.e. co…
-
comment
Comment #49117793
Yes, I've seen this too and how Luna xhigh is so good that Terra doesn't really serve a purpose because beyond that you can continue at Sol medium. This can be the most cost effici…
-
comment
Comment #49117776
Kimi K3 is fairly cheap per token but thinks like a madman with poor self esteem.
-
comment
Comment #49108831
So is this how Opus 5 ran it? https://arcprize.org/results/anthropic-claude-opus-5
-
comment
Comment #49069140
In my opinion, next step is to cut down on reasoning tokens while maintaining intelligence. The Chain of Thought and looping can still be an issue with these Chinese models. They i…
-
comment
Comment #49061829
I agree and this is why I think open models will win in the end. There is just so much to gain on being 10% behind the curve. Especially when the curve is far beyond your needs.
-
comment
Comment #49046409
This will also make it harder to compete because it holds true for everyone, not just you. I am already, this year, seeing vibe coders put out some decent stuff on App Stores but t…
-
comment
Comment #48961942
Yeah I've noted this behavior with best in class open weight models. They said K3 would have token efficiency improvements and I was hoping especially solving the thinking loop iss…
-
comment
Comment #48956697
I think they're less and less advertised as true generalists these days, as they pivot to profits that obviously lie (for the time being) first and foremost in agentic coding. It's…
-
comment
Comment #48847595
Addiction due to the dopamine hits of occasional struggle and then churning out apps that work: https://leaddev.com/ai/ai-coding-is-addictive-engineers-are-...
-
comment
Comment #48847470
This article is also related to exhausting AI through generating pressure and posted here recently: AI coding is addictive. Engineers are paying the price https://leaddev.com/ai/ai…
-
comment
Comment #48839330
You probably have it backwards. It's Grok that is shoving right wing ideology down your throat. Research has shown that without specific guidance to otherwise, LLM's tend to be sli…
-
comment
Comment #48829531
Yet, in a month we'll be fine. We were fine with Anthropic naming models by music. I'm sure celestial bodies will be OK too. Larger = better. It's simple. As for the why? Marketing…
-
comment
Comment #48818110
Free tier of Google Gemini can summarize and let you ask questions about pasted YT links.
-
comment
Comment #48778314
Alternative 1 isn’t all that unlikely given Opus 4.8 couldn’t do this. So it’s a recently possible hack. Not something LLM corps have been blindsided by for years. I also strongly …
-
comment
Comment #48765339
I often feel like we're nowadays mostly pushing AI developments in the ways of finetuning differences. Like how new editions of Claude are tuned for agentic coding which might even…
-
comment
Comment #48389973
”If you don't cannibalize yourself, someone else will." — Steve Jobs
-
comment
Comment #48309724
Looks like an ongoing theme and a very poor benchmark. Not at all the claims I expected.
-
comment
Comment #48297728
It's also very surprising to me. This whole deal where humans instantly started taking AI answers at face value, as sources standing on their own legs, or delegating their own mind…
-
comment
Comment #48256201
While oldest source of it, note that the 86-DOS v0.1-C binaries are even earlier (and v0.34 has also been found) than this v1.00 source and can be downloaded and used in an emulato…
-
comment
Comment #48239791
This is a risk although then this is fortunately a model that isn't tied to Chinese hosting. But indeed something to consider if using straight DeepSeek.com.
-
comment
Comment #48216151
I found this thought provoking and just had to see how the new Gemini 3.5 Flash reasoned about this (I find it fun to go meta on modern AI like this), and I'm happy that I did! Als…
-
comment
Comment #48215242
I think that's what the Omniscience Index is for: https://artificialanalysis.ai/evaluations/omniscience#aa-omn... It rewards correct answers and penalizes hallucinations, and final…
-
comment
Comment #48028765
They have now been released on e.g Hugging Face with model suffixes "-assistant".
-
comment
Comment #47902090
Shouldn't one use e.g a Wolfram Alpha MCP endpoint for math in AI? From what I've seen on even premium non-quantized models, I would never ever trust the innate ability of a LLM to…