Live data from Hacker News

Viewing profile — jug

jug

HN member
Joined
Mon, Nov 17, 2014, 10:53 AM UTC
HN karma
4,502
Public activity
1,449 items

About jug

No profile information was provided.

Recent public activity

  1. comment
    Comment #49139824

    If that's marketed well it feels like it should cause a system shock like R1 did. It would also be interesting to see the reaction with code models becoming so good already i.e. co…

  2. comment
    Comment #49117793

    Yes, I've seen this too and how Luna xhigh is so good that Terra doesn't really serve a purpose because beyond that you can continue at Sol medium. This can be the most cost effici…

  3. comment
    Comment #49117776

    Kimi K3 is fairly cheap per token but thinks like a madman with poor self esteem.

  4. comment
    Comment #49108831

    So is this how Opus 5 ran it? https://arcprize.org/results/anthropic-claude-opus-5

  5. comment
    Comment #49069140

    In my opinion, next step is to cut down on reasoning tokens while maintaining intelligence. The Chain of Thought and looping can still be an issue with these Chinese models. They i…

  6. comment
    Comment #49061829

    I agree and this is why I think open models will win in the end. There is just so much to gain on being 10% behind the curve. Especially when the curve is far beyond your needs.

  7. comment
    Comment #49046409

    This will also make it harder to compete because it holds true for everyone, not just you. I am already, this year, seeing vibe coders put out some decent stuff on App Stores but t…

  8. comment
    Comment #48961942

    Yeah I've noted this behavior with best in class open weight models. They said K3 would have token efficiency improvements and I was hoping especially solving the thinking loop iss…

  9. comment
    Comment #48956697

    I think they're less and less advertised as true generalists these days, as they pivot to profits that obviously lie (for the time being) first and foremost in agentic coding. It's…

  10. comment
    Comment #48847595

    Addiction due to the dopamine hits of occasional struggle and then churning out apps that work: https://leaddev.com/ai/ai-coding-is-addictive-engineers-are-...

  11. comment
    Comment #48847470

    This article is also related to exhausting AI through generating pressure and posted here recently: AI coding is addictive. Engineers are paying the price https://leaddev.com/ai/ai…

  12. comment
    Comment #48839330

    You probably have it backwards. It's Grok that is shoving right wing ideology down your throat. Research has shown that without specific guidance to otherwise, LLM's tend to be sli…

  13. comment
    Comment #48829531

    Yet, in a month we'll be fine. We were fine with Anthropic naming models by music. I'm sure celestial bodies will be OK too. Larger = better. It's simple. As for the why? Marketing…

  14. comment
    Comment #48818110

    Free tier of Google Gemini can summarize and let you ask questions about pasted YT links.

  15. comment
    Comment #48778314

    Alternative 1 isn’t all that unlikely given Opus 4.8 couldn’t do this. So it’s a recently possible hack. Not something LLM corps have been blindsided by for years. I also strongly …

  16. comment
    Comment #48765339

    I often feel like we're nowadays mostly pushing AI developments in the ways of finetuning differences. Like how new editions of Claude are tuned for agentic coding which might even…

  17. comment
    Comment #48389973

    ”If you don't cannibalize yourself, someone else will." — Steve Jobs

  18. comment
    Comment #48309724

    Looks like an ongoing theme and a very poor benchmark. Not at all the claims I expected.

  19. comment
    Comment #48297728

    It's also very surprising to me. This whole deal where humans instantly started taking AI answers at face value, as sources standing on their own legs, or delegating their own mind…

  20. comment
    Comment #48256201

    While oldest source of it, note that the 86-DOS v0.1-C binaries are even earlier (and v0.34 has also been found) than this v1.00 source and can be downloaded and used in an emulato…

  21. comment
    Comment #48239791

    This is a risk although then this is fortunately a model that isn't tied to Chinese hosting. But indeed something to consider if using straight DeepSeek.com.

  22. comment
    Comment #48216151

    I found this thought provoking and just had to see how the new Gemini 3.5 Flash reasoned about this (I find it fun to go meta on modern AI like this), and I'm happy that I did! Als…

  23. comment
    Comment #48215242

    I think that's what the Omniscience Index is for: https://artificialanalysis.ai/evaluations/omniscience#aa-omn... It rewards correct answers and penalizes hallucinations, and final…

  24. comment
    Comment #48028765

    They have now been released on e.g Hugging Face with model suffixes "-assistant".

  25. comment
    Comment #47902090

    Shouldn't one use e.g a Wolfram Alpha MCP endpoint for math in AI? From what I've seen on even premium non-quantized models, I would never ever trust the innate ability of a LLM to…