Live data from Hacker News

Viewing profile — anana_

anana_

HN member
Joined
Mon, Dec 05, 2022, 8:37 AM UTC
HN karma
49
Public activity
16 items

About anana_

No profile information was provided.

Recent public activity

  1. comment
    Comment #48970087

    Apparently agentic performance in Gemma was improved recently: https://x.com/googlegemma/status/2077449152062247219 Too little too late imo

  2. comment
    Comment #48724527

    They keep mentioning a 31B dense model, but there are no benchmarks or weights for it anywhere?

  3. comment
    Comment #48662696

    It looks like the purpose of this model is to i. generate environmental sim data for doing RL on other models or ii. act as a foundation model (they trained it to select actions as…

  4. comment
    Comment #48656531

    I believe the benchmark listed is about simulating the environment for the various tasks, rather than doing them. It seems that the point of this model is to generate sim data to i…

  5. comment
    Comment #48545336

    Unfortunately on Strix Halo or any similar unified memory set up, dense models are gonna be dirt slow due to the tiny memory bandwidth... But I agree, 27B is superior.

  6. comment
    Comment #48545298

    Perhaps try a different model? Just from anecdotal experience, I find that the Gemma models smaller than 31B do not tool call as often as they should. Some of the benchmarks appear…

  7. comment
    Comment #48544987

    I have one too and it never occurred to me to use it for anything other than games. Would be interested in seeing how you did it!

  8. comment
    Comment #47969018

    They do now - https://support.mozilla.org/en-US/kb/use-sidebar-access-tool...

  9. comment
    Comment #47838643

    Hypothesizing here, but maybe the idea is sort of a form of technological/economic warfare? Releasing performance equivalent yet more cost efficient open weight models should in th…

  10. comment
    Comment #47750523

    I'm not saying it's the latest Qwen iteration - that would be Qwen3.6. I'm saying it's the latest iteration of the finetuned model mentioned in the parent comment. I'm also not sug…

  11. comment
    Comment #47749755

    It's rather surprising that a solo dev can squeeze more performance out of a model with rather humble resources vs a frontier lab. I'm skeptical of claims that such a fine-tuned mo…

  12. comment
    Comment #47738136

    Upon rereading, I'd agree. Fits with the tone of the rest of the write up.

  13. comment
    Comment #47737613

    > Sometimes you need the absolute cutting-edge reasoning of Claude 3.5 Sonnet or GPT-4o Dead giveaway

  14. comment
    Comment #47256123

    https://huggingface.co/Qwen/Qwen3.5-27B I wasn't aware of that, which page mentions that?

  15. comment
    Comment #47252823

    I've had even better results using the dense 27B model -- less looping and churning on problems

  16. comment
    Comment #47064392

    They own GEICO...