Live data from Hacker News

Viewing profile — wolttam

wolttam

HN member
Joined
Fri, Nov 04, 2022, 9:21 PM UTC
HN karma
924
Public activity
357 items

About wolttam

https://codeberg.org/mlow/lmcli

Recent public activity

  1. comment
    Comment #49224173

    As written it sounds like you're saying that it was incompetent of the labs to disable the sandbox internet access? They tried to disable open internet access but the models zero-d…

  2. comment
    Comment #49221830

    Automated defence is going to use so many tokens.

  3. comment
    Comment #49219013

    Yours is the only mention of GPT-OSS in the whole thread. I count about 2-3 comments saying the model is meh and many more saying it’s a big step up. I am among those with real lif…

  4. comment
  5. comment
    Comment #49211920

    I think we will eventually reach a point where this is the case, but at the moment it seems like you can throw virtually any non-trivial use-case at a model today and end up being …

  6. comment
    Comment #49210374

    It feels like the rate of stupid we have achieved in the last (checks watch) ~18 months has surely eclipsed anything that may have occurred during the roman empire. By sheer volume…

  7. comment
  8. comment
    Comment #49203171

    It's a terrible moat. You etch the silicon then nobody wants to run it in 6 months because models have advanced that much further.

  9. comment
    Comment #49183645

    We’re already seeing incredible knowledge compression out of the models the GP mentioned like Qwen 35B-A3B, which feels well within the realm of “runs on a phone” in the next handf…

  10. comment
    Comment #49183473

    Essentially, MCP is the harness’ wheelhouse, not the model’s

  11. comment
    Comment #49183380

    MCP is the protocol by which the harness can connect with additional tools. The harness does the work of discovering the tools, the model is then fed a description of what tools it…

  12. comment
    Comment #49181314

    MCP is invisible to the LLM, or should be, under normal use

  13. comment
    Comment #49155897

    The problem is that there is no signal to the folks involved whether you have bothered to understand what you’re posting or are simply offloading that work to others. Same issue as…

  14. comment
    Comment #49129623

    2 sparks currently run this model at 60 t/s single session, up to just over 100 t/s aggregate with concurrency of 4. Going local has as opened up a world of use-cases I never would…

  15. comment
  16. comment
    Comment #49119893

    Hooray! This model makes me very optimistic about the future of local inference. The CyberGym score stands out to me.

  17. comment
    Comment #49119714

    It runs really well on 2 DGX Sparks - 60t/s

  18. comment
  19. comment
    Comment #49118501

    > and used tiktoken to approximate tokens, by dividing character count by four. This is just the type of thing that stands out. You used a tokenizer library to determine the number…

  20. comment
    Comment #49083575

    Not always. An email has a valid SPF when its return path email’s domain permits the sending server’s IP. But that email may have a forged From: header (which causes an SPF mis-ali…

  21. comment
    Comment #49077379

    > smaller model that doesn't have any domain knowledge or facts built into its weights I’m not sure it works this way. Language modelling itself is a “domain”, and if it didn’t hav…

  22. comment
    Comment #49023007

    Well, one thing it will still have is a boat-load of compute.

  23. comment
  24. comment
    Comment #49021271

    Each one of those employees maps to several fold times more spending on compute.

  25. comment
    Comment #49021218

    No way this can go wrong for the U.S. economy! Do everything you can to protect the existing incumbents, pay no mind to the wider consequences. Winning strategy.