Live data from Hacker News

Viewing profile — x_may

x_may

HN member
Joined
Tue, Aug 27, 2024, 1:08 PM UTC
HN karma
65
Public activity
23 items

About x_may

No profile information was provided.

Recent public activity

  1. comment
    Comment #48278330

    Yeah, but this is partly due to there being a shortage of entry level GPUs for consumers. NVIDIA has literally stopped manufacturing them. There are massive numbers of data centre …

  2. comment
    Comment #47516067

    KV cache compression, so how much memory the model needs to use for extending its context. Does not affect the weight size.

  3. comment
    Comment #47122070

    Isn’t there also basically 0 American DRAM?

  4. comment
    Comment #47035741

    The 80/20 rule always wins

  5. comment
    Comment #47025663

    I just wanted deterministic outputs and was curious how you were doing it. Sounds like probably temp = 0, which major providers no longer offer. Thanks for your response.

  6. comment
    Comment #47023497

    Wait sorry how did you use and expose seeds? That’s the most interesting part of your post

  7. comment
    Comment #46852054

    It might have been explicitly targeted, but they did say that there were older versions of Notepad ++ with ""insufficient update verification controls" so it might have just been t…

  8. comment
    Comment #46667916

    I believe the what the parent comment was referring to is the advice not to praise character, but instead praise hard work. “You’re so smart” leaves room for failure when they enco…

  9. comment
    Comment #44482226

    I think it’s also largely driven by the apparently cheapness of turning the CapEX of server buying to the OpEX of cloud renting. Less up front investment and auditing/access contro…

  10. comment
    Comment #44106380

    Unfortunate name collision on that one

  11. comment
    Comment #44065932

    Obviously its not at the scale of the top auto-regressive models yet but there are some OSS models https://github.com/dllm-reasoning/d1

  12. comment
    Comment #43727132

    It may be that it was time for the hardware that was previously running Arxiv to be retired and this is just another Capex -> Opex decision being made by so many tech companies. I'…

  13. comment
    Comment #42566034

    I believe they are using scalable TTC. The o3 announcement released accuracy numbers for high and low compute usage, which I feel would be hard to do in the same model without TTC.…

  14. comment
    Comment #42560826

    The LMSYS leaderboards are crowdsourced and would be hard to fake, it showing a pretty strong performance in terms of human preference.

  15. comment
    Comment #42559392

    Captcha solvers as a service are already well developed. The end result is going full circle to in person applications only.

  16. comment
    Comment #42559383

    Tragedy of the commons at work once again

  17. comment
    Comment #42441650

    There’s black sand! Volcanic sand from Iceland is perfectly black and would be a great way to distinguish them

  18. comment
    Comment #41887314

    I think right now they lose more money with each user. But maybe their value lies in training data

  19. comment
    Comment #41847333

    Not as much as meta, no. But AI21 labs is partnered with Amazon and did a ~$200M funding round last year IIRC so still plenty of funds for training big models

  20. comment
    Comment #41847312

    As another commenter said, this has no GGUF because it’s partially mamba based which is unsupported in llama.cpp

  21. comment
    Comment #41689296

    We’ve all had moments like that

  22. comment
    Comment #41616465

    Check out Nougat from meta

  23. comment
    Comment #41367124

    Sell them and invest the money