Live data from Hacker News

Viewing profile — djsjajah

djsjajah

HN member
Joined
Sun, Nov 26, 2023, 11:03 PM UTC
HN karma
90
Public activity
37 items

About djsjajah

No profile information was provided.

Recent public activity

  1. comment
    Comment #49191938

    You have to power it all the time, but the amount of power it uses while it’s on will change by up to a few orders of magnitude depending on the gpu. It’s not uncommon for a gpu to…

  2. comment
    Comment #49077905

    but what is the point? A ban is supposed to make a certain thing less likely to occur. Does a ban of open source models do that? Presumably, the behavior you are trying to limit is…

  3. comment
  4. comment
    Comment #48812454

    I think they were making a joke. In the future, you might consider the advice you are giving as well as giving it.

  5. comment
    Comment #48743258

    Yes. Wait a day

  6. comment
    Comment #48625024

    Not with 800 examples. If you are going to consider an ngram model, I think you are better off getting a frontier llm to write you an absurd regex.

  7. comment
    Comment #48582070

    Except, that won’t help. By the time a new fab is up and running, we will probably have a massive surplus.

  8. comment
    Comment #48534458

    You need to think this thought through all the way to the end. What it has said also influences what it will say. If it has consistently made combative responses, then the most lik…

  9. comment
    Comment #48468560

    It’s amusing that a lot of the agents have worked out that sampling doesn’t change ppl.

  10. comment
    Comment #48452026

    I think what they mean by “now” is the stuff announced today.

  11. comment
    Comment #47772549

    I don't follow. Can you explain how your comment is relevant to mine? It might help if you also explain how you interpreted my comment.

  12. comment
    Comment #47761344

    You just failed the Turing test.

  13. comment
    Comment #47750516

    I have 2 of them. I would advise against if you want to run things like vllm. I have had the cards for months and I still have not been able to create a uv env with trl and vllm. F…

  14. comment
    Comment #47748722

    > or by the community Hmmm

  15. comment
    Comment #47525710

    yes, but the difference between one model and one 4x larger is usually a lot more than that. It is not a question of do a run Qwen 8b at bf16 or a quantized version. It more of a q…

  16. comment
    Comment #47473019

    trl. give me a uv command to get that working. But even in the amd stack things (like ck and aiter) consumer cards are not even second class citizens. They are a distance third at …

  17. comment
    Comment #47460910

    No. It seems to me that the comment is objectively incorrect. The original comment was talking about inference and from what I can tell, it is strictly going to run slower than the…

  18. comment
    Comment #46983567

    That’s kind of a moot point. Even if none of those overheads existed you would still be getting a a fractions of the mfu. Models are fundamental limited by memory bandwidth even wi…

  19. comment
    Comment #46697922

    > including all previous experiments How far back do you go? What about experiments into architecture features that didn’t make the cut? What about pre-transformer attention? But m…

  20. comment
    Comment #46672841

    Not only can it be streamed, but lz4 will probably make things quicker.

  21. comment
    Comment #46497316

    You just ruined my day. The post makes it sound like gel is now dead. The post by Vercel does not give me much hope either [1]. Last commit on the gel repo was two weeks ago. [1] h…

  22. comment
    Comment #46369368

    > Do you really though? Yes. It stays in on the hbm but it need to get shuffled to the place where it can actually do the computation. It’s a lot like a normal cpu. The cpu can’t d…

  23. comment
    Comment #46369247

    GPUs might not be bandwidth starved most of the time, but they absolutely are when generating text from an llm. It’s the whole reason why low precision floating point numbers are b…

  24. comment
    Comment #45989197

    I can't tell if you are making a joke or not. They are not even remotely equivalent. tinygrad is a toy. If you are serious, I would be interested to hear how you see tinygrad repla…

  25. comment
    Comment #45964253

    I went to check how many services are being impacted on down detector, but it was down.