Live data from Hacker News

Viewing profile — WhitneyLand

WhitneyLand

HN member
Joined
Wed, Mar 13, 2013, 7:53 PM UTC
HN karma
6,495
Public activity
2,143 items

About WhitneyLand

I have steadily endeavored to keep my mind free so as to give up any hypothesis, however much beloved, as soon as the facts are shown to be opposed to it - C. Darwin

Recent public activity

  1. comment
    Comment #49214925

    The DeepSeek team is so strong, very impressive. Imagine if they had GPU resources of western labs.

  2. comment
    Comment #49188503

    They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it. They left Opus in and got beat in all but one benchmark. Nothi…

  3. comment
    Comment #49169500

    Nowhere in the paper do they mention the reasoning level or budget used for the experiments? You’ve got to be kidding me. That one variable could make a huge difference in the resu…

  4. comment
    Comment #49169347

    Another headline of “model runs on x”, which usually means “let’s list how much you give up to run on x”. Dumbed down quantization? No. Full intended inference weights preserved, s…

  5. comment
    Comment #49169203

    How do you figure that? When they just loaded the weights alone, it was taking 156GB in vLLM. After warm-up and adding a KV cache pool, it took over 200GB. And this implementation …

  6. comment
    Comment #49128973

    If speed is a metric for you, tokens required to solve a problem affects that metric. All else being equal passing triple the amount of tokens through a model to solve the same pro…

  7. comment
    Comment #49125289

    It’s not outdated at all to use tokens to estimate performance, it’s directly related.

  8. comment
    Comment #49122449

    It’s exciting that a model scoring this high is dirt cheap. It’s also so inefficient, when they release the full performance numbers it’s not going to be good. One example, it take…

  9. comment
    Comment #49085009

    False dichotomy right? Are Chinese labs impressively innovating? Clearly. However this doesn’t rule out possible gains due to distillation. I don’t know the degree of the latter bu…

  10. comment
    Comment #49077837

    The problem is what you’re skeptical about, the true cost, is probably the least important part. Did it really cost $1 million instead of the 150k that’s been floating around? If y…

  11. comment
    Comment #49028040

    Cool worlds says no? https://youtube.com/shorts/qreL7htXp98?is=G9zJVbpMT0yZsLQ-

  12. comment
    Comment #48993884

    The silence is deafening. Google watches over the last few months a flat out assault on the Pareto curve from American and Chinese companies. Release after release pushing the boun…

  13. comment
    Comment #48978014

    The world doesn’t need advisors like Moh that are bitter enough to gossip about a family funeral.

  14. comment
    Comment #48946740

    Not convinced just using top n sigma is going to beat a SOTA detector. The fingerprints Pangram uses should in principle be able to detect style above the token selection level. Ot…

  15. comment
    Comment #48943086

    I’m personally interested in it as part of the research to improve LLM writing. Detecting “AI voice” is part of understanding what’s wrong with it in the first place and how to imp…

  16. comment
    Comment #48943040

    I’m not sure what you’re saying I’m wrong about. The comment I was responding to, about changing LLM output, referred to prompting, not temp/sampling tricks. I’m not aware of Pangr…

  17. comment
    Comment #48940454

    That’s a different point. I’d want detectors to be as accurate as possible, false positives of 1 in 10000 seems like a good starting point. I believe their results have been indepe…

  18. comment
    Comment #48939495

    "It's really easy to have a false positive" Not really. The false positives for the SOTA detector are very very low. "It's also very easy to change the pattern of LLM output." Not …

  19. comment
    Comment #48939447

    "Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it...it's a bad fiction to perpetuate that any of this is anything more t…

  20. comment
    Comment #48907904

    I find it hard to see a language as beautiful that’s grown too complex for a single person to hold a complete mental model of. I used to think that was a personal limitation, until…

  21. comment
    Comment #48900400

    Everyone will still need to use Xcode for at least some debugging, no way around that. As for the builds, your agent probably already knows how to do a lot of this from the command…

  22. comment
    Comment #48866757

    Your value is intrinsic as a human being. We’re capable of love and shared experiences that a machine will never know.

  23. comment
    Comment #48866220

    The prompt does matter. They specifically told it to assume a proof exists so it would not too easily dismiss the possibility.

  24. comment
    Comment #48864305

    If all checks out this is a huge milestone. AI has now solved one of the most famous open problems in graph theory, using an off the shelf model, in one hour. It might be a better …

  25. comment
    Comment #48849682

    Codex is supported well on iPhone/iPad, it’s inside the ChatGPT app. It’s amazing how much work you can get done on your phone now, especially if you already have a design mapped o…