Live data from Hacker News

Viewing profile — npn

npn

HN member
Joined
Mon, Jun 07, 2021, 5:25 AM UTC
HN karma
180
Public activity
199 items

About npn

No profile information was provided.

Recent public activity

  1. comment
    Comment #49214346

    weak argument. deepseek v4 flash is open weight, you can easily find other providers with competitive price with Deepseek (except for input caching), some even half as cheap.

  2. comment
    Comment #49198011

    From the private conversation of deepseek CEO and the investors, I'm under impression that they take pride for the low API price though. So it is not only training data. If they re…

  3. comment
    Comment #49192851

    btw, GA usually means it is suitable for production, and available globally. the last time a gemini model was available globally in multiple datacenters was gemini 2.0 era.

  4. comment
    Comment #49186728

    Similar to aistudio vs vertex. Or rather, Alibaba cloud is equivalent to google cloud or aws.

  5. comment
    Comment #49133753

    More like it is a generic idea that has been implemented 1000 times before so LLM already have the perfect solution. Well, not like I'm saying the product is doomed though, because…

  6. comment
    Comment #49106571

    weird, for so long I'm damn sure that researchers are obligatory to release research papers to fulfill their quotas. otherwise they will lost their titles. so unless those companie…

  7. comment
    Comment #49088354

    well usually you can just generate the data using LLM. use 2 or 3 different frontier models from different providers, then compare the results and pick the consensus. yes it is not…

  8. comment
    Comment #49088235

    There is a list of a few websites that I subconsciously think only have garbage articles and never really pay serious attention to them: quora, medium and substack. Surely dev.to a…

  9. comment
    Comment #49083013

    fine tuning a small LLMs or even real small language models (like bert) is what the recommended way since the introduction of LLM. the benefit is pretty much obvious: faster to run…

  10. comment
    Comment #49066845

    even deepseek, with their current (dirt cheap) price, can earn enough profit to cover the cost (hardware investment?) in 10 months.

  11. comment
    Comment #49054460

    Huawei chips need advanced 3d packaging in order to keep up. Since the process is too complex the yield is still bad. Not to mention there is a lot of demand from various factors, …

  12. comment
    Comment #49044151

    They spend $40mil for lobbying, I'm sure they can also spare some millions to this place (and other places like reddit). They all do.

  13. comment
    Comment #49020700

    AI means it has npu, Max+ is marking the memory channel, PRO is a normal label for chips that have extra security baked in, it has been this way since forever. I don't see anything…

  14. comment
    Comment #49003290

    update: the knowledge cut off date is "unknown" now. funny because some people downvoted me believed that there is no relation between knowledge cut off date and real world events.…

  15. comment
    Comment #48993607

    tested the models on aistudio. despite that the knowledge cut off is march 2026 it still knows nothing about 2025! you can check by asking "list notable world events in 2025, only …

  16. comment
    Comment #48988252

    what a horrible article. full of misinformation and dishonesty. 1. training new base models are expensive for sure, but fine-tuning them are relatively inexpensive enough the labs …

  17. comment
    Comment #48957822

    Did they cry about it? No right? Don't apply your own standard then judge them about it, petty people. Stackoverflow aimed to be a knowledge base. And knowledge base has a ceiling …

  18. comment
    Comment #48936100

    Not worth it. I have just tried a single prompt in the web interface and it is still not finish reasoning. It thinks too much and often repeats the same stuff over and over. Combin…

  19. comment
    Comment #48782938

    it is funny because nobody ever bother points out that they overcharge you for text input token price. sure it was pretty resource intensity a few years before, but with turbo quan…

  20. comment
    Comment #48477797

    That's for the long term. Anthropic only needs short term solutions for the sake of IPO. They will do whatever they can to sabotage other companies (specially the Chinese ones) to …

  21. comment
    Comment #48449428

    I doubt you can do that. MTP magic happens because for texts, we have a lot of low value fixed tokens that almost always get generated in the sequence (like punctuation, function w…

  22. comment
    Comment #48447050

    How? edit: now I read the article fully, seems like they utilize some very effective MTP algorithm. and somehow the quality is still decent enough. though, I doubt that the quality…

  23. comment
    Comment #48426145

    On the other hand, google does not lose all the money in that deal. Computation is still expensive. So at most they lose like 200M each month. Peanut compares to the potentially ga…

  24. comment
    Comment #48426070

    Is this somehow satire? This is just the dgx spark with keyboard and monitor in a convenient format. Since it has more stuff, I'm sure that the price mark up will increase too. Up …

  25. comment
    Comment #48376246

    from what I understand, it's because unlike the other models, MAI models haven't yet fine-tuned against the synthetic datasets specifically designed to boost the benchmark scores.