Viewing profile — npn
npn
HN member- Joined
- Mon, Jun 07, 2021, 5:25 AM UTC
- HN karma
- 180
- Public activity
- 199 items
- HN profile
- View on Hacker News ↗
About npn
No profile information was provided.
Recent public activity
-
comment
Comment #49214346
weak argument. deepseek v4 flash is open weight, you can easily find other providers with competitive price with Deepseek (except for input caching), some even half as cheap.
-
comment
Comment #49198011
From the private conversation of deepseek CEO and the investors, I'm under impression that they take pride for the low API price though. So it is not only training data. If they re…
-
comment
Comment #49192851
btw, GA usually means it is suitable for production, and available globally. the last time a gemini model was available globally in multiple datacenters was gemini 2.0 era.
-
comment
Comment #49186728
Similar to aistudio vs vertex. Or rather, Alibaba cloud is equivalent to google cloud or aws.
-
comment
Comment #49133753
More like it is a generic idea that has been implemented 1000 times before so LLM already have the perfect solution. Well, not like I'm saying the product is doomed though, because…
-
comment
Comment #49106571
weird, for so long I'm damn sure that researchers are obligatory to release research papers to fulfill their quotas. otherwise they will lost their titles. so unless those companie…
-
comment
Comment #49088354
well usually you can just generate the data using LLM. use 2 or 3 different frontier models from different providers, then compare the results and pick the consensus. yes it is not…
-
comment
Comment #49088235
There is a list of a few websites that I subconsciously think only have garbage articles and never really pay serious attention to them: quora, medium and substack. Surely dev.to a…
-
comment
Comment #49083013
fine tuning a small LLMs or even real small language models (like bert) is what the recommended way since the introduction of LLM. the benefit is pretty much obvious: faster to run…
-
comment
Comment #49066845
even deepseek, with their current (dirt cheap) price, can earn enough profit to cover the cost (hardware investment?) in 10 months.
-
comment
Comment #49054460
Huawei chips need advanced 3d packaging in order to keep up. Since the process is too complex the yield is still bad. Not to mention there is a lot of demand from various factors, …
-
comment
Comment #49044151
They spend $40mil for lobbying, I'm sure they can also spare some millions to this place (and other places like reddit). They all do.
-
comment
Comment #49020700
AI means it has npu, Max+ is marking the memory channel, PRO is a normal label for chips that have extra security baked in, it has been this way since forever. I don't see anything…
-
comment
Comment #49003290
update: the knowledge cut off date is "unknown" now. funny because some people downvoted me believed that there is no relation between knowledge cut off date and real world events.…
-
comment
Comment #48993607
tested the models on aistudio. despite that the knowledge cut off is march 2026 it still knows nothing about 2025! you can check by asking "list notable world events in 2025, only …
-
comment
Comment #48988252
what a horrible article. full of misinformation and dishonesty. 1. training new base models are expensive for sure, but fine-tuning them are relatively inexpensive enough the labs …
-
comment
Comment #48957822
Did they cry about it? No right? Don't apply your own standard then judge them about it, petty people. Stackoverflow aimed to be a knowledge base. And knowledge base has a ceiling …
-
comment
Comment #48936100
Not worth it. I have just tried a single prompt in the web interface and it is still not finish reasoning. It thinks too much and often repeats the same stuff over and over. Combin…
-
comment
Comment #48782938
it is funny because nobody ever bother points out that they overcharge you for text input token price. sure it was pretty resource intensity a few years before, but with turbo quan…
-
comment
Comment #48477797
That's for the long term. Anthropic only needs short term solutions for the sake of IPO. They will do whatever they can to sabotage other companies (specially the Chinese ones) to …
-
comment
Comment #48449428
I doubt you can do that. MTP magic happens because for texts, we have a lot of low value fixed tokens that almost always get generated in the sequence (like punctuation, function w…
-
comment
Comment #48447050
How? edit: now I read the article fully, seems like they utilize some very effective MTP algorithm. and somehow the quality is still decent enough. though, I doubt that the quality…
-
comment
Comment #48426145
On the other hand, google does not lose all the money in that deal. Computation is still expensive. So at most they lose like 200M each month. Peanut compares to the potentially ga…
-
comment
Comment #48426070
Is this somehow satire? This is just the dgx spark with keyboard and monitor in a convenient format. Since it has more stuff, I'm sure that the price mark up will increase too. Up …
-
comment
Comment #48376246
from what I understand, it's because unlike the other models, MAI models haven't yet fine-tuned against the synthetic datasets specifically designed to boost the benchmark scores.