Viewing profile — fspeech
fspeech
HN member- Joined
- Mon, Jan 14, 2013, 7:06 PM UTC
- HN karma
- 2,703
- Public activity
- 1,275 items
- HN profile
- View on Hacker News ↗
About fspeech
No profile information was provided.
Recent public activity
-
comment
Comment #49090189
Linear layers use decays (like IIR filters) that naturally provide relative positions. Full attention layers can then be free to develop concepts that attend to each other regardle…
-
comment
Comment #49053773
That's not true. First there is still a licensing and quota scheme on the US side for the H200s. Secondly China blocked them for use in inferencing. Thirdly Chinese companies don't…
-
comment
Comment #48989136
Their kv cache is smaller so they can use less vram and also keep your prefix cached for longer. https://deepseek.ai/blog/deepseek-v4-compressed-attention
-
comment
Comment #47773041
Well, so far as the governments get VAT from the manufacturers, they are getting a return on their investments. They are more like mutualizing the companies than subsidizing them, …
-
comment
Comment #47435628
Performance is generally limited by the process. Yield not so much. Assuming you can make it at a meaningful level at all, yield is generally a learning process.
-
comment
Comment #47407742
Yield is generally not an issue over time, at least not as big as someone outside of the industry would think, if you get enough chips to work in the first place. For high volume c…
-
comment
Comment #47406197
What do you mean? Yield is a function of the chip size and density as well as the process. Plus it's a commercial secret so your bet can't be adjudicated.
-
comment
Comment #47113237
Yes division is a poor example. It's a poor separation of concerns to try to wrap at this level without usage context. To see the point try to wrap overflows on arithmetic function…
- story
-
comment
Comment #46881973
Try ZhiHu(知乎).
-
comment
Comment #46801095
If you are just interested in a structural description (so-called netlist) the standard is EDIF.
-
comment
Comment #46612658
A lot of Chinese internet commentators are very ignorant of the reality in the US, but the Economist's riposte is weird too. For example how is the Chinese property malaise, which …
-
comment
Comment #46405423
Price of a commodity metal can do whatever they want without causing a big problem. It is just a resource allocation signal. However if you base your currency value on it suddenly …
-
comment
Comment #46383049
Do you know what questions Newton was asking? https://en.wikipedia.org/wiki/Religious_views_of_Isaac_Newto... Being right is often hindsight and luck.
-
comment
Comment #46322793
Your point is right on. And additionally, why would an average Indian refuse the pay package to work in China? The top r&d guy at SMIC is from Taiwan after all. Liang got both Sams…
-
comment
Comment #46249259
Humans are also not rewarded for making pronouncements all the time. Experts actually have a reputation to maintain and are likely more reluctant to give opionions that they are no…
-
comment
Comment #46236211
I don't think people mind having bigger spaces but market is not clearing. In the US you have slums and bombed out building shells in prime urban locations as well. It is fascinati…
-
comment
Comment #46233930
Cheap housing isn't the problem. The problems are people speculating on the appreciation of property, banking system depending on property value as loan collaterals and local gover…
- story
-
comment
Comment #46024446
I don't think the repair could be done. It's not about plugging a hole in space. It's about surviving reentry. They can't guarantee the integrity of the glass. Anyway to your point…
-
comment
Comment #45989373
It has a fixed capacity of how many different things it can pay close attention to. If it fails on a seemingly less important but easy to follow instruction it is an indicator that…
-
comment
Comment #45950980
A human rated vehicle would be much more expensive than one rated for cargo. And there are not many use cases for the vehicle other than rescue missions to the two space stations.
-
comment
Comment #45903645
Whatever the reason California averages $4000 in labor cost per Tesla vs less than $300 in Shanghai. That's quite a difference.
-
comment
Comment #45843378
It uses 75% linear attention layers so it is inherently lower cost. And it is MOE so active parameters are far lower.
-
comment
Comment #45843309
There is also Minimax M2 https://huggingface.co/MiniMaxAI/MiniMax-M2