Viewing profile — ycui1986
ycui1986
HN member- Joined
- Sun, Dec 23, 2012, 12:24 AM UTC
- HN karma
- 190
- Public activity
- 99 items
- HN profile
- View on Hacker News ↗
About ycui1986
No profile information was provided.
Recent public activity
-
comment
Comment #49125909
exactly. we need branch predictor for expert weight.
-
comment
Comment #49100546
There are a lot of SSD streaming engines these days. But few to actually try some hard features. There is one that could really improve the speed. Given almost all major models com…
-
comment
Comment #48811171
i always thought Ryzen AI Halo, together with DGS Spark, has mismatched compute capacity with memory size. Given 128GB VRAM, people would want to run large models, but the GPU comp…
-
comment
Comment #47905589
[flagged]
-
comment
Comment #47885687
So, dual RTX PRO 6000
-
comment
Comment #47885656
I really like the pro version. The pelican is so cute.
-
comment
Comment #47865168
32GB RAM on mac also need to host OS, software, and other stuff. There may not even be 24GB VRAM left for the model.
-
comment
Comment #47858468
just because a bunch of rockets went up without blowing up, does not mean they are profitable. it cost money to shot rocket, and it is very expensive, reusable or not. most launche…
-
comment
Comment #47858437
another 60 billion to save a failed AI endeavor.
-
comment
Comment #47843967
outputting docx files does not have much to do with model capability. it is about whether tool calling has be configured .
-
comment
Comment #47843131
There are also many Chines AI-target GPU/NPU producers. You can get a hold of some boards on taobao.com. They are usable in some way. No, nVidia and AMD are not the only ones benef…
-
comment
Comment #47843066
i give it in real ubuntu, no vm, no docker. so long I don't ask it to organize files, it will behave. it has not screw me so far.
-
comment
Comment #47843060
qwen3.5 and qwen3.6 are both good at tool calling.
-
comment
Comment #47746356
For many LLM load, it seems ROCm is slower than vulkan. What’s the point?
-
comment
Comment #47669059
he won't. if anything, openai is falling behind recently. the trend won't change easily. it is like the old time Netscape.
-
comment
Comment #47645862
only works if the users are evenly distributed around the globe (which is likely more of less the case). if the user concentrates in on century, the token rate will be terrible.
-
comment
Comment #47596886
i hope someone do a 100b 1-bit parameter model. that should fit into most 16GB graphics cards. local AI democratized.
-
comment
Comment #47473118
i am guessing, without any proof, that, when one breaker fails the server lose it all, or loose two GPUs, depending on whether one connected to the cpu side failed.
-
comment
Comment #47473102
9070XT provide roughly same inference performance at double the power, half the cost, as RTX PRO 4500. So this one is optimized for total BOM cost.
-
comment
Comment #47473092
they could had gone with the Max-Q version RTX PRO 6000 and only require 120V circuit. 10% performance hit, but half the power. fundamentally, looks like they are shipping consumer…
-
comment
Comment #47450751
for all past years, I have been told wayland is the future. but the decade long dragged out rolling out did not made much sense to me. neither did I investigate why. until today, I…
-
comment
Comment #47109054
the reality is no where to get the fuel. hydrogen stations are shutting down not building up.
-
comment
Comment #46983130
China tests crewed spacecraft abort and rocket recovery in major lunar milestone
- story
-
comment
Comment #46968647
it is bizarre that a notepad app can have remote code execution. how much unnecessary function did MS add to get to this point?