Viewing profile — ssheng
ssheng
HN member- Joined
- Fri, Aug 28, 2015, 3:58 AM UTC
- HN karma
- 3
- Public activity
- 9 items
- HN profile
- View on Hacker News ↗
About ssheng
No profile information was provided.
Recent public activity
-
comment
Comment #41250940
To get a full A10, one must also get 36 vCPU and 440 GB of memory. I must be missing something.
- comment
- story
-
comment
Comment #40740836
Quality loss with quantization is expected. It seems like with GPTQ the loss is within acceptable range based on the perplexity score shown.
-
comment
Comment #40740252
How does Exllama rank among these? Heard good things about it.
-
comment
Comment #40703857
Creative idea. Any data on how much it takes to load the LoRAs and how much latency it adds to the generation speed?
-
comment
Comment #40601960
Hello! We're the authors of this blog post. Please let us know if there are other models and inference backends you'd like us to benchmark next.
-
comment
Comment #32088731
The user feedback seems really encouraging. https://www.linkedin.com/posts/eric-riddoch_its-early-to-say...
-
comment
Comment #31769795
FastAPI is great building block but can't expect it to work for model serving out of box.