Live data from Hacker News

Viewing profile — ssheng

ssheng

HN member
Joined
Fri, Aug 28, 2015, 3:58 AM UTC
HN karma
3
Public activity
9 items

About ssheng

No profile information was provided.

Recent public activity

  1. comment
    Comment #41250940

    To get a full A10, one must also get 36 vCPU and 440 GB of memory. I must be missing something.

  2. comment
  3. story
  4. comment
    Comment #40740836

    Quality loss with quantization is expected. It seems like with GPTQ the loss is within acceptable range based on the perplexity score shown.

  5. comment
    Comment #40740252

    How does Exllama rank among these? Heard good things about it.

  6. comment
    Comment #40703857

    Creative idea. Any data on how much it takes to load the LoRAs and how much latency it adds to the generation speed?

  7. comment
    Comment #40601960

    Hello! We're the authors of this blog post. Please let us know if there are other models and inference backends you'd like us to benchmark next.

  8. comment
    Comment #32088731

    The user feedback seems really encouraging. https://www.linkedin.com/posts/eric-riddoch_its-early-to-say...

  9. comment
    Comment #31769795

    FastAPI is great building block but can't expect it to work for model serving out of box.