Live data from Hacker News

Viewing profile — philipkiely

philipkiely

HN member
Joined
Mon, Aug 20, 2018, 11:10 PM UTC
HN karma
1,120
Public activity
260 items

About philipkiely

DevRel @ https://baseten.co

Email me: username at baseten.co

Recent public activity

  1. story
  2. story
  3. comment
    Comment #48645742

    Good thing we have GLM-5.2

  4. story
  5. comment
    Comment #47537783

    https://github.com/AliesTaha/polar_quant

  6. story
  7. story
    Show HN: Inference Engineering

    There is a ton of demand for inference, but there are relatively few engineers working in the space. This leaves novel, interesting, and deeply technical challenges left to solve a…

  8. story
  9. story
  10. story
  11. story
  12. comment
    Comment #46362683

    GLM 4.6 has been very popular from my perspective as an inference provider with a surprising number of people using it as a daily driver for coding. Excited to see the improvements…

  13. comment
    Comment #46222525

    The Information link, for those with a subscription: https://www.theinformation.com/articles/inference-provider-b...

  14. story
  15. comment
    Comment #45923934

    You give it a text prompt and optional image. What you get is a 3D room based on the prompt/image. It rewrites your prompt to a specific format. Overall the rooms tend to be detail…

  16. comment
    Comment #45923544

    I played with Marble yesterday, Fei-Fei/World Labs' new product. It is the most impressed I've been with an AI experience since the first time I saw a model one-shot material code.…

  17. story
  18. story
  19. story
  20. comment
    Comment #44823974

    We have built a ton of tooling on top of TRT-LLM and use it not just for LLMs but also for TTS models (Orpheus), STT models (Whisper), and embedding models.

  21. comment
    Comment #44823953

    Yeah the custom hardware providers are super good at TPS. Kudos to their teams for sure, and the demos of instant reasoning are incredibly impressive. That said, we are serving the…

  22. comment
    Comment #44823930

    Yeah we have tried to build calculators before it just depends so much. Your equation is roughly correct, but I tend to multiply by a factor of 2 not 1.2 to allow for highly concur…

  23. comment
    Comment #44823907

    TRT-LLM has its challenges from a DX perspective and yeah for Multi-modal we still use vLLM pretty often. But for the kind of traffic we are trying to serve -- high volume and late…

  24. comment
    Comment #44823877

    This comment made my day ty! Yeah definitely speaking from a datacenter perspective -- fastest piece of hardware I have in the parts drawer is probably my old iPhone 8.

  25. comment
    Comment #44823840

    Went to bed with 2 votes, woke up to this. Thank you so much HN!