Viewing profile — philipkiely
philipkiely
HN member- Joined
- Mon, Aug 20, 2018, 11:10 PM UTC
- HN karma
- 1,120
- Public activity
- 260 items
- HN profile
- View on Hacker News ↗
About philipkiely
Email me: username at baseten.co
Recent public activity
- story
- story
-
comment
Comment #48645742
Good thing we have GLM-5.2
- story
-
comment
Comment #47537783
https://github.com/AliesTaha/polar_quant
- story
-
story
Show HN: Inference Engineering
There is a ton of demand for inference, but there are relatively few engineers working in the space. This leaves novel, interesting, and deeply technical challenges left to solve a…
- story
- story
- story
- story
-
comment
Comment #46362683
GLM 4.6 has been very popular from my perspective as an inference provider with a surprising number of people using it as a daily driver for coding. Excited to see the improvements…
-
comment
Comment #46222525
The Information link, for those with a subscription: https://www.theinformation.com/articles/inference-provider-b...
- story
-
comment
Comment #45923934
You give it a text prompt and optional image. What you get is a 3D room based on the prompt/image. It rewrites your prompt to a specific format. Overall the rooms tend to be detail…
-
comment
Comment #45923544
I played with Marble yesterday, Fei-Fei/World Labs' new product. It is the most impressed I've been with an AI experience since the first time I saw a model one-shot material code.…
- story
- story
- story
-
comment
Comment #44823974
We have built a ton of tooling on top of TRT-LLM and use it not just for LLMs but also for TTS models (Orpheus), STT models (Whisper), and embedding models.
-
comment
Comment #44823953
Yeah the custom hardware providers are super good at TPS. Kudos to their teams for sure, and the demos of instant reasoning are incredibly impressive. That said, we are serving the…
-
comment
Comment #44823930
Yeah we have tried to build calculators before it just depends so much. Your equation is roughly correct, but I tend to multiply by a factor of 2 not 1.2 to allow for highly concur…
-
comment
Comment #44823907
TRT-LLM has its challenges from a DX perspective and yeah for Multi-modal we still use vLLM pretty often. But for the kind of traffic we are trying to serve -- high volume and late…
-
comment
Comment #44823877
This comment made my day ty! Yeah definitely speaking from a datacenter perspective -- fastest piece of hardware I have in the parts drawer is probably my old iPhone 8.
-
comment
Comment #44823840
Went to bed with 2 votes, woke up to this. Thank you so much HN!