Three types of LLM workloads and how to serve them
1–6 of 6 posts
Re: Three types of LLM workloads and how to serve them
#2OCD-driven fix: The correct Latin quote is "Gallia est omnis divisa in partes tres".
Re: Three types of LLM workloads and how to serve them
#3lord
Re: Three types of LLM workloads and how to serve them
#4> Gallia est omnis divisor in partes tres. OCD-driven fix: The correct Latin quote is "Gallia est omnis divisa in partes tres".
Re: Three types of LLM workloads and how to serve them
#5> we recommend using SGLang with excess tensor parallelism and EAGLE-3 speculative decoding on live edge Hopper/Blackwell GPUs accessed via low-overhead, prefix-aware HTTP proxies lord
The technical terms there are later explained and diagrammed, and the recommendations derived from something close to first principles (e.g. roofline analysis).
Re: Three types of LLM workloads and how to serve them
#6Do you have benchmarks for the SGLang vs vLLM latency and throughput question? Not to challenge your point, but I’d like to reproduce these results and fiddle with the configs a bit, also on different models & hardware combos.
(happy modal user btw)