Live data from Hacker News

Cutting LLM Batch Inference Time in Half: Dynamic Prefix Bucketing at Scale

daft.ai

1–2 of 2 posts