Cutting LLM Batch Inference Time in Half: Dynamic Prefix Bucketing at Scale
1–2 of 2 posts
Re: Cutting LLM Batch Inference Time in Half: Dynamic Prefix Bucketing at Scale
#2Part of the Daft team here! Happy to answer any questions
1–2 of 2 posts