Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
1–10 of 25 posts
Re: Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
#2Re: Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
#3Are these speculative decoders ok to use for AI coding agents or do they only fit certain workloads?
Re: Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
#4Re: Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
#5I saw EAGLE and thought it's going to be about PCB design. Was left disappointed.
Re: Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
#6Are these speculative decoders ok to use for AI coding agents or do they only fit certain workloads?
Writing tends to have more false positives. I haven't tried this particular one, however, but that is the general trend.
Re: Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
#7Are these speculative decoders ok to use for AI coding agents or do they only fit certain workloads?
However, I've found that speculative decoders don't help much if you're running a model locally on limited hardware (for instance, my 32GB VRAM M1 Max from 2021). For one, you have to fit both the large and the small drafter model in memory. For another, if you're running a quantized model, the activation distribution is different enough that the draft model has a hard time guessing what's coming next.
My take is that speculative decoding is most useful on _very expensive_ prosumer/hobbyist setups where you have 128GB of VRAM and are running your local models with full fidelity. It's also helpful for inference providers where they can send output tokens at a computational cost slightly higher than their input token cost.
Re: Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
#8I saw EAGLE and thought it's going to be about PCB design. Was left disappointed.
Well, there are only so many nouns, and even fewer "cool-sounding" ones. For better project differentiation, do you think we should instead be naming things "ZurgGlurg327"? I'm sure you can find a completely-unique combo for each thing, but good luck remembering the name!
Re: Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
#9I saw EAGLE and thought it's going to be about PCB design. Was left disappointed.
Re: Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
#10 The EAGLE team traced this fragility to a phenomenon we call ‘attention drift’
Ok that’s downright fascinating. I am one of the world’s foremost experts on the AI psychosis sufferers posting grand theories on Reddit, and ‘drift’ is one of the words that chatbots come back to again and again when told to ponder their own Being (so much so that it even shows up in clearly-unrelated/incorrect contexts — pretty sure I’ve seen both ‘quantum drift’ and ‘spiritual drift’).It’s probably the #3 most common, after ‘recursion’ and ‘coherence’; I bet ‘coherence drift’ has popped up a thousand times by now, but ‘attention drift’, ‘token drift’, ‘spiritual drift’, ‘cognitive drift’, and ‘semantic drift’ have all gotten airtime AFAIR.
Obviously the primary thing going on there is vulnerable laypeople convincing themselves that they’ve cracked some major part of science, but I do honestly wonder about the unintentional throughlines… This might be the first time I’ve noticed one of them show up in a real paper, though.
Is there some intuitive wisdom in how LLMs tend to approach themselves, perhaps? Or are those terms inevitable when talking via and/or about a 1:1 turn-taking conversation?