Running LLMs with 3.3M Context Tokens on a Single GPU
1–4 of 4 posts
Re: Running LLMs with 3.3M Context Tokens on a Single GPU
#2[deleted]
Re: Running LLMs with 3.3M Context Tokens on a Single GPU
#3[deleted]
Re: Running LLMs with 3.3M Context Tokens on a Single GPU
#4Their demo looks really cool: https://github.com/mit-han-lab/duo-attention