Zero-Mem: Zero-Token Memory Operations for LLM Agents
11–17 of 17 posts
Re: Zero-Mem: Zero-Token Memory Operations for LLM Agents
#12I am working on the same thing right now. However, unlike storing conversations in an external retrieval system, I use a local LLM to store the conversation's KV cache and perform retrieval directly on that cache. The method involves running a prefill pass and, after obtaining the attention scores, filtering for the corpus segments that received attention. This aligns with the "zero tokens" approach described in this…
I was doing something similar where I saved user input/model output in a multi-depth node style storage system (each depth having more precise details) with the focus on the model having accurate user fed information. I was mostly focused on retrieval of accurate / useful information based on user query (injecting the database node as additional, high confidence information) Once this(Zero-mem) passes it's peer revie…
Re: Zero-Mem: Zero-Token Memory Operations for LLM Agents
#13Re: Zero-Mem: Zero-Token Memory Operations for LLM Agents
#14This is actually quite easy to implement at the harness level and the NER can be way more naive because of the typical nature of LLM dialogue (programming, long running tasks etc).
Re: Zero-Mem: Zero-Token Memory Operations for LLM Agents
#15Re: Zero-Mem: Zero-Token Memory Operations for LLM Agents
#16I am working on the same thing right now. However, unlike storing conversations in an external retrieval system, I use a local LLM to store the conversation's KV cache and perform retrieval directly on that cache. The method involves running a prefill pass and, after obtaining the attention scores, filtering for the corpus segments that received attention. This aligns with the "zero tokens" approach described in this…
Some custom kernels and I was able to find all the relevant paragraphs with full force of qwen reasoning within 0.3s, and with a summary round within 0.7s.
Downside - required 200GB ram/vram ;) A few GBs for model and most of it for caching kvs.