The models aren't actually capable of taking into account everything in their context window with industrial yields.
These are stochastic processes, not in the "stochastic parrot" sense, but in the sense of "you are manufacturing emissions and have some measurable rate of success." Like a condom factory.
When you reduce the amount of information you inject, you both decrease cost and improve yield.
"RAG" is application specific methods of estimating which information to admit to the context window. In other words, we use domain knowledge and labor to reduce computational load.
When to do that is a matter of economy.
The economics of RAG in 2024 differ from 2022, and will differ in 2026.
So the question that matters is, "given my timeframe, and current pricing, do I need RAG to deliver my application?"
The second question is, "what's an acceptable yield, and how do I measure it?"
You can't answer that for 2026, because, frankly, you don't even know what you'll be working on.