Earlier quoted context omitted.
Use neural indexes to find the code that most closely matches the output. Explainable AI should be able to tell you where the autocompletion results came from, even if it is a weighted set of files.
That's a good idea in theory, but the smarter the agent gets, the less direct the derivation and the harder to explain it (and to check the explanation). We're already a long way from a nearest-neighbor model. Yet the equivalent problem for humans gets addressed by the clean-room approach. This seems unfair.
at some point it should be different enough to stand on its own, right? then we have no problem with copyrights