Earlier quoted context omitted.
I have this strange sensation that I can't put into words that somehow we are on the brink of unveiling an entirely new paradigm of AIs or perhaps even of combining AI with classical algorithms in a way to rapidly iterate between each other (and sensor data) that will instantly 10x or 100x current capabilities. Anyone else feel this?
I think part of it is the feeling of false understanding that comes from using llms regularly. They let you operate at a higher conceptual level, and they paper over enough of the actual details that your conceptual model might not actually be correct. I'm a mechanical engineer by training, and have similar vibes with the similarities I see between llm training and metallurgy. I could probably put together a formal c…
Matrix Orthogonalization Improves Memory in Recurrent Models
11–20 of 34 posts
Re: Matrix Orthogonalization Improves Memory in Recurrent Models
#12Earlier quoted context omitted.
no. we're approach a sigmoid. AI is bloated carcass and we're tweaking out the size of the models and speed they'll run on smaller hardware. I think to feel what you're feeling, you've bought into "all we need is more context". I think evolution demonstrates that's not really true.
would you really bet that this is it? there is nothing beyond this? reminds me of the famous anecdote of a 19th century physics professor who said "there is nothing left to be discovered in physics, only minor corrections" then came Einstein...
Re: Matrix Orthogonalization Improves Memory in Recurrent Models
#13https://github.com/adrianjav/pogo — POGO: A Proximal One-step Geometric Orthoptimizer
https://arxiv.org/abs/2602.14656 — An Embarrassingly Simple Way to Optimize Orthogonal Matrices at Scale; Adrián Javaloy, Antonio Vergari
Re: Matrix Orthogonalization Improves Memory in Recurrent Models
#14I can't help but think of orthogonal frequency-division multiplexing and it's use in encoding data on multiple carrier frequencies, and it makes me wonder what other parallels we will discover between digital transmission technology for cross-domain stuff like this.
Linear algebra is used everywhere, orthogonalization, SVD, eigenvalues etc are valuable because the resulting properties are very useful in many places.
Re: Matrix Orthogonalization Improves Memory in Recurrent Models
#15I can't help but think of orthogonal frequency-division multiplexing and it's use in encoding data on multiple carrier frequencies, and it makes me wonder what other parallels we will discover between digital transmission technology for cross-domain stuff like this.
I feel like this is an inverted interpretation? Transmission tech uses those methods because the math shows the desired properties. Linear algebra is used everywhere, orthogonalization, SVD, eigenvalues etc are valuable because the resulting properties are very useful in many places.
Re: Matrix Orthogonalization Improves Memory in Recurrent Models
#16Now I’m wondering what is the eigenspace of an LLM? If I take a set of LLM’s with the same number of parameters, then what are the eigenvectors? Do they have different personalities?
The concept of nonlinear eigenvalues exists, but it is a bit more exotic.
Re: Matrix Orthogonalization Improves Memory in Recurrent Models
#17If it can be made orthogonal, can you go a step further and diagonalize it? The storage and performance improvement from that would be huge.
Re: Matrix Orthogonalization Improves Memory in Recurrent Models
#18Earlier quoted context omitted.
no. we're approach a sigmoid. AI is bloated carcass and we're tweaking out the size of the models and speed they'll run on smaller hardware. I think to feel what you're feeling, you've bought into "all we need is more context". I think evolution demonstrates that's not really true.
would you really bet that this is it? there is nothing beyond this? reminds me of the famous anecdote of a 19th century physics professor who said "there is nothing left to be discovered in physics, only minor corrections" then came Einstein...
I don't need to bet anything. I'm not a sociopath who thinks the AI god needs to be built, appeased, etc. That's the torment nexus.
So, it's pretty easy to see realistically if you are satisified with local models and how they affect what you actually do.
I can see the POV of a software engineer that isn't specialized to any specific topic being replaced by various models.
But again, I see the sigmoid, not the "AGI" or the "this baby has grow very big in 1 year, urely it'll become a giant in 5.
Re: Matrix Orthogonalization Improves Memory in Recurrent Models
#19Earlier quoted context omitted.
would you really bet that this is it? there is nothing beyond this? reminds me of the famous anecdote of a 19th century physics professor who said "there is nothing left to be discovered in physics, only minor corrections" then came Einstein...
That wasn’t just a physics professor that was William Thompson aka Lord Kelvin (the dude the temperature unit is named after and one of the most important mathematical physicists of the 19th century [1]), who also said that heavier than air flight was physically impossible only a couple of weeks before the Wright Brothers (and presumably in spite of having at least once in his lifetime seen a bird). Proof that you ca…
This means we can just jump over to mars, then explore other planets, etc, etc.
We know tons of regimes where there is non-continuous progress. Finding a smart dude with an anecdote does not invalidate the breadth and width of all human experience with non-continuous systems.
Some dude thought all fluid was newtonian, and then we discovered non-newtonian fluid. It does exactly what yuou don't expect. Which basically demos physics is complex but that still doesn't mean progress is fluid.
Re: Matrix Orthogonalization Improves Memory in Recurrent Models
#20Earlier quoted context omitted.
I have this strange sensation that I can't put into words that somehow we are on the brink of unveiling an entirely new paradigm of AIs or perhaps even of combining AI with classical algorithms in a way to rapidly iterate between each other (and sensor data) that will instantly 10x or 100x current capabilities. Anyone else feel this?
no. we're approach a sigmoid. AI is bloated carcass and we're tweaking out the size of the models and speed they'll run on smaller hardware. I think to feel what you're feeling, you've bought into "all we need is more context". I think evolution demonstrates that's not really true.