The universal weight subspace hypothesis
141–146 of 146 posts
Re: The universal weight subspace hypothesis
#142Re: The universal weight subspace hypothesis
#143Earlier quoted context omitted.
But there is a steep gap in the spectrum at 16 on page 7
That's the spectrum of LoRAs, which are LoW RAnk by design.
I also don’t understand what they write under figure 2, since resnet50 has 50 layers, not 31.
Re: The universal weight subspace hypothesis
#144Earlier quoted context omitted.
I don't think its that surprising actually. And I think the paper in general completely oversells the idea. The ResNet results hold from scratch because strict local constraints (e.g., 3x3 convolutions) force the emergence of fundamental signal-processing features (Gabor/Laplacian filters) regardless of the dataset. The architecture itself enforces the subspace. The Transformer/ViT results rely on fine-tunes because…
You’ve explained this in plain and simple language far more directly than the linked study. Score yet another point for the theory that academic papers are deliberately written to be obtuse to laypeople rather than striving for accessibility.
Re: The universal weight subspace hypothesis
#145Re: The universal weight subspace hypothesis
#146Earlier quoted context omitted.
Understand the mind to then exploit it. Why else would they put so much money into something if not to try and get more out of it? Capitalists' morals are driven by their social position. To them this is right becauae its rewarding. To us its an akin abomination we create that destroys us But the problem isnt inherently tech. Its how society is structured around it that allows it to be used against us.
A lot of the basic research including this paper is out of academia rather than business.