Honey, I shrunk the embeddings: Matryoshka vs. PCA
dylancastillo.co
Honey, I shrunk the embeddings: Matryoshka vs. PCA
1–10 of 20 posts
Re: Honey, I shrunk the embeddings: Matryoshka vs. PCA
#2In my experiments, I used lots of embedding models and the results were not nearly as uniform as this curve, just FYI. I didn’t use any of the API-based models though
I also wrote about this exact comparison when using PCA and MRL to quantize static models, see: https://stephantul.github.io/blog/mrl-pca/
Re: Honey, I shrunk the embeddings: Matryoshka vs. PCA
#3Nice! I’ve been working on something similar and found similar results. In my experiments, I used lots of embedding models and the results were not nearly as uniform as this curve, just FYI. I didn’t use any of the API-based models though I also wrote about this exact comparison when using PCA and MRL to quantize static models, see: https://stephantul.github.io/blog/mrl-pca/
I couldn't find much when I first looked into this, which is why I ended up writing the article.
Re: Honey, I shrunk the embeddings: Matryoshka vs. PCA
#4Nice! I’ve been working on something similar and found similar results. In my experiments, I used lots of embedding models and the results were not nearly as uniform as this curve, just FYI. I didn’t use any of the API-based models though I also wrote about this exact comparison when using PCA and MRL to quantize static models, see: https://stephantul.github.io/blog/mrl-pca/
Thank you! Will take a look at your results. I couldn't find much when I first looked into this, which is why I ended up writing the article.
Re: Honey, I shrunk the embeddings: Matryoshka vs. PCA
#5Feels like a “just use logistic regression” moment :)
Re: Honey, I shrunk the embeddings: Matryoshka vs. PCA
#6Nice! I’ve been working on something similar and found similar results. In my experiments, I used lots of embedding models and the results were not nearly as uniform as this curve, just FYI. I didn’t use any of the API-based models though I also wrote about this exact comparison when using PCA and MRL to quantize static models, see: https://stephantul.github.io/blog/mrl-pca/
Thank you! Will take a look at your results. I couldn't find much when I first looked into this, which is why I ended up writing the article.
Re: Honey, I shrunk the embeddings: Matryoshka vs. PCA
#7Earlier quoted context omitted.
Thank you! Will take a look at your results. I couldn't find much when I first looked into this, which is why I ended up writing the article.
Did you find much difference in inference latency or throughput between baseline and PCA?
So I guess the answer is: no
Re: Honey, I shrunk the embeddings: Matryoshka vs. PCA
#8> You can push this further by combining quantization with truncation or PCA. The resulting vectors can be dramatically smaller while still preserving a surprising amount of retrieval quality.
Counterintuitively - quantisation can also be combined with a random rotation step before the quantisation. A random rotation spreads information across more dimensions, allowing more aggressive quantisation without losing accuracy. Ironically - almost the opposite of a PCA.
I do wonder if relevant here though. It relies on the embeddings having "structure", i.e. that principal components point along basis vectors, which may not be the case with text embeddings.
Source: https://research.google/blog/turboquant-redefining-ai-effici...
Re: Honey, I shrunk the embeddings: Matryoshka vs. PCA
#9This is fascinating if it works as well as the experiments make it seem. For example, how does it compare to classic image resize algorithms like Seam Carving or Inpainting: https://en.wikipedia.org/wiki/Seam_carving, https://en.wikipedia.org/wiki/Inpainting
(can they be compared?)
Compression/loss, and the opposite - scaling up and oversmoothing - are fascinating in that any time even the tiniest innovation happens in those areas, all this other technology improves overnight, and a bunch of new technology becomes possible.
Re: Honey, I shrunk the embeddings: Matryoshka vs. PCA
#10That is where something like Matryoshka embeddings has appeal. You trade a little bit of performance for a guarantee of training + validation set coverage.