I put this paper into 4o so i can check if it is relevant, so that you do not have to do this too here are the bullet points: - Vision Transformers can be parallelized to reduce latency and improve optimization without sacrificing accuracy. - Fine-tuning only the attention layers is often sufficient for adapting ViTs to new tasks or resolutions, saving compute and memory. - Using MLP-based patch preprocessing improve…
Three things everyone should know about Vision Transformers
11–19 of 19 posts
Re: Three things everyone should know about Vision Transformers
#12There's something that tickles me about this paper's title. The thought that everyone should know these three things. The idea of going to my neighbor who's a retired K-12 teacher and telling her about how adding MLP-based patch pre-processing layers improves Bert-like self-supervised training based on patch masking.
Clickbait titles are something of a tradition in this field by now. Some important paper titles include "One weird trick for parallelizing convolutional neural networks", "Attention is all you need", and "A picture is worth 16x16 words". Personally I still find it kind of irritating, but to each their own I guess.
For modest incremental improvements, I greatly prefer boring technical titles. Not everything needs to a stochastic parrot. We see this dynamic with building luxury condos. On any individual project, making that pick will help juice profit. When the whole city follows that , it leads to a less desirable outcome.
Re: Three things everyone should know about Vision Transformers
#13I put this paper into 4o so i can check if it is relevant, so that you do not have to do this too here are the bullet points: - Vision Transformers can be parallelized to reduce latency and improve optimization without sacrificing accuracy. - Fine-tuning only the attention layers is often sufficient for adapting ViTs to new tasks or resolutions, saving compute and memory. - Using MLP-based patch preprocessing improve…
just read the abstract
Re: Three things everyone should know about Vision Transformers
#14There's something that tickles me about this paper's title. The thought that everyone should know these three things. The idea of going to my neighbor who's a retired K-12 teacher and telling her about how adding MLP-based patch pre-processing layers improves Bert-like self-supervised training based on patch masking.
Clickbait titles are something of a tradition in this field by now. Some important paper titles include "One weird trick for parallelizing convolutional neural networks", "Attention is all you need", and "A picture is worth 16x16 words". Personally I still find it kind of irritating, but to each their own I guess.
Re: Three things everyone should know about Vision Transformers
#15Earlier quoted context omitted.
just read the abstract
You would think. I don't know about this paper in particular, but I'm continually surprised about how much more I get out of LLM summaries of papers than the abstracts of papers written by the authors.
Re: Three things everyone should know about Vision Transformers
#16Earlier quoted context omitted.
just read the abstract
You would think. I don't know about this paper in particular, but I'm continually surprised about how much more I get out of LLM summaries of papers than the abstracts of papers written by the authors.
If you’ve already decided you’re interested in the paper, then the Introduction and/or Conclusion sections are what you’re looking for.
Re: Three things everyone should know about Vision Transformers
#17Earlier quoted context omitted.
just read the abstract
You would think. I don't know about this paper in particular, but I'm continually surprised about how much more I get out of LLM summaries of papers than the abstracts of papers written by the authors.
Re: Three things everyone should know about Vision Transformers
#18Earlier quoted context omitted.
You would think. I don't know about this paper in particular, but I'm continually surprised about how much more I get out of LLM summaries of papers than the abstracts of papers written by the authors.
Paper abstracts are not optimized by drive-by readers like you and me. They are optimized for active researchers in the field reading their daily arXiv digest that lists all the new papers across the categories they work in, and needing to take the read/don't-read decision for each entry there as efficiently as possible. If you’ve already decided you’re interested in the paper, then the Introduction and/or Conclusion…
Re: Three things everyone should know about Vision Transformers
#19There's something that tickles me about this paper's title. The thought that everyone should know these three things. The idea of going to my neighbor who's a retired K-12 teacher and telling her about how adding MLP-based patch pre-processing layers improves Bert-like self-supervised training based on patch masking.
Clickbait titles are something of a tradition in this field by now. Some important paper titles include "One weird trick for parallelizing convolutional neural networks", "Attention is all you need", and "A picture is worth 16x16 words". Personally I still find it kind of irritating, but to each their own I guess.