Search ArXiv Fluidly
searchthearxiv.com
Search ArXiv Fluidly
1–10 of 16 posts
Re: Search ArXiv Fluidly
#2Re: Search ArXiv Fluidly
#3This seems interesting. I always thought of doing something similar but for specific topic in two specific arxiv categories. But for 300k papers like what is in here, Does anyone have an estimation on how much this will cost using OpenAI ada embeddings model?
It depends on whether you do full text search, or abstract only. If you do full text, I’d guess about 1k tokens per page, 10 pages per paper? So that would be 3B tokens, which would cost you 60$ if you use the cheapest embedder.
If you just do abstracts, the costs will be negligible.
Re: Search ArXiv Fluidly
#4Re: Search ArXiv Fluidly
#5This seems interesting. I always thought of doing something similar but for specific topic in two specific arxiv categories. But for 300k papers like what is in here, Does anyone have an estimation on how much this will cost using OpenAI ada embeddings model?
Ada is deprecated: use the text embedding models instead. It depends on whether you do full text search, or abstract only. If you do full text, I’d guess about 1k tokens per page, 10 pages per paper? So that would be 3B tokens, which would cost you 60$ if you use the cheapest embedder. If you just do abstracts, the costs will be negligible.
Plus you could use a mixed system: first you index the abstract of the most relevant 50 papers, then embedd the text of those 50 in order to asses which are truly relevant and/or meaningful.
Re: Search ArXiv Fluidly
#6Re: Search ArXiv Fluidly
#7If this is your project, please talk about how you made it.
Re: Search ArXiv Fluidly
#8Did they calculate embeddings for the entire archive? That must have cost a fortune.
Re: Search ArXiv Fluidly
#9Did they calculate embeddings for the entire archive? That must have cost a fortune.
What are embeddings and why are they expensive?
Re: Search ArXiv Fluidly
#10Earlier quoted context omitted.
What are embeddings and why are they expensive?
Embeddings are vectors of chunks of documents, lists of 1024 (depending on a model) float numbers that represent that short snippet of text. This kind of search works by finding the most similar vectors, calculating them cost fractions of the cent, but when you need to do it billions to trillions of times, it adds up.
Searching the embeddings is a different problem, but there are lots of specialised databases that can make it efficient.