This is interesting but what problem does it solve better than CTRL+F-ing a transcript? It seems like this would be a worse solution for when the precise way someone says something could be important (ex. journalists parsing an interview, students studying their recorded lectures) and that it would be most useful if you were working with a large volume of recorded audio, such as customer service calls. This makes me…
After All Is Said and Indexed – Unlocking Information in Recorded Speech
11–17 of 17 posts
Re: After All Is Said and Indexed – Unlocking Information in Recorded Speech
#12Hadn't heard of the thing they were putting their data into, Marqo, a "tensor search for humans" , https://github.com/marqo-ai/marqo
Re: After All Is Said and Indexed – Unlocking Information in Recorded Speech
#13How does this compare to using Whisper and feeding that into a vector DB and querying with a LLM Pardon the dumb question I only have an elementary understanding
Re: After All Is Said and Indexed – Unlocking Information in Recorded Speech
#14A really interesting blog post I found using LLMs for audio search which I think is a pretty nifty/new idea. I've found it cumbersome using some of the new vector DBs (chroma, faiss, etc) to make end to end systems, but with Marqo it doesn't seem too hard.
> I've found it cumbersome using some of the new vector DBs (chroma, faiss, etc) to make end to end systems What parts are cumbersome?
Chroma, Pinecone, I guess FAISS/HNSWlib/etc only handle vector operations. Really what I'd want, which Marqo does, is handle everything end to end.