Run structured extraction on documents/images locally with Ollama and Pydantic
1–10 of 32 posts
Re: Run structured extraction on documents/images locally with Ollama and Pydantic
#2Re: Run structured extraction on documents/images locally with Ollama and Pydantic
#3Re: Run structured extraction on documents/images locally with Ollama and Pydantic
#4Interesting. We're using a SAAS solution for document extraction right now. I don't know if it's in our interest to build out more but I do like the idea of keeping extraction local.
Re: Run structured extraction on documents/images locally with Ollama and Pydantic
#5Interesting. We're using a SAAS solution for document extraction right now. I don't know if it's in our interest to build out more but I do like the idea of keeping extraction local.
Re: Run structured extraction on documents/images locally with Ollama and Pydantic
#6Interesting. We're using a SAAS solution for document extraction right now. I don't know if it's in our interest to build out more but I do like the idea of keeping extraction local.
Our customers insist we run everything on their docs locally.
Re: Run structured extraction on documents/images locally with Ollama and Pydantic
#7Re: Run structured extraction on documents/images locally with Ollama and Pydantic
#8Re: Run structured extraction on documents/images locally with Ollama and Pydantic
#9Not really this application, but QvQ for visual reasoning is also impressive. https://qwenlm.github.io/blog/qvq-72b-preview/
Meta has used Qwen as the basis for their Apollo research. https://arxiv.org/abs/2412.10360
Re: Run structured extraction on documents/images locally with Ollama and Pydantic
#10I'd really like to play with Qwen2.5-VL at some point, perhaps for reading data-sheets for microchips. Nicely for some applications, it's also very good at reporting position of what it finds, which many ML tools are pretty mediocre at. https://qwenlm.github.io/blog/qwen2.5-vl/ Not really this application, but QvQ for visual reasoning is also impressive. https://qwenlm.github.io/blog/qvq-72b-preview/ Meta has used Qw…
We’ve locally tested with Llama 3.2 11B Vision on Ollama: https://github.com/vlm-run/vlmrun-hub/blob/main/tests/benchm...
FWIW I think Ollama structured outputs API is quite buggy compared to the HF transformers variant.