Turning a pile of documents into a searchable useable knowledge base
11–20 of 58 posts
Re: Turning a pile of documents into a searchable useable knowledge base
#12I had an issue. A documents folder with over 12k objects in it. A hodgepodge of folders and sub-folders. That over time had created a mess that no amount of file movement was ever going to make it usable. I wanted: 1) To keep my data local 2) be able to filter out PII and other data 3) Be able to find and delete duplicates 4) Get short synopsis of what a document is 5) Semantic and keyword search 6) All of this kept…
Also need to search git repos including all branches and history (TIL/xkcd#153'd GitLab's web search can basically only do one branch at a time).
Re: Turning a pile of documents into a searchable useable knowledge base
#13Earlier quoted context omitted.
Could it be extended so it also extracts pictures from pptx and xlsx and run vision to get a description to be added to the text content before indexing?
Let me look into this
Re: Turning a pile of documents into a searchable useable knowledge base
#14Nice, what are you hoping to accomplish with this project?
Re: Turning a pile of documents into a searchable useable knowledge base
#15Nice, what are you hoping to accomplish with this project?
Re: Turning a pile of documents into a searchable useable knowledge base
#16I had an issue. A documents folder with over 12k objects in it. A hodgepodge of folders and sub-folders. That over time had created a mess that no amount of file movement was ever going to make it usable. I wanted: 1) To keep my data local 2) be able to filter out PII and other data 3) Be able to find and delete duplicates 4) Get short synopsis of what a document is 5) Semantic and keyword search 6) All of this kept…
Sounds similar to https://docs.paperless-ngx.com/ Key difference I see is that you point it to a folder instead of uploading to a system.
It's pretty cool, I've set up a share where the scanner scans, and it automatically picks it up from there and ingests it into the system.
Re: Turning a pile of documents into a searchable useable knowledge base
#17Re: Turning a pile of documents into a searchable useable knowledge base
#18We need projects like this. Automatically classifying the files is smart. I'm working on a similar application called Hister ( https://github.com/asciimoo/hister ). I should borrow some of your ideas. =]
Re: Turning a pile of documents into a searchable useable knowledge base
#19I had an issue. A documents folder with over 12k objects in it. A hodgepodge of folders and sub-folders. That over time had created a mess that no amount of file movement was ever going to make it usable. I wanted: 1) To keep my data local 2) be able to filter out PII and other data 3) Be able to find and delete duplicates 4) Get short synopsis of what a document is 5) Semantic and keyword search 6) All of this kept…
Personal use? I need this at work, dragging useful info from tarpits like Teams and GitLab. Also need to search git repos including all branches and history (TIL/xkcd#153'd GitLab's web search can basically only do one branch at a time).
Re: Turning a pile of documents into a searchable useable knowledge base
#20Nice, what are you hoping to accomplish with this project?
Pretty much in that order