For those interested in related/alternative approaches, one or more of the following established open-source libraries might appeal to you: - Snorkel (training data curation, weak supervision, heuristic labeling functions, uncertainty sampling, relation extraction): https://github.com/snorkel-team/snorkel - AllenNLP (many pretrained NLP research models for tasks beyond text classification, model training and serving,…
I've tried multiple times (although mostly with DeepDive) and it was pretty complicated to get to do anything outside the demos.
The Spacy link is here BTW: https://spacy.io/