AmpliGraph: A TensorFlow-Based Library for Knowledge Graph Embeddings
11–14 of 14 posts
Re: AmpliGraph: A TensorFlow-Based Library for Knowledge Graph Embeddings
#12Can you help me understand, what are possible inputs to ampligraph? I think the main use-case is plugging in an existing knowledge graph, and it filling in the gaps, correct? Can I augment this will really high-quality embeddings for the nodes, that were learned over auxiliary unlabelled text? What are other ways I can augment the data set? Is this useful only when there are many edge-types, or is it also good when t…
> the main use-case is plugging in an existing knowledge graph, and it filling in the gaps Correct. That is known as Link Prediction. There are other machine learning tasks you can do, though: for example you can generate embeddings and then cluster them. Or you can use embeddings to see if distinct entities are indeed the same.
> Can I augment this will really high-quality embeddings for the nodes, that were learned over auxiliary unlabelled text? I know there is a handful of papers in literature that do that, but we have not implemented any of them yet in AmpliGraph. Examples:
* Xie, Ruobing, et al. "Representation Learning of Knowledge Graphs with Entity Descriptions." AAAI 2016. * Xu, Jiacheng, et al. "Knowledge Graph Representation with Jointly Structural and Textual Encoding." arXiv preprint arXiv:1611.08661 (2016). * [Han16] Han, Xu, Zhiyuan Liu, and Maosong Sun. "Joint Representation Learning of Text and Knowledge for Knowledge Graph Completion." arXiv preprint arXiv:1611.04125 (2016).
> What are other ways I can augment the data set? I would try first with a dataset with no literals (no strings, no numbers, no geo coordinates) as these are treated as entities, for now. I suggest generating embeddings first on your current graph, and measuring the predictive power using http://docs.ampligraph.org/en/1.0.1/generated/ampligraph.eva... Merging additional datasets would be another option, to get more data to work on.
> Is this useful only when there are many edge-types, or is it also good when there are very few?
Also when there are a few.
Let us know how you likeit, and if you need assistance, we have a public Slack channel - happy to answer any question! https://join.slack.com/t/ampligraph/shared_invite/enQtNTc2NT...
Re: AmpliGraph: A TensorFlow-Based Library for Knowledge Graph Embeddings
#13I've been doing some work on link prediction in knowledge graphs recently with poor results on real-world data. These methods don't necessarily require a huge amount of data but they are very sensitive to noise and the 'density' of dataset. The benchmark datasets are, in essence, very easy to get good performance on. It's a real shame that metrics for these methods' tolerance of noise and sparsity are not reported be…
As a general rule of thumb, it is important your graph has enough redundancy in it, i.e. the more relations, the better. Also, bear in mind these models do not support multi-modality, i.e. literals such as numbers, strings, geo coordinates, timestamps are simply treated as entities. In most cases it is probably better to filter literals out before generating the embeddings.