Earlier quoted context omitted.
Tried this in the past, it's too limited... There are too many ways certain locations can be referred to. Take: New York City, NYC, NY, New York, NYCity, so on...
Wikipedia handles “New York City” and “NYC” as intended. “NY” and “New York” are ambiguous to both machines and humans (are you referring to the city or the state?) and if you have a resolution strategy for this then Wikipedia gives you the options to disambiguate. I’ve never seen “NYCity” used by anybody.
Show HN: Open-source text-to-geolocation models
11–20 of 21 posts
Re: Show HN: Open-source text-to-geolocation models
#12Re: Show HN: Open-source text-to-geolocation models
#13Earlier quoted context omitted.
Wikipedia handles “New York City” and “NYC” as intended. “NY” and “New York” are ambiguous to both machines and humans (are you referring to the city or the state?) and if you have a resolution strategy for this then Wikipedia gives you the options to disambiguate. I’ve never seen “NYCity” used by anybody.
If you start processing web articles on the scale of millions you'll be surprised by how creative people can be. Not talking about tweets, just news and blog articles.
Re: Show HN: Open-source text-to-geolocation models
#14Earlier quoted context omitted.
If you start processing web articles on the scale of millions you'll be surprised by how creative people can be. Not talking about tweets, just news and blog articles.
Not surprised, just not relevant. The criteria here is “you can get pretty good results”, not “you must be able to process millions of articles without failure”.
When processing text at large scale, the usefulness of heuristic approaches like the one we're discussing diminishes rapidly.
Re: Show HN: Open-source text-to-geolocation models
#15Earlier quoted context omitted.
Not surprised, just not relevant. The criteria here is “you can get pretty good results”, not “you must be able to process millions of articles without failure”.
If a method is not generalizable to the entire dataset, it's not that useful. When processing text at large scale, the usefulness of heuristic approaches like the one we're discussing diminishes rapidly.
No, in many situations, something doesn’t have to be perfect to be useful.
Again, I think you are missing the original point being made:
> Depending upon your use-case, you can get pretty good results by…
You seem to be responding as if I said:
> For all use-cases, you can get flawless results by…
Pointing out that this is not perfect is irrelevant to the point I was making. “Good enough” is usually good enough.
Re: Show HN: Open-source text-to-geolocation models
#16There are no weights and no data, only some code to create a pytorch character based network and train it. Will you provide weights or data in the future? Do you have any benchmark over Nominatim or Google maps? I think something like this (but with more substance) could be helpful for some people, especially in the social sciences.
Yea, I was expecting a general-purpose model or dataset to train a model. The idea is great, but - as it currently stands - of no use to most people.
Re: Show HN: Open-source text-to-geolocation models
#17This is _really_ cool. Early in the pandemic I released a local news aggregation tool that aimed to aggregate COVID-related content and score it for relevance using an ensemble of ML classification models, including one that would attempt to infer an article's geographic coordinates. Accuracy peaked at about ~70-80%, which was just not quite high enough for this use case. With a large enough dataset of geotagged docu…
Re: Show HN: Open-source text-to-geolocation models
#18This does look interesting but as other comments have pointed out without data or weights it's not clear how well this works. The training notebook seems to suggest it is not actually improving all that much on the training data
We are working on adding more data as well - feel free to create a GitHub issue if there's more you need - we're going to be working on everything there is to do to help the developers here:)
Re: Show HN: Open-source text-to-geolocation models
#19Depending upon your use-case, you can get pretty good results by using spaCy for named entity recognition then matching on the titles of Wikipedia articles that have coördinates.
Re: Show HN: Open-source text-to-geolocation models
#20Has anyone got this working? Curious if someone could PR a dependencies file that can be used to run this?