Live data from Hacker News

In Defense of OpenStreetMap's Data Model

stevecoast.substack.com

71–80 of 130 posts

Re: In Defense of OpenStreetMap's Data Model

#71
post #61

Earlier quoted context omitted.

What exactly you are trying to do? Import data? Map something manually on your own? Something else?

Add keys to existing nodes mostly. Possibly using tasks.openstreetmap.org and/or possibly doing something in a batch if I can get data from the city to use. These structures seem well defined, thankfully. And the crossings and signal locations look to be complete.

In this case I would strongly encourage to start from manual mapping. StreetComplete Android app may be useful here (disclaimer: I am involved in making it).

See also https://wiki.openstreetmap.org/wiki/Import/Guidelines before importing data

Re: In Defense of OpenStreetMap's Data Model

#72
post #31

Earlier quoted context omitted.

The date example is a good one. No one has fun by choosing their own date format. This is putting the burden of choice onto the user. They might like to think about some map stuff and now they have to think about data format stuff. Of course projects like these have to strike a balance between the strictest bureaucratic nightmare and such a structure so loose that people are overburdened by the available options at e…

> The date format in the backend should be fixed and then you should offer flexibility in the frontend for user input. Agreed. However, it might be not so easy for historical dates, because doing it correctly requires great diligence on the part of the tool developer as well as from the user to choose the correct calendar system. For example: Q: What is the correct representation of the date of Caesar's death, 15 Mar…

This is a good point in a general sense, but I don't think it would be a problem in this particular case for the date some imported OSM data was sourced, which is similar to the "date accessed" for a website in a bibliography.

Re: In Defense of OpenStreetMap's Data Model

#73

The proposed improvements would obsolete a bunch of problems such as broken polygons [1] which happen regularly. They would also make processing OSM more accessible without needing to randomly seek over GBs of node locations just to assemble geometries which takes a significant runtime percentage of osm2pgsql. For me Steve Coast lost his credibility when he joined the closed and proprietary what3words. [1] https://wi…

But you don't need to complicate the storage format to fix a problem like that. You can build validation tools that will check whether the stored data conforms to the correct specified geometry, and only emit valid polygons to later tools in the pipeline when they do. "Be liberal in what you accept and strict in what you send" is still a good principle. The problem with rejecting invalid structures at the data storag…

I don't have horribly strong opinions here, but the argument feels circular to me:

- The format should be kept simple to encourage more people to build tools on top of it, and users will be more likely to work with it.

- We should deal with the emergent complexity of bad validation by making tools more complicated and having them detect errors on their end.

If users are going to use a validation tool to work with data, then they can also use a helper tool to generate data. And if the goal is to make it easier to build on top of data, import it, etc... allowing developers to do less work validating everything makes it easier for them to build things.

I'm going over the various threads on this page, and half of the critics here are saying that user data should be user facing, and the other half are saying that separate tools/validators should be used when submitting data. I don't know how to reconcile those two ideas; particularly a few comments that I'm seeing that validation should be primarily clientside embedded in tools.

Again, no strong opinions, and I'll freely admit I'm not familiar enough with OSM's data model to really have an opinion on whether simplification is necessary. But one of the good things about user facing data should be that you can confidently manipulate it without requiring a validator. If you need a validator, then why not also just use a tool to generate/translate the data?

To me, "just use a tool" doesn't seem like a convincing argument for making a data structure more error prone, at least not if the idea is that people should be able to work directly with that data structure.

----

> you'll need to create a new version of the file format and update all tools reading it even if they won't handle the new type, instead of just having old tools silently ignoring the new format that they don't understand.

Again, not sure that I understand the full scope of the problem here, and I'm not trying to make a strong claim, but extensible/backwards-compatible file formats exist. And again, I don't really see how validation solves this problem, you're just as likely to end up with a validator in your pipeline that rejects extensions as invalid, or a renderer that doesn't know how to handle a data extension that used to be invalid or impossible.

Wouldn't be nicer to have a clear definition of what's possible that everyone is aware of and can reason about without inspecting the entire validation stack? Wouldn't it be nice to not finish a big mapping project and then only find out that it has errors when you submit it? Or to know that if your viewer supports vWhatever of the spec that it is guaranteed to actually work, and not fall over when it encounters a novel extension to the data format that it doesn't understand or that it didn't think was possible? Personally, I'd rather be able to know right off the bat what a program supports rather than have to intuit it by seeing how it behaves and looking around for missing data.

Part of what's nice about trying to do extensions explicitly rather than implicitly through assumptions about data shape, is that it's easier to explicitly identify what is and isn't an extension.

Re: In Defense of OpenStreetMap's Data Model

#74

> Let us pray that the EWG is just throwing Jochen a bone to go play in the corner and stop annoying the grownups. That is neither helpful, not useful, nor making me more likely to treat this diatribe more seriously.

It is the kind of disrespectful rhetoric that defines the OSM community though.

I do not consider it as defining and definitely nor desirable or improving ones standing.

For reference: I am extremely active in OSM community. On channels that I moderate this would result in user being warned/kicked (but not banned, unless in case of repetitive insults).

Re: In Defense of OpenStreetMap's Data Model

#75

Earlier quoted context omitted.

The best thing is not to allow invalid geometries to begin with. Any validation would need to be done in an off-line fashion for a number of reasons (such as needing to retrieve any referred OSM elements), and by that time you can't automatically revert offending changes as any revert carries a chance of an object version conflict.

> The best thing is not to allow invalid geometries to begin with. The best thing for whom? The developer? Certainly not for the end user, who needs to have invalid geometries while the drawing is being made and the data is still incomplete . Having a file format that won't admit that temporary state means that either the user can't save incomplete draft work, or that an entirely different format will be needed to re…

>> The best thing is not to allow invalid geometries to begin with.

> The best thing for whom?

Mappers. Noone enjoys untangling broken mulipolygons.

And invalid geometries, by definition, are never desirable or intentional.

I guess that future data consumers (including authors of editors) also would benefit.

Re: In Defense of OpenStreetMap's Data Model

#76
post #64

> The harder you make it for them to edit, the less volunteers you’ll get. And that is why dedicated area type (rather than representing areas with lines or special relations[0]) could help new mappers and new users of data. There would be very significant transition costs, but maybe it would be overall beneficial. It is possible to have objects that are both area and line at once. Or area according to one tool/map/e…

1000 times yes! I am a spatial data expert but only a some-time OSM editor and I still have yet to figure out how to create a polygonal feature more complex than a single building footprint. The theoretical advantage of a unified topology model of just nodes/edges where polygons and lines share core geometry is nullified by cultural rules that say "don't do that" to editors (I had a bunch of parks that shared a bound…

> how to create a polygonal feature more complex than a single building footprint

In ID (default editor) you can mark area and area inside or select two disjointed areas and press right click on the and select "Merge". Or press "c" while selecting areas for combining.

In JOSM there is equivalent "create multipolygon" (or "update multipolygon")

https://wiki.openstreetmap.org/wiki/Relation:multipolygon#Ho...

> parks that shared a boundary with a road

FYI, that is because highway=* road line represents centerline of carriageway - and unless park somehow ends in the middle of road and includes half of its surface it will be not correct.

It also makes future editing quite nasty.

Re: In Defense of OpenStreetMap's Data Model

#77
post #35

Author states: >And that’s the point, rules and complexity have completely unknowable downsides. Downsides like the destruction of the whole project. With each rule and added complexity you make the system less human and less fun. You make it a Computer Scientists rube goldberg machine while sterilizing it of all the joy of life. While too much rules and complexity can certainly be bad, some basic amount of standardi…

why not use the universal date format [1] that works for everyone? yyyy mm dd mm yyyy [1] https://twitter.com/dan_abramov/status/1447710863960551433

you'll have to round up those 86,400ths!

Re: In Defense of OpenStreetMap's Data Model

#78

Earlier quoted context omitted.

Why make a good point and follow it with a bad joke which: both has a class of victim and does not work as a joke? Nobody has ever heard of "the French" undertaking such a project. And if it had ever happened, it would most surely have been a particular Academy or the like. Not the 65M people under your careless swipe.

https://en.wikipedia.org/wiki/Decimal_time#France

Along with most of the metric system, which they did somehow manage to persuade the rest of the world to use.

Re: In Defense of OpenStreetMap's Data Model

#79

The exact same argument that praises OSM's super flexible tagged node data model should also praise MS Excel for the number of things that can be achieved in the world of business with just a grid of boxes. Both have been hugely successful, and both have the same pile of downsides.

> Both have been hugely successful, and both have the same pile of downsides. Exactly. And the solution should not be to throw away spreadsheets completely and turn them into relational databases, but to create new tools to alleviate the downsides and reduce their impact (possible by exporting the spreadsheet information into a relational database, but without taking away the user's option to continue working with it…

> but without taking away the user's option to continue working with it

You keep bringing this up, but I still don't understand what about this change would prevent people from working with intermediary/local formats, or why tools working with intermediary/local formats would be harder than building validators at every step of the submission process?

How is this change taking away anybody's ability to do anything with local data on their device? And if the point is that they should be able to submit that data, then validators will be just as much of a problem for them as a file format will be.

Lots of programs work with their own temp formats locally that are specific to their needs. I mean, you don't need to even make a new one, if you like the existing format so much, save temp changes to it, and publish finished changes to the new format.

What am I missing here, why is any of this a problem?

Re: In Defense of OpenStreetMap's Data Model

#80
post #44

Heh, their data model is 99% of the reason why I don't use OSM. It's scattered all over the place with so many tables ! It's such a nice project, but damn is it impossible to work with programmatically, let alone poke around it to discover what's all in there.

> so many tables ? There are various complaints about OSM data model but this is a new one to me. In OSM basically everything is mixed together and there is no real separation into layers. What you mean by "many tables"?

He probably means what the osm2pgsql import tool creates.
Post reply on HN