Earlier quoted context omitted.
What's an alternative you would recommend?
Anything prefix length encoded, with no schema.
BuffDB is a Rust library to simplify multi-plexing on edge devices
21–30 of 33 posts
Re: BuffDB is a Rust library to simplify multi-plexing on edge devices
#22Why don't you make up your own bullshit words instead of randomly picking stuff from other places? That's not even multiplexing.
Re: BuffDB is a Rust library to simplify multi-plexing on edge devices
#23Earlier quoted context omitted.
Protobuf is a tag-length-value (TLV) encoding. It's bad. TLV is the thing that everyone loves to hate about ASN.1's DER.
> It's bad I'm not sure that's a helpful way to look at things. There are tradeoffs. Can you elaborate more about what aspects of a TLV encoding you find problematic? Is it decoding speed? The need to copy the encoded value into a native value in order to make use of it? Something else?
Re: BuffDB is a Rust library to simplify multi-plexing on edge devices
#24Earlier quoted context omitted.
> It's bad I'm not sure that's a helpful way to look at things. There are tradeoffs. Can you elaborate more about what aspects of a TLV encoding you find problematic? Is it decoding speed? The need to copy the encoded value into a native value in order to make use of it? Something else?
TLV encodings are always redundant and take more space than non-TLV encodings. Therefore they are a pessimization. As well, definite-length TLV encodings require two passes to encode the data (one to compute the length of the data to be encoded, and one to encode it), thus they are a) slower than non-TLV binary encodings, b) not on-line for encoding.
Now, you may disagree that tolerating unknown fields is a features (as many people do), but one must understand the context where protobuf has been designed, namely the situation where it takes time to roll out new versions of binaries that process the data (either in API calls or on stored files) and thus the ability to design a schema evolution with backward and forward compatibility is worth a few more cycles during encoding.
Not all users have that need and hence there exist other formats, but I wouldn't dismiss the protobuf encoding as flatly "wrong" just because you don't have the requirements it has been designed for.
Re: BuffDB is a Rust library to simplify multi-plexing on edge devices
#25Earlier quoted context omitted.
TLV encodings are always redundant and take more space than non-TLV encodings. Therefore they are a pessimization. As well, definite-length TLV encodings require two passes to encode the data (one to compute the length of the data to be encoded, and one to encode it), thus they are a) slower than non-TLV binary encodings, b) not on-line for encoding.
Yes all those are indeed pessimizations during encoding but are features when decoding: the decoder can skip decoding fields and tolerate unknown fields. Now, you may disagree that tolerating unknown fields is a features (as many people do), but one must understand the context where protobuf has been designed, namely the situation where it takes time to roll out new versions of binaries that process the data (either…
See ASN.1's extensibility rules. If you mark a record type (SEQUENCE, in ASN.1 parlance) as extensible, then when you later add fields (members, in ASN.1 parlance) the encoding rules have to make it possible to skip/ignore those fields when decoded by software implementing the pre-extension type. PER/OER will include a length for all the extension fields for this purpose, but it can be one length for all the extension fields in each round of extensions rather than one per (which would only save on the type "tag" in TLV).
> the decoder can skip decoding fields
This is mainly true for on-line decoders that produce `{path, leaf value}` tuples _and_ which take paths or path filters as arguments.
> Now, you may disagree that tolerating unknown fields is a features (as many people do), but one must understand the context where protobuf has been designed, namely the situation where it takes time to roll out new versions of binaries that process the data (either in API calls or on stored files) and thus the ability to design a schema evolution with backward and forward compatibility is worth a few more cycles during encoding.
This is the sort of thing I mean when I complain about the ignorant reinvention of the wheel that we all seem to engage in. It's natural and easy to do that, but it's not really a good thing.
Extensibility, versioning, and many many other issues in serialization are largely not new, and have been well-known and addressed for decades. ASN.1, for example, had no extensibility functionality in the early 1980s, but the only encoding rules at the time (BER/DER/CER), being TLV encodings, naturally supported extensibility ("skipping over unknown fields"). Later formal support for extensibility was added to ASN.1 so as to support non-TLV encodings.
ASN.1 also has elaborate support for "typed holes", which is what is referred to as "references" in [0].
ASN.1 gets a lot of hate, mainly for
a) its syntax being ugly (true) and hard to parse (true-ish)
b) the standard being non-free (it used to be non-free, but it's free now, in PDF form anyways)
c) lack of tooling (which is not true anymore).
(c) in particular is silly because if one invents a new syntax and encoding rules then one has to write the non-existent tooling.And every time someone re-invents ASN.1 they miss important features that they were unaware of.
Meanwhile ASN.1 is pluggable as to encoding rules, and it's easy enough to extend the syntax too. So ASN.1 covers XML and JSON even. There's no other syntax/standard that one can say that for!
Next time anyone invents a new syntax and/or encoding rules, do please carefully look at what's come before.
[0] https://en.wikipedia.org/wiki/Comparison_of_data-serialization_formatsRe: BuffDB is a Rust library to simplify multi-plexing on edge devices
#26Earlier quoted context omitted.
> Self-describing is point-less for serializations you couldn't be more wrong. what happens when you lose the schema, or never had access to it in the first place? think from the point of view of reverse engineering
I imagine when someone is choosing a serialization solution they probably don't really care about people trying to reverse engineer it... And if they did, they would just make schemas available instead. And if you lose your own schema, then you probably have more serious underlying problems.
Re: BuffDB is a Rust library to simplify multi-plexing on edge devices
#27Earlier quoted context omitted.
I imagine when someone is choosing a serialization solution they probably don't really care about people trying to reverse engineer it... And if they did, they would just make schemas available instead. And if you lose your own schema, then you probably have more serious underlying problems.
just imagine a world where protobuf replaces JSON, would you really want that? you're not thinking big picture
Ideally it is capnproto instead of protobuf, but in general the same applies.
It is a _transport_ protocol. You can still write your config or data in any format you like.
Re: BuffDB is a Rust library to simplify multi-plexing on edge devices
#28Earlier quoted context omitted.
Yes all those are indeed pessimizations during encoding but are features when decoding: the decoder can skip decoding fields and tolerate unknown fields. Now, you may disagree that tolerating unknown fields is a features (as many people do), but one must understand the context where protobuf has been designed, namely the situation where it takes time to roll out new versions of binaries that process the data (either…
> the decoder can [...] tolerate unknown fields. See ASN.1's extensibility rules. If you mark a record type (SEQUENCE, in ASN.1 parlance) as extensible, then when you later add fields (members, in ASN.1 parlance) the encoding rules have to make it possible to skip/ignore those fields when decoded by software implementing the pre-extension type. PER/OER will include a length for all the extension fields for this purpo…
One of the design requirements was simplicity and ease of implementation and for all the love in the world I can muster for ASN.1 I must admit it's far from simple.
IIRC complete and open implementations of ASN.1 were/are rare and the matrix of covered features didn't quite overlap between languages.
Re: BuffDB is a Rust library to simplify multi-plexing on edge devices
#29Earlier quoted context omitted.
just imagine a world where protobuf replaces JSON, would you really want that? you're not thinking big picture
I actually don't think it'd be so bad. Ideally it is capnproto instead of protobuf, but in general the same applies. It is a _transport_ protocol. You can still write your config or data in any format you like.
Re: BuffDB is a Rust library to simplify multi-plexing on edge devices
#30Earlier quoted context omitted.
I imagine when someone is choosing a serialization solution they probably don't really care about people trying to reverse engineer it... And if they did, they would just make schemas available instead. And if you lose your own schema, then you probably have more serious underlying problems.
just imagine a world where protobuf replaces JSON, would you really want that? you're not thinking big picture
Do I want protobufs replacing JSON? No. I want flatbufs augmenting JSON (and XML, and ...).
I maintain an ASN.1 compiler and run-time that supports transliteration of DER to JSON. This is possible because the syntax/schema language/IDL and the encoding rules are separable. This is how you get the best of both worlds. You can use optimized binary encoding rules for interchange but convert to/from JSON/XML/whatever as needed for inspection.