Live data from Hacker News

Reverse Engineering Apple's typedstream Format

chrissardegna.com

21–26 of 26 posts

Re: Reverse Engineering Apple's typedstream Format

#22

Earlier quoted context omitted.

This is stuff is such a PIA to parse. I assume it's just different teams doing different features over the years, and being alternately repulsed/seduced by each format. Probably features are implemented as libraries so there isn't a master oversight - they aren't trying to make iMessage's internal formats follow a consistent plan, just let all the libs coexist...

As someone who used to work on that team, it’s so interesting hearing thoughts from external public on the team.

I would love to hear your thoughts as an insider.

Re: Reverse Engineering Apple's typedstream Format

#23
post #9

Earlier quoted context omitted.

Grandfather of Protobuf is ASN.1

Very much so. Pretty much all of these protocols are simplifications of asn1 and in some cases (like protobuf) there are a handful of things that got lost because the wire formats didn’t have them as they didn’t need them. A schema indicator being the single biggest flaw in protobuf.

Why is the lack of a schema indicator the biggest flaw of protobuf?

Re: Reverse Engineering Apple's typedstream Format

#24
post #23

Earlier quoted context omitted.

Very much so. Pretty much all of these protocols are simplifications of asn1 and in some cases (like protobuf) there are a handful of things that got lost because the wire formats didn’t have them as they didn’t need them. A schema indicator being the single biggest flaw in protobuf.

Why is the lack of a schema indicator the biggest flaw of protobuf?

If you parse a serialized protobuf byte array without having a .proto file, you have no way to dustinguish a byte string field from a nested message field. Thus you have no way to know how deep your parser should go.

Re: Reverse Engineering Apple's typedstream Format

#25
post #24
post #23

Earlier quoted context omitted.

Why is the lack of a schema indicator the biggest flaw of protobuf?

If you parse a serialized protobuf byte array without having a .proto file, you have no way to dustinguish a byte string field from a nested message field. Thus you have no way to know how deep your parser should go.

Semi-related, one of the `imessage-exporter` contributors provided a great write-up on reverse engineering the handwritten and digital touch message protobufs [0]. The reconstructed proto files are [1] [2].

[0]: https://github.com/trymoose/handwriting2svg/blob/0eb56cf4582...

[1]: https://github.com/ReagentX/imessage-exporter/blob/beeb853b2...

[2]: https://github.com/ReagentX/imessage-exporter/blob/beeb853b2...

Re: Reverse Engineering Apple's typedstream Format

#26
post #23

Earlier quoted context omitted.

Very much so. Pretty much all of these protocols are simplifications of asn1 and in some cases (like protobuf) there are a handful of things that got lost because the wire formats didn’t have them as they didn’t need them. A schema indicator being the single biggest flaw in protobuf.

Why is the lack of a schema indicator the biggest flaw of protobuf?

It makes it impossible to write a general purpose dissector that takes captured messages or bytes and figure out how to parse it.

All they needed was a varint at the head of any marshaled from to at least provide some scoping clue.

Post reply on HN