Live data from Hacker News

W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

w3.org

221–230 of 287 posts

Re: W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

#221
post #169
post #59

Earlier quoted context omitted.

The biggest impact of WHATWG's work wasn't about features, but about defining HTML itself in a way that matched how web pages were written. Before HTML5, every browser parsed HTML in subtly incompatible ways. A document written for one browser could fail to render properly in another. Now, the process for converting bytes into a DOM tree is completely and precisely specified ( https://html.spec.whatwg.org/multipage/p…

> It's hard to overstate how important this sort of compatibility work has been to the strength of the modern web. There are fewer compatible browsers than at any time in the past 20 years. The compatibility problem has been solved by drowning browser maintainers in complexity and convincing them to give up. All the small browsers lack the manpower to implement the rapidly churning spec. This leaves the Google-funded…

I don't think that is a fair analysis.

Firstly, much of the complexity is an emergent property of interaction between a spec and reality: the engineering choices developers make when implementing a spec.

Pixel examples: (a) mitering of borders, (b) sizing four 25% width divs within a 99px div.

Pick just about any old spec, then look at the corner cases where developers have discovered different browsers act differently ("bugs"). The programming differences are often emergent and are not covered by the spec.

Developers create web pages that depend on the differences in a browser: that is a hard reality.

Secondly: Chrome mostly works better, follows specifications faster, and it is marketed better. Firefox has a significant budget, but it tools around with a bunch of shit that doesn't make their browser better. My interactions with Mozilla trying to get real bugs fixed have been poor. Safari and Microsoft were worse. Chrome cares about bugs, and fixes them in my experience.

Re: W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

#222
post #27

Earlier quoted context omitted.

It's not like anybody of us voted for W3C either (so it's not like one body is the "legitimate" one and other is not). And the W3C had stalled progress so much in the 00s and 10s, that it deservedly got sidelined.

They were focused on data organization and security instead of rounded corners or querySelectors, but most frontend developers saw that as stalling.

Neither data organization nor security is their role...

One belongs to the backend to decide, the other is for HTTP/S-level specs and Javascript, none of which are or should be W3C's concern.

Re: W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

#223
post #200

Earlier quoted context omitted.

XHTML 2.0 was very different from HTML semantically. Some of its tags (ARTICLE, SECTION, MENU, etc) wound up smashed into HTML5 later and became semantically meaningless again, but XHTML 2.0 tried to be more semantically sound than HTML and is a large part of how W3C lost the war to HTML5, because semantics are hard and most of the browsers didn't care about semantics.

Why to you consider the XHTML 2.0 elements more "semantically sound" than the equivalent HTML5 elements?

A reason XHTML 2.0 got so bogged down in committee and never actually finished a standard was that the attempt was made to define semantically what an ARTICLE would be, how SECTIONS work, what things a browser or semantic web crawler could infer/summarize build from such things. For instance, one group of the committee argued you couldn't have SECTIONs outside of an ARTICLE; that an ARTICLE consisted of zero or more SECTIONs (and maybe SECTIONs could be nested inside of each other). Folks argued for SECTIONs to have concepts of names that could be listed in auto-generated Tables of Contents.

HTML5 mostly just defines ARTICLE and SECTION as optional block-level content elements, with no other real importance. This leaves them as merely fancier synonyms for DIV. Semantics is almost entirely left to ARIA, and while HTML5 has come back around to ARTICLE tag should imply, for instance, ARIA role="article", there's still a bunch of interesting reasons that people concerned with ARIA semantics continue to write "redundant" things like <ARTICLE aria-role="article"...

Re: W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

#224

Earlier quoted context omitted.

The point is that HTML is almost 30 years old, and based on SGML which is much older (even though ISO 8879 is officially "only" from 1986), where SGML is just a formalization of typesetting practices established in the 1960's and 1970's. Given the depth of usage of HTML in everyday life (laws, contracts in ecommerce, medical records, personal communication, education, etc., etc.), I think HTML deserves better than be…

> I think HTML deserves better than being tinkered with all the time for no good reasons other than job security and/or achieving Webkit dominance, or other nebulous reasons at this point. That's good, because these aren't the reasons the HTML standard is changed, and to claim they are is absurd. HTML may have had its origins in SGML, but it has long, long since grown past beyond those origins to become the web platf…

There is a gulf of nuance between these two extremes you describe.

The web has "evolved" from a medium for simple self-publishing into a medium of mass surveillance and manipulation, big media, privacy-invading ads, uncalled-for browser monopoly, information oligopoly, and arbitrary crap code being sent to you in ridiculous quantities, not only draining your batteries and showing no respect for planet earth wrt energy efficiency, but also actively putting you in danger through fishing, xss and whatnot, and making your future ability to even read your legal, personal, study, business, or banking documents dependent on a needlessly over-complicated technology stack that no-one has the ability to influence in meaningful ways except Google, an ad company.

What it has not evolved into is a medium for long-term preservation of digital information, information autonomy, for simple ecommerce transactions and payments for everyone (as a merchant), for letting content producers thrive with quality content, or one that fosters free speech and diversity.

It has "evolved" by being captured to serve the interests of very few players, and fails the criteria of not having to appease computers or software programs that seem to be at war with one another.

Re: W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

#225

Earlier quoted context omitted.

The Web Authentication standard [0] seems super useful and something we really need in the web. [0] https://developer.mozilla.org/en-US/docs/Web/API/Web_Authent...

Yeah, this is pretty great. It's also comes from the W3C, not the WHATWG. (it might be a bit lost in the hierarchy of the thread by now, but the original comment was about the WHATWG taking over and monopolising the normal considered and democratic standardisation process of the W3C with their HTML5 "living" spec.)

W3C was a good fit for WebAuthn because the W3C is a body for corporations and by its nature WebAuthn is built and primarily implemented by corporations.

Not a criticism, by the way, sometimes that's just the right fit.

Re: W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

#226
post #9

I still maintain that the "living standard" is an oxymoron. It's a collaborative browser dev document. Don't get me wrong, that's great. However for everyone else an unversioned document, any part of which can change at any moment, is not what's usually thought of as a standard.

This obsession with "living" and constant change seems to be mostly confined to the web --- instead of settling on a spec and then leaving it alone and "doing what you can with what you have", those working on this stuff seem more inclined with continuing to make browsers change. I suspect at least part of the reason is to build a high barrier to entry and preserve the monopoly, keeping out competitors, given who the…

>I suspect at least part of the reason is to build a high barrier to entry and preserve the monopoly

That isn't very logical. New specs like Grid are introduced because they're insanely useful; not because browser vendors just want to make the spec more complex and hard to implement.

Re: W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

#227
post #169

Earlier quoted context omitted.

> It's hard to overstate how important this sort of compatibility work has been to the strength of the modern web. There are fewer compatible browsers than at any time in the past 20 years. The compatibility problem has been solved by drowning browser maintainers in complexity and convincing them to give up. All the small browsers lack the manpower to implement the rapidly churning spec. This leaves the Google-funded…

I don't think that is a fair analysis. Firstly, much of the complexity is an emergent property of interaction between a spec and reality: the engineering choices developers make when implementing a spec. Pixel examples: (a) mitering of borders, (b) sizing four 25% width divs within a 99px div. Pick just about any old spec, then look at the corner cases where developers have discovered different browsers act different…

> Firstly, much of the complexity is an emergent property of interaction between a spec and reality: the engineering choices developers make when implementing a spec.

Ah, clearly that makes it easier for third parties to maintain browsers. Or, wait, no -- that's yet another way to push players that don't have hundreds of millions per year worth of funding.

> Chrome mostly works better, follows specifications faster, and it is marketed better.

That's largely because of the way the specs are developed: They rubber stamp Chrome features.

Re: W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

#228

Earlier quoted context omitted.

> But programming has strict rules itself and will also fail if not adhered to so I never understood the complaint of "draconian" error checking in XML/XHTML. In most cases, if your code has some syntax error, the author of the code sees the syntax error; in the web case, if your code has some syntax error, the user sees the syntax error. That's the dramatic difference. The other reality is unlike program code, there…

Not true. XHTML and HTML have validators to check your markup for proper syntax and usage. You know that. You made one of them! Good writers of markup will always check with those before they ship it. In the case of user supplied markup, that's still an issue today with HTML.

I've been involved with multiple HTML and XML parsers, but never validators. :)

The reality ten years ago, when a number of prominent XML advocates were using XHTML (and actually using it as such, serving it as such), almost all of their sites had user input means where the input was sanitized well enough for HTML to be secure (and not have any markup injection), but not for XML well-formedness (they got all the markup injection risks in XML, but not all the other WF requirements). If the very people who claim XML is easy can't get it right, can everyone else?

Re: W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

#229

Earlier quoted context omitted.

This isn't against XHTML 1.1, which was HTML 4.01 shoved into a container of "if it isn't valid XML, display an error page instead;" rather, a lot of the hate is against XHTML 2.0, which decided to rip out HTML features such as forms, frames, most of the old elements such as or , and generally screw compatibility completely. For a reaction from a browser developer, see https://dbaron.org/log/20090707-ex-html

XHTML 2.0 also did screwy stuff with MIME types I believe. It required specific MIME types that a lot of browsers at the time didn't support, and because of that most browsers would straight-up refuse to render XHTML 2.0.

That… wasn't even the screwy stuff.

Yes, it was a requirement that XHTML 2.0 documents be served as application/xhtml+xml (which IE didn't support at the time), but that was really a non-issue (with so much renamed and moved around v. HTML 4.01 and its XML reformulations (XHTML 1.0, XHTML 1.1) because there was no graceful fallback story.

The bigger problem was that it required content to be served as application/xhtml+xml, and gave the same element and attribute names, in the same namespace, different semantics to what XHTML 1.0/XHTML5 gave them, with different implementation requirements (and it being impossible to satisfy both).

AFAIK this essentially got resolved by XHTML 2.0 being abandoned (in 2009, years after HTML 5 had moved to being jointly developed with the W3C).

Re: W3C and the WHATWG sign agreement to collaborate on single version of HTML, DOM

#230

Earlier quoted context omitted.

After writing XHTML for several years (giving it a full-faith attempt) I never understood the point. You keep repeating "semantically sound" but I can't fathom what that means in your context. I never saw any indication that XHTML brought significant practical semantic information or standardization over what HTML5 can do. It did add significant gratuitous verbosity that made XHTML documents much harder to read and e…

> You keep repeating "semantically sound" but I can't fathom what that means in your context. Documents built on the principle that their structure should relate to the meaning of their content rather than its presentation are what I consider "semantically sound". By negative example, an HTML document filled with div pyramids just to apply layout information, and obtuse class and id names are not what I consider sema…

One of the interesting things you could do with XHTML was ditch the HTML entirely and write an XML document expressing the semantic content, and then couple that with an XSLT stylesheet to convert it into XHTML for display. This way the exact same resource could be read by machines to get the semantic content, and then read by browsers and transformed into the display content.

I'm not sure if anyone ever actually used this seriously though. Definitely very "ivory tower" design. But in the abstract it's a cool idea.

Post reply on HN