Live data from Hacker News

Ladybird passes the Apple 90% threshold on web-platform-tests

twitter.com

121–130 of 272 posts

Re: Ladybird passes the Apple 90% threshold on web-platform-tests

#121

Earlier quoted context omitted.

While I haven't tried it myself, I've seen a few of the monthly summaries videos. Passing the tests and being fast enough for daily usage is two very different things and right now Ladybird doesn't appear to be all that speedy. Still an amazing feat of development from the entire team.

I was going to say the same thing. Why are the tests so disconnected from the usability? My assumption is the tests are closer to a unit test, while browsing a page is essentially an E2E test, and if anything in the pipeline goes wrong (especially given that we use complex JS everywhere) the result is essentially useless.

There's not a linear relationship between the tests and usability. There are many tests for various character encodings, but viewing a web page means you're only "using" one of them, for example.

As such, 90% test pass rate but low usability simply means that 10% of the tests cover a lot of very visible usability features that ladybird hasn't addressed yet.

Re: Ladybird passes the Apple 90% threshold on web-platform-tests

#122

I think it’s just fantastic that the Ladybird browser is close to being usable. I was under the impression this was going to take many years before it became competitive.

Three years ago I was very skeptical of Ladybird. But two things have changed. First, they have funding for 8 full time engineers, which I definitely wasn’t expecting. Second, it’s been three years. So given that, I am more optimistic. There’s still a very long way before they can compete with Chrome, of course. And I’m not sure I ever understood the value proposition compared to forking an existing engine.

The value proposition is not having vendor lockin and having WebKit/Blink be the defacto behaviour. For example the Ladybird team have found and raised spec issues in the different specs.

Another example is around ad blockers -- if Blink is the only option, they can make it hard for ad blockers to function whereas having other engines allows different choices to be made.

Re: Ladybird passes the Apple 90% threshold on web-platform-tests

#123

I think it’s just fantastic that the Ladybird browser is close to being usable. I was under the impression this was going to take many years before it became competitive.

It really goes to show what a dedicated team can accomplish. Before Ladybird it was taken for granted that building an entirely new browser engine would take decades and people would laugh at you for even bringing it up.

Well, it is going to take decades…

It’s a valuable, ambitious project, but it is going to take a while before it can be used for anything real.

Re: Ladybird passes the Apple 90% threshold on web-platform-tests

#124
post #108

I think it’s just fantastic that the Ladybird browser is close to being usable. I was under the impression this was going to take many years before it became competitive.

I do wonder if it's the case of "90% of completeness takes 90% of time; the remaining 10% takes another 90%". Though, I suppose even if true, it would still be a pretty good timeframe.

I'll guess that the remaining 10% will take more than another 90%, and also that it will keep growing as time goes on. Web standards are becoming more complex every day.

Re: Ladybird passes the Apple 90% threshold on web-platform-tests

#125
post #115

Earlier quoted context omitted.

No support for uBlock Origin and other tools that make the web sane

Orion is doing it somehow on iOS in a way I still don’t really understand.

As far as I know, they just emulate the Chrome extension API right?

Re: Ladybird passes the Apple 90% threshold on web-platform-tests

#126
post #100

Earlier quoted context omitted.

such as? I consider myself a power user and I've never run into anything I couldn't handle or get around. Genuinely curious.

No support for uBlock Origin and other tools that make the web sane

Chrome doesn't allow the full version of uBlock Origin on desktop, or any version of it on mobile.

How does Chrome have so much market share?

Re: Ladybird passes the Apple 90% threshold on web-platform-tests

#127
post #91

Earlier quoted context omitted.

The wpt score is not that well balanced. Look at https://staging.wpt.fyi/results/?product=servo&product=ladyb... : out of about 2 million tests, more than half are for the "encoding" category. Good encoding support is needed for sure, but likely not at that level of prevalence.

It seems my communication did not adequately convey the fact that I have no problem with the Ladybird team doing this. It makes perfect sense and is the right thing to do. However, a jump like that means precisely and exactly what I said it means; very suddenly, that metric became much more important to the team. It is written straight into the graph. A large number of encoding-related tests that were probably relati…

fwiw, I'm not imputing you any assumptions. I'm just pointing out that using wpt score as a criteria is not necessarily a good proxy for browser readiness. So I'm not sure why Apple uses that, other than... there's no other objective measure? Of course it's fair game for browser engines to improve their score!

Re: Ladybird passes the Apple 90% threshold on web-platform-tests

#128
post #91

Earlier quoted context omitted.

The wpt score is not that well balanced. Look at https://staging.wpt.fyi/results/?product=servo&product=ladyb... : out of about 2 million tests, more than half are for the "encoding" category. Good encoding support is needed for sure, but likely not at that level of prevalence.

It seems my communication did not adequately convey the fact that I have no problem with the Ladybird team doing this. It makes perfect sense and is the right thing to do. However, a jump like that means precisely and exactly what I said it means; very suddenly, that metric became much more important to the team. It is written straight into the graph. A large number of encoding-related tests that were probably relati…

Dude at this point just raise your knickers up and criticize the thing. You’ve got the most valuable observation about this topic on your side. The graph is jarring and for someone only recently made familiar with Goodhart’s Law this is a great example of it in practice. You must be further well-informed enough to defend any issues you actually have with the project outright instead of this small war of attrition playing out waaay down here.

Too much useful insight is withheld or misappropriated these days.

Re: Ladybird passes the Apple 90% threshold on web-platform-tests

#129
post #91

Earlier quoted context omitted.

The wpt score is not that well balanced. Look at https://staging.wpt.fyi/results/?product=servo&product=ladyb... : out of about 2 million tests, more than half are for the "encoding" category. Good encoding support is needed for sure, but likely not at that level of prevalence.

It seems my communication did not adequately convey the fact that I have no problem with the Ladybird team doing this. It makes perfect sense and is the right thing to do. However, a jump like that means precisely and exactly what I said it means; very suddenly, that metric became much more important to the team. It is written straight into the graph. A large number of encoding-related tests that were probably relati…

> However, a jump like that means precisely and exactly what I said it means; very suddenly, that metric became much more important to the team. It is written straight into the graph.

Not really, though. The latest jump was from implementing some CSS Typed OM features, which has been in-progress work for a while now. The 6k increase in the test score was a bit of a happy surprise. It's also not that much of a jump when you zoom out and see it's "just" a continuation of a steady increase in score over a long period.

Re: Ladybird passes the Apple 90% threshold on web-platform-tests

#130
post #45

Thank you for the belly-laugh. It's Goodhart's Law in graph form. "Oh, is this metric important ? Let me get right on that." No shade intended towards the Ladybird team. You were given the terms and you're behaving rationally in response to them. More power to you. It's just a fantastic demonstration of what it looks like to very suddenly be developing against a very specific metric.

While it is kinda unfortunate to have one unbalanced test suite as the major external progress indicator, there are.. like no other good options that don't involve someone manually going through like the top 1000 sites or something and checking whether they look good. That leaves having no progress indication whatsoever, which is also pretty bad.

And, in any case, implementing more of the standards is just simply good, and would need to be done at some point anyway.

Post reply on HN