Live data from Hacker News

This Week In Servo 69

blog.servo.org

61–70 of 70 posts

Re: This Week In Servo 69

#61

Earlier quoted context omitted.

> zero-copy parallel parsing of html, zero copy yes, parallel I don't know We do selector matching in parallel (parsing is probably serial?), layout in parallel, and offload almost all of the rendering work onto the GPU. Not sure what you mean by "in-going data duplication". We try to be zero-copy when we can, and share as much as possible. Rust helps here; because you are free to try and share things without worryin…

Thanks! "in-going data duplication" ummm... i guess i meant ;) duplication of the incoming data (network packets->buffers->resources like html-files etc) in the sense of trying to minimize the amount of data duplicated and moved around to be used by various code-parts. Most probably the most interesting parts are those areas where even Rust can't really "help" to prevent that. thanks

.....

grins sheepishly

As Jack mentioned in another comment (https://news.ycombinator.com/item?id=11994607), a lot of our glue is slow just because we haven't paid attention to it. Some of this is duplication of the kind you specify.

So we do do a lot of extra copying. Especially once we switched to multiprocess -- there's a ton of data being passed between processes through a copy that really should use shared memory or something. It's pretty straightforward to fix in most cases, but we just haven't gotten to it. Like many other issues we have. Servo is mostly a testbed for the smaller components, so this is to some degree okay.

My favorite example of this is that we probably contain the worst cache implementation ever. Before multiprocess, we still used threads and senders a lot. So switching to multiprocess just involved replacing certain threads with processes, and using IPC instead of regular sync::mpsc senders.

One thing that got caught in the mix was our font cache. This cache loads fonts from disk and shares them with content threads till they aren't needed anymore. It was a simple cache, implemented as an atomic refcounted weak pointer being stored in the cache while strong refcounted pointers are being handed out to content threads.

In the process world, refcounted pointers are duplicated across IPC with their refcounts reset to one. So this involves copying the whole font across IPC. Inefficient in itself.

Also, on Linux, our IPC currently assumes /tmp is tempfs (it isn't in many cases), and uses it for shared memory. So, our cache, to save us from loading fonts from disk, loads a font from disk, and each time a process wants it, stores the file to disk in /tmp and has that process load it from disk. Which is unequivocally worse than what we had :)

It's not hard to fix, though. But it's stuff like this that we still have. It doesn't have to do with Rust, just has to do with priorities. My plan here is to first get a sane IPC shared memory implementation for Linux (to be used whenever we send gobs of data across a process), and build a better caching layer over it for the font cache. (If you like OS stuff and want to help ; let me know!)

However, the individual components themselves are pretty great!

Re: This Week In Servo 69

#62
post #44

Are they doing security testing as they go? I'm really looking forward to seeing how it fares against the typical "use after free" javascript errors that pwn to own always demo's. I'm not particularly a fan of Rust, but I certainly like the ideas that they are trying to incorporate in it. Great stuff. Looking forward to see how it evolves.

I enjoyed this 2014 blog post on how they avoid use-after-frees by design, by letting SpiderMonkey have responsibility for all DOM garbage collection: https://blog.mozilla.org/research/2014/08/26/javascript-serv...

Ooh. Nice one, mate. Thanks.

Re: This Week In Servo 69

#63
post #55

Earlier quoted context omitted.

Not to be negative, but I don't know if the benchmarks tell you much given the fact that Servo still has a ways to go before it renders sites properly. I periodically test it out (most recently yesterday, in fact) and on popular sites like espn.com and cnn.com it doesn't render them properly. Meaning, not the padding is a little off, but parts of the page simply don't render at all. This is not a criticism of their e…

You are generally correct, and this is why we don't often publish numbers. However we have made an effort to implement things we expect will have an effect on our performance first to counteract this. There are also sites included in our page load suite that Servo renders quite well, and should be a fair comparison. Another reason benchmarks are hard is that many of Servo's individual pieces are blazing fast, and we…

[deleted]

Re: This Week In Servo 69

#64
post #57

Earlier quoted context omitted.

Services which allow their users to post custom HTML and JavaScript to their own subdomains (without filtering to exclude scripts) need to go on that list to prevent eg evil.blogspot.com from stealing cookies that were set on innocent.blogspot.com

Why is that the responsibility of the browser and not the website's owner?

I don't understand either, with user generated subdomains I thought it was common practice to use a completely different domain for all trusted activity.

Re: This Week In Servo 69

#65

Earlier quoted context omitted.

Thanks! "in-going data duplication" ummm... i guess i meant ;) duplication of the incoming data (network packets->buffers->resources like html-files etc) in the sense of trying to minimize the amount of data duplicated and moved around to be used by various code-parts. Most probably the most interesting parts are those areas where even Rust can't really "help" to prevent that. thanks

..... grins sheepishly As Jack mentioned in another comment ( https://news.ycombinator.com/item?id=11994607 ), a lot of our glue is slow just because we haven't paid attention to it. Some of this is duplication of the kind you specify. So we do do a lot of extra copying. Especially once we switched to multiprocess -- there's a ton of data being passed between processes through a copy that really should use shared mem…

Thanks for the insights!

I did not mean to criticize servo or rust in any way... Any serious and non-tiny project will have areas like those you pointed out. I just think servo (and rust as its main 'tool') is a really great endeavor on the track to multi-process/parallel/concurrent (system)programming on the larger scheme of things ;). Any 'issues' you guys trip over are for sure nice lessons... which even potentially can feed back into rust. (which was the plan all along, afaik ;) Cheers

Re: This Week In Servo 69

#66
post #57

Earlier quoted context omitted.

Services which allow their users to post custom HTML and JavaScript to their own subdomains (without filtering to exclude scripts) need to go on that list to prevent eg evil.blogspot.com from stealing cookies that were set on innocent.blogspot.com

Why is that the responsibility of the browser and not the website's owner?

Nothing profound, just historical reasons.

To really show the problem, you have to do something like contrast how "blogspot.com" is a top-level site, one level below a TLD, but so is bbc.co.uk, one level below what "co.uk". The naive "count one element" doesn't work, or all of "co.uk" would share cookies. And it turns out that now there just isn't much you can do other than have a huge table. Sure, we'd probably do it differently if we had it to do all over again, but, we don't.

Re: This Week In Servo 69

#67
post #17

Earlier quoted context omitted.

So they achieved a 25% performance increase all from better parsing and a better algorithm for this[1] list? That's unexpected indead. I would love to see a blogpost with details on that. [1] https://publicsuffix.org/list/public_suffix_list.dat

They changed the implementation from always iterating over this 6000 length array: https://github.com/fduraffourg/servo/blob/8bb853f64354b2cc1b... to a HashSet which is only filled once based on a text file. The domain list also more easily updated now with a python script.

Given that they know the list at compile time I wonder if they could do faster e.g. by using https://github.com/sfackler/rust-phf to generate a perfect hash function over the set.

Re: This Week In Servo 69

#68
post #66

Earlier quoted context omitted.

Why is that the responsibility of the browser and not the website's owner?

Nothing profound, just historical reasons. To really show the problem, you have to do something like contrast how "blogspot.com" is a top-level site, one level below a TLD, but so is bbc.co.uk, one level below what "co.uk". The naive "count one element" doesn't work, or all of "co.uk" would share cookies. And it turns out that now there just isn't much you can do other than have a huge table. Sure, we'd probably do i…

You misunderstood. I fully understand why a count based approach cannot work. I don't understand why, should I want to create a service like blogspot, I would have to have my URL added in there.

Re: This Week In Servo 69

#69

Earlier quoted context omitted.

..... grins sheepishly As Jack mentioned in another comment ( https://news.ycombinator.com/item?id=11994607 ), a lot of our glue is slow just because we haven't paid attention to it. Some of this is duplication of the kind you specify. So we do do a lot of extra copying. Especially once we switched to multiprocess -- there's a ton of data being passed between processes through a copy that really should use shared mem…

Thanks for the insights! I did not mean to criticize servo or rust in any way... Any serious and non-tiny project will have areas like those you pointed out. I just think servo (and rust as its main 'tool') is a really great endeavor on the track to multi-process/parallel/concurrent (system)programming on the larger scheme of things ;). Any 'issues' you guys trip over are for sure nice lessons... which even potential…

Oh, I didn't take it as criticism. Was just slightly amused :)

Not sure the specific situations I mention can be fixed by Rust (or any language), those are design issues basically.

Re: This Week In Servo 69

#70
post #66

Earlier quoted context omitted.

Nothing profound, just historical reasons. To really show the problem, you have to do something like contrast how "blogspot.com" is a top-level site, one level below a TLD, but so is bbc.co.uk, one level below what "co.uk". The naive "count one element" doesn't work, or all of "co.uk" would share cookies. And it turns out that now there just isn't much you can do other than have a huge table. Sure, we'd probably do i…

You misunderstood. I fully understand why a count based approach cannot work. I don't understand why, should I want to create a service like blogspot, I would have to have my URL added in there.

You don't have to add it there. You can make it secure anyway. Public suffix list will mean that should your security get messed up, the browser prevents this anyway.
Post reply on HN