Earlier quoted context omitted.
The most commonly used optimization is to have just a load, or just a store, rather than a load-and-branch. Basically you access a page that you mprotect to trigger the handshake. Unrolling loops is also super common. Recognizing loops that have a bounded runtime is also common. My favorite technique to try one day is: 1. just record where the pollcheck points are but don't emit code there 2. to handshake with a thre…
Memory protection games just don't scale. You're modifying global structures, resulting in tons of contention. That's also why compacting GCs prefer to use card marking rather than mprotect. About 2. - that's exactly what I meant. The downside is that now you have to write a full x86 emulator, maybe including SIMD instructions. Good luck. And that's also what I want to try to avoid. Imagine generating a copy of the i…
I’m describing what production JVMs do. They do it because it’s a speedup for them. Like, those compacting GCs that you’re saying use card marking are using page faults for safe points more often than not.
It’s true that page protection games don’t result in speedups for generational store barriers. A safe point is not that.