Live data from Hacker News

System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

humprog.org

21–30 of 34 posts

Re: System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

#21

Linux is unusual in OS kernels in that direct system calls from arbitrary userspace code are supported and ABI-stable. This model has always been a terrible idea. It robs the system of an ability to intercept system calls in userspace before doing an expensive privilege-mode transition. If, instead, as on OpenBSD, the kernel enforced the rule that all system calls had to go through libc (or perhaps a big ntdll.dll-li…

> all system calls had to go through libc (or perhaps a big ntdll.dll-like Which makes containers crap on Windows and *BSD as they have to run the currect libc or equivalent. Thus you need to build a different container per OS version which sucks compared to Linux.

You understand that your container is using the VDSO today, right? A UAPI requirement to issue system calls through it wouldn't hurt your deployment story at all.

But sure, keep using SYSCALL, THE DEPENDENCY MUTILATOR. It's got what containers crave!

Re: System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

#22

Earlier quoted context omitted.

[flagged]

You might enjoy my work on the lone lisp language. I got rid of the libc and implemented an entire interpreter with nothing but Linux system calls. Been working on it and blogging about it for about 3 years now. http://github.com/lone-lang/lone/

[flagged]

Re: System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

#23
post #9

Earlier quoted context omitted.

> all system calls had to go through libc (or perhaps a big ntdll.dll-like Which makes containers crap on Windows and *BSD as they have to run the currect libc or equivalent. Thus you need to build a different container per OS version which sucks compared to Linux.

Windows doesn't even have its own libc.

In Window,s the last-userspace-before-kernel-mode layer is called ntdll.dll. Unlike msvcrt or any other libc, ntdll is universal and loaded into every process.

Re: System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

#24

Linux is unusual in OS kernels in that direct system calls from arbitrary userspace code are supported and ABI-stable. This model has always been a terrible idea. It robs the system of an ability to intercept system calls in userspace before doing an expensive privilege-mode transition. If, instead, as on OpenBSD, the kernel enforced the rule that all system calls had to go through libc (or perhaps a big ntdll.dll-li…

> This model has always been a terrible idea. I disagree. It's an amazing idea. It allows me to write freestanding programs without any C libraries. It allows compilers to have Linux system call builtins that directly generate the calling convention. I created an entire lisp interpreter with nothing but Linux system calls, completely freestanding. I've written a sort of manifesto around this: https://www.matheusmorei…

> It allows me to write freestanding programs without any C libraries.

KERNEL32.dll is not a C library (for once, its exported functions don't even use any of the default C calling conventions on x86).

> I created an entire lisp interpreter with nothing but Linux system calls, completely freestanding.

"Freestanding", as in "standing on top of an OS but nothing else"? Then using the OS-provided shared object that is the documented interface between the userspace and the kernel doesn't violate your free stand.

I mean, I too had written small interpreters that had only LoadLibraryW/GetProcAddress from kernel32.dll as their imports and nothing else.

> The instruction set is the correct abstraction for the system call entry point.

Why? A function call seems a much more appropriate abstraction for the system call entry point.

> There should be no "required C libraries".

There is no required C library on Windows, yet it doesn't use direct system calls.

> Forcing all programs to use the vDSO would force them all to not only implement the ELF spec but also to implement a small ELF linker.

Not really. Neither Windows nor UEFI require you to reimplement any linking functionality. The OS can simply give your program a pointer to a table of function pointers at your entry point... which it already can do, see the aux vector on Linux.

Re: System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

#25

Linux is unusual in OS kernels in that direct system calls from arbitrary userspace code are supported and ABI-stable. This model has always been a terrible idea. It robs the system of an ability to intercept system calls in userspace before doing an expensive privilege-mode transition. If, instead, as on OpenBSD, the kernel enforced the rule that all system calls had to go through libc (or perhaps a big ntdll.dll-li…

Direct system calls are an amazing idea. The NtDll and bsd models are worse. The whole libc becomes a security boundary without the protection of kernel space. So much windows malware and process tampering happens because now you have a library (ntdll) fully in userspace that is given special privileges, which now becomes a huge attack surface. Then you have to deal with breakages between the built in libc versions a…

Your argument about libc/ntdll having "special privileges" is a bit weird in that the alternate option is everything having those privileges. The ntdll tampering doesn't exist on Linux because it's not necessary. It's not better due to this.

Re: System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

#26

Earlier quoted context omitted.

> This model has always been a terrible idea. I disagree. It's an amazing idea. It allows me to write freestanding programs without any C libraries. It allows compilers to have Linux system call builtins that directly generate the calling convention. I created an entire lisp interpreter with nothing but Linux system calls, completely freestanding. I've written a sort of manifesto around this: https://www.matheusmorei…

> It allows me to write freestanding programs without any C libraries. KERNEL32.dll is not a C library (for once, its exported functions don't even use any of the default C calling conventions on x86). > I created an entire lisp interpreter with nothing but Linux system calls, completely freestanding. "Freestanding", as in "standing on top of an OS but nothing else"? Then using the OS-provided shared object that is t…

> "Freestanding", as in "standing on top of an OS but nothing else"?

Freestanding as in freestanding C.

> Then using the OS-provided shared object that is the documented interface between the userspace and the kernel doesn't violate your free stand.

Correct. I'm just saying it shouldn't be required.

> I mean, I too had written small interpreters that had only LoadLibraryW/GetProcAddress from kernel32.dll as their imports and nothing else.

And where are LoadLibraryW and GetProcAddress coming from? What if you had to implement those functions yourself?

> A function call seems a much more appropriate abstraction for the system call entry point.

The system call entry point is essentially its own calling convention. It pretty much is a function call. The function just happens to be identified by a stable number rather than function address.

> The OS can simply give your program a pointer to a table of function pointers at your entry point

It's not "a table of function pointers", it's a complete ELF object which you have to parse and resolve symbols from. That's a lot more work than putting the system call number and arguments in specific registers, executing one instruction and retrieving the return value from a specific register.

Re: System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

#27
post #25

Earlier quoted context omitted.

Direct system calls are an amazing idea. The NtDll and bsd models are worse. The whole libc becomes a security boundary without the protection of kernel space. So much windows malware and process tampering happens because now you have a library (ntdll) fully in userspace that is given special privileges, which now becomes a huge attack surface. Then you have to deal with breakages between the built in libc versions a…

Your argument about libc/ntdll having "special privileges" is a bit weird in that the alternate option is everything having those privileges. The ntdll tampering doesn't exist on Linux because it's not necessary . It's not better due to this.

Yeah. On Linux it's just an optimization. What user space really wanted was a way to memory map some kernel data into the process address space in order to avoid switching to kernel mode while accessing it. Instead Linux memory mapped an entire ELF whose only purpose is to wrap the data. Newer system calls like io_uring are doing it right.

Re: System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

#28
post #25

Earlier quoted context omitted.

Your argument about libc/ntdll having "special privileges" is a bit weird in that the alternate option is everything having those privileges. The ntdll tampering doesn't exist on Linux because it's not necessary . It's not better due to this.

Yeah. On Linux it's just an optimization. What user space really wanted was a way to memory map some kernel data into the process address space in order to avoid switching to kernel mode while accessing it. Instead Linux memory mapped an entire ELF whose only purpose is to wrap the data. Newer system calls like io_uring are doing it right.

Strongly disagree that providing the vDSO in ELF file format is somehow harmful or inefficient. You'll need a compatibility mechanism in any case since the exposed features will change over time, and doing that through normal symbol resolution avoids a whole bunch of extra effort. And after ld.so is done with relations on executable startup, it makes no difference in performance either.

Look at the Linux architectures that have a vDSO in non-ELF format. It's seriously ugly.

(I don't think the comparison with io_uring is valid either, very different kind of API.)

Re: System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

#29

Earlier quoted context omitted.

> It allows me to write freestanding programs without any C libraries. KERNEL32.dll is not a C library (for once, its exported functions don't even use any of the default C calling conventions on x86). > I created an entire lisp interpreter with nothing but Linux system calls, completely freestanding. "Freestanding", as in "standing on top of an OS but nothing else"? Then using the OS-provided shared object that is t…

> "Freestanding", as in "standing on top of an OS but nothing else"? Freestanding as in freestanding C. > Then using the OS-provided shared object that is the documented interface between the userspace and the kernel doesn't violate your free stand. Correct. I'm just saying it shouldn't be required. > I mean, I too had written small interpreters that had only LoadLibraryW/GetProcAddress from kernel32.dll as their imp…

> I'm just saying it shouldn't be required.

I'm still not entirely sure why you want to use instructions from the privilleged subset of your ISA instead of plain old "call fun_addr".

> And where are LoadLibraryW and GetProcAddress coming from?

They're provided by the OS. Their addresses are patched into your executable's image during the loading. It's a very ancient technology, one of the very first software technologies invented, in fact — predates FORTRAN.

> What if you had to implement those functions yourself?

What if you had to implement exec(2) yourself? As a matter of fact, why is exec even provided as a syscall? Almost all of it (except for locking the text segment IIRC) can be done in the user space, including the parsing of the program headers and relocating stuff. Which, again, I've done once and I appreciate the OS giving it to me already implemented.

> The function just happens to be identified by a stable number rather than function address.

Or you can identify it as a stable offset into a large table of function addresses; or even as a stable character string!

    intptr_t fd = invoke_ffi("kernel32.dll!CreateFileW", fn, GENERIC_READ | GENERIC_WRITE, etc.);
    if (fd 
> it's a complete ELF object which you have to parse and resolve symbols from.

You don't have to parse it. And UEFI environment in fact does give your program's entry point a table of function pointers: you put the arguments into the registers, take an offset into this table of functions, load the address, and call it, with one instruction, "call"/"branch-and-link", and it will give you the return value in a specific register. No need to parse anything by yourself.

I personally think this kind of dependency injection is pretty neat; you can intercept your own syscalls by passing pointer a modified table down your call stack. Trapping "sysenter" instruction in the userspace is way harder.

Re: System call instrumentation on Linux/x86‑64 using memory‑indirect calls, part I

#30
post #28

Earlier quoted context omitted.

Yeah. On Linux it's just an optimization. What user space really wanted was a way to memory map some kernel data into the process address space in order to avoid switching to kernel mode while accessing it. Instead Linux memory mapped an entire ELF whose only purpose is to wrap the data. Newer system calls like io_uring are doing it right.

Strongly disagree that providing the vDSO in ELF file format is somehow harmful or inefficient. You'll need a compatibility mechanism in any case since the exposed features will change over time, and doing that through normal symbol resolution avoids a whole bunch of extra effort. And after ld.so is done with relations on executable startup, it makes no difference in performance either. Look at the Linux architecture…

Well, no more harmful or inefficient than ELF itself. :-) I really wish we'd ended up with PE or something with a two-level namespace.

And yeah, nothing wrong with using ELF for the vDSO. People have strange intuitions about what's expensive and what's cheap.

Post reply on HN