Live data from Hacker News

Multi-Core by Default

rfleury.com

11–20 of 61 posts

Re: Multi-Core by Default

#11
I think the actually interesting non-obvious part starts at "Redesigning Algorithms For Uniform Work Distribution". All the prior stuff done you basically get for free in a functional language that has some thread pool or futures built in, doing FP. The real question is how you write algorithms or parts of programs in a way that they lend themselves to be run in parallel as small units with results that easily merge again (parallel map reduce) or maybe don't even need to be merged again. That is the real difficult part, aside from transforming some mutating program into FP style and having appropriate data structures.

And then of course the heuristics start to become important. How much parallelism, before overhead eats the speedup?

Another question is energy efficiency. Is it more important to finish calculation as quickly as possible, or would it be OK to need some longer time, but in total calculate less, due to less overhead and no/less merging?

Re: Multi-Core by Default

#12
This post underscores how traditional imperative language syntax just isn't that well-suited to elegantly expressing parallelism. On the other hand, this is exactly where array languages like APL/J/etc. or array-based frameworks like NumPy/PyTorch/etc. really shine.

The list summation task in the post is just a list reduction, and a reduction can automatically be parallelized for any associative operator. The gory parallelization details in the post are only something the user needs to care about in a purely imperative language that lacks native array operations like reduction. In an array language, the `reduce` function can detect whether the reduction operator is associative and if so, automatically handle the parallelization logic behind-the-scenes. Thus `reduce(values, +)` and `reduce(values, *)` would execute seamlessly without the user needing to explicitly implement the exact subdivision of work. On the other hand, `reduce(values, /)` would run in serial, since division is not associative. Custom binary operators would just need to declare whether they're associative (and possibly commutative, depending on how the parallel scheduler works internally), and they'd be parallelized out-of-the-box.

Re: Multi-Core by Default

#13
post #5

The thing I struggle with is that most userland applications simply don't need multiple physical cores from a capacity standpoint. Proper use of concepts like async/await for IO bound activity is probably the most important thing. There are very few tasks that are truly CPU bound that a typical user is doing all day. Even in the case of gaming you are often GPU bound. You need to fire up things like Factorio, Cities…

I find async a terrible way to write interactive apps, because eventually something will take too long, and then suddenly your app jerks. So I have to keep figuring out manually which tasks need sending to a thread pool, or splitting my tasks into smaller and smaller pieces. I’m obviously doing something wrong, as the rest of the world seems to love async. Do their programs just do no interesting CPU intensive work?

Rest of the world doesn't love async. Just the loud opinionated people.

Re: Multi-Core by Default

#15
post #5

The thing I struggle with is that most userland applications simply don't need multiple physical cores from a capacity standpoint. Proper use of concepts like async/await for IO bound activity is probably the most important thing. There are very few tasks that are truly CPU bound that a typical user is doing all day. Even in the case of gaming you are often GPU bound. You need to fire up things like Factorio, Cities…

I find async a terrible way to write interactive apps, because eventually something will take too long, and then suddenly your app jerks. So I have to keep figuring out manually which tasks need sending to a thread pool, or splitting my tasks into smaller and smaller pieces. I’m obviously doing something wrong, as the rest of the world seems to love async. Do their programs just do no interesting CPU intensive work?

Probably. In web development it's usually get data, transform data, send data. That's in both directions, client to server and viceversa. Transformations are almost always simple. Native apps maybe do something more client side on average but I won't bet anything more than a cup of coffee on that.

Re: Multi-Core by Default

#16

Earlier quoted context omitted.

I find async a terrible way to write interactive apps, because eventually something will take too long, and then suddenly your app jerks. So I have to keep figuring out manually which tasks need sending to a thread pool, or splitting my tasks into smaller and smaller pieces. I’m obviously doing something wrong, as the rest of the world seems to love async. Do their programs just do no interesting CPU intensive work?

Rest of the world doesn't love async. Just the loud opinionated people.

[deleted]

Re: Multi-Core by Default

#17

Earlier quoted context omitted.

I find async a terrible way to write interactive apps, because eventually something will take too long, and then suddenly your app jerks. So I have to keep figuring out manually which tasks need sending to a thread pool, or splitting my tasks into smaller and smaller pieces. I’m obviously doing something wrong, as the rest of the world seems to love async. Do their programs just do no interesting CPU intensive work?

Are you using multiple threads or just a single one? Not sure why your application would "jerk" because something takes long time? If it's in a separate thread, it being async or not shouldn't matter, or if it's doing CPU intensive work or just sleeping.

If I’m using threads for each of my tasks, then why do I need async at all? I find mixing async and threads is messy, because it’s hard to take a lock in async code, as that blocks other async code from running. I’m sure this can be done well, but I failed when I tried.

Re: Multi-Core by Default

#18

I think this is less innovative than it seems. The approach described in this article is to reverse the good old fork/join, but it would only be practical for simple sub tasks or basic CLI tools, not entire programs. In the end, using this style is almost the same as doing fork/join, except the setup is somewhat hidden.

Based on the title I would have assumed that the programming model would be inverted, but it wasn't. What is needed is something akin to the Haskell model where the evaluation order is unspecified while simultaneously allowing mutation. The way to do this would be a Rust style linear type system where you are allowed to acquire exclusive write access to a region in memory, but not be allowed to perform any side effect and all modifications must be returned as if the function was referentially transparent. This is parallel by default, because you actively have to opt into a sequential execution order if you want to perform side effects.

The barriers to this approach are the same old problems with automatic parallelization.

Current hardware assumes a sequential instruction stream with hardware threads and cores and no hardware primitive in the microsecond range to rapidly schedule code to be executed on another core. This means you must split your program into two identical programs that then are managed by the operating system. This kills performance due to excessive amount of synchronization overhead.

The other problem is that even if you have low latency scheduling, you still need to gather a sufficient amount of work for each thread. Too fine grained and you run into synchronization overhead (no matter how good your hardware is), too coarse grained and you won't be able to spread the load onto all the processors.

There is also a third problem that is lurking in the dark and many developers with the exception of the Haskell community are underestimating: Running programs in a suboptimal order can lead to a massive increase in the instantaneous memory usage to the point where the program can no longer run. Think of a program allocating memory for each request, processing it and then deallocating, then allocating again. What if it accepts all requests in parallel? It will first allocate everything, then process everything and then deallocate everything.

Re: Multi-Core by Default

#19
If the author has not already, I would commend to them a search of the literature (or the relevant blog summaries) for the term "implicit parallelism". This was an academic topic from a few years back (my brain does not do this sort of thing very well but I want to say 10-20 years go) where the hope was that we could just fire some sort of optimization technique at normal code which would automatically extract all the implicit parallelism in the code and parallelize it, resulting in massive essentially-free gains.

In order to do this, the first thing that was done was to analyze existing source code and determine what the maximum amount of implicit parallelism was that was in the code, assuming it was free. This attempt then basically failed right here. Intuitively we all expect that our code has tons of implicitly parallelism that can be exploited. It turns out our intuition is wrong, and the maximum amount of parallelism that was extracted was often in the 2x range, which even if the parallelization was free it was only a marginal improvement.

Moreover, it is also often not something terribly amenable to human optimization either.

A game engine might be the best case scenario for this sort of code, but once you start putting in the coordination costs back into the charts those charts start looking a lot less impressive in practice. I have a sort of rule of thumb that the key to high-performance multithreading is that the cost of the payload of a given bit of coordination overhead needs to be substantially greater than the cost the coordination, and a games engine will not necessarily have that characteristic... it may have lots of tasks to be done in parallel, but if they

Re: Multi-Core by Default

#20
There is a deep literature on this in the High Performance Computing (HPC) field, where researchers traditionally needed to design simulations to run on hundreds to thousands of nodes with up to hundreds of CPU threads each. Computation can be defined as dependency graphs at the function or even variable level (depending on how granular you can make your threads). Languages built on top of LLVM or interpreters that expose AST can get you a long way there.
Post reply on HN