Administrative Things
First, thanks Steffi for tackling async. It’s a big task with a lot sitting on top of it.
Since this is so important, I’ll ask everyone to try to keep this thread on track since Steffi is going to have to read the collective output of everyone with opinions on this.
Secondly, I’ll ask everyone to please respect that the compiler team is going to need substantial reasons to deviate from the points in the Decisions we have committed to section. The lowest levels of this interface are unlikely to be pretty, since they will have to handle every possible use-case for Mojo. There is interest in models of asynchronous computation which are strictly better than the proposed design, but I mean strictly in an academic sense. This means that if the model currently proposed can do express all of the same concepts, even if it requires 20 linear zero-sized types, a horrible abuse of origins, then you need to show that either your new model is better in at least one of the following categories and equivalent or extremely close to it in others; Runtime performance, compilation cost, ease of use for a particular use-case (performance may not suffer for this unless it’s otherwise impossible), memory usage, composability, and portability. I’ll ask that discussions of other designs which receive push-back under those criteria move to other threads where a resolution can be hashed out and, if the result is positive, then the model can be returned back to the main thread.
Open Questions
Async on the GPU. Should async def be usable inside GPU kernels, or only on the host to coordinate device work?
I think it absolutely should. async def and friends all form an abstraction over fairly primitive programming constructs (while and switch for one common mapping) which currently work on almost all of the MAX’s targets. Over time we may discover some targets are too limited to allow for many forms of async constructs, similar to those targets which do not allow for dynamic control flow the language feature will not be available there.
Type erasure. Storing different futures in one collection needs some form of existential or boxing, like Box<dyn Future> in Rust. How much does this depend on existentials landing in Mojo?
I think that, if we phrase existentials as the “make me a vtable please” feature, then I’d say we’d need some form of existentials for true dynamism. However, in many cases there is a compile-time known set of top-level coroutines such that a tagged union may be sufficient to de-virtualize everything. I think this means we can land “MVP” async without existentials, although its usage will be a bit un-ergonomic until existentials arrive, potentially depending on large hand-written TypeList values.
One other thing I’d like to poke at is the idea of something like a inline dyn Coroutine + size_bytes_leq[64]. Effectively, have the vtable be inline instead of behind another layer of indirection (yes, more binary size and memory consumption which isn’t great but it should be fine), and potentially the coroutine body as well, since this provides the guarantee that the actual data which matters is no larger than some amount, meaning that there is a safe way to have inline data without knowing the type. The “struct size” trait should be pretty cheap for the compiler to compute, given that it would be the amount of space needed for the vtable (potentially just a pointer) and space for the data, which is just size_of.
Function coloring. Given stackless coroutines, what can we do to reduce the cost of coloring? For example, a blocking wait() bridge, or APIs that are generic over sync and async.
My vote, although it might be a bit un-ergonomic in some cases, would be the following:
First, create def, sync def and async def. def means “I don’t care”, meaning that the function trait def() -> None is the union of async def() -> None and sync def -> None. Most functions in both user and library code are expected to be def functions, since most code doesn’t need to do fancy things with async and doesn’t really care about whether or not it makes a blocking syscall. sync def is for functions which are explicitly not async, and you may not call them in an async context. Similarly, async def is explicitly async and must be called in an async context.
From there, we have two options. In a def, either any function call which returns an awaitable has it automatically awaited (+ergonomics, -explicitness, potentially +confusion), or def must be written as “async correct” code where function calls to def or async def functions occur on await (potentially a lot of compiler transformation here). This second method is less ergonomic but is likely to cause less surprises.
Finally, at the lowest levels of code, you’ll run into why async def and sync def exist. These are for functions which are fundamentally one or the other, either due to making blocking syscalls themselves, or due to doing things like awaiting multiple things at once to increase productive CPU time (see Rust’s FuturesUnordered). Right before them, you have a bridge, of some form like the following:
def simple_print(s: String):
comptime if async:
await async_print(s)
else:
sync_print(s)
This async can be implemented as an “invisible parameter” passed down the call stack. It starts out sync, and then async executors can provide a primitive to “set” the “flag”. Additionally, awaiting a def function would make it definitely async since that means you’re in an async context. This means that, while we do technically end up with whole-program monomorphization to support this, it should be relatively cheap to support since it’s a boolean flag and I expect most programs mostly be written in one “dialect” (sync or async). This enables libraries to mostly be “synchness”-agnostic with a low-level IO layer handling the rest. This has the additional benefit of allowing the standard library to make std.atomic.Mutex or wherever we put it to behave “correctly” given the context, getting rid of a problem that Rust has where executors need to provide unique async synchronization primitives even for simple operations (although I expect executors will need to provide hooks in order to make this performant).
Use-cases
DPUs
DPUs, although they are nominally a fancy NIC, typically come with normal CPU cores running Linux. They also exist pretty much entirely for doing IO. As such, in my opinion they represent an easy target for “MAX should probably let me shove async code here so I can do data loading.”
Accelerator Driven IO
NVMe is a fundamentally asynchronous protocol, and given that most disks are moving towards using NVMe (we even have NVMe hard drives), it’s unlikely to go anywhere for a bit. Even if it does, I don’t expect that we as an industry going to suddenly decide that sync IO is a good idea again.
Given that production-grade KV caching now involves tiering to local and network storage, giving the Accelerator the ability to express the async protocol that is NVMe directly is very useful, because while we do have a lot of warps, there’s no reason to stall an entire warp per block requested from disk. If we’re already going to do async IO, then I think we should probably plumb async/await up to it.
Networking is also fundamentally asynchronous on pretty much every network in common use today. Having interactions with your scale-out network be able to be represented the same way on both the host and accelerator is quite nice.
DMA Devices
Another place this is useful is in expressing async transfers inside of accelerators. Once off of CPUs, it becomes fairly common to have some of “data mover” or DMA device you can ask to move things around for you, and on some platforms it’s mandatory if you want to access the accelerator’s main memory, such as AMD XDNA and Tenstorrent Tensix. In the latter case, where you may only have a few KiB of local memory, I don’t think anything other than precise control over the memory locations of coroutines is going to work.
General use in embedded devices
Async is a natural abstraction for cooperative multi-tasking, which is great for embedded devices which may not have the processing power to run a full threading abstraction or which may not have a large number of spare cores to support many threads.
No-allocator on larger servers
In networking, as traffic volume goes up, tighter and tighter coordination with the NIC is desirable for performance reasons. As a result, once one starts to push the server hard enough, manual control over memory becomes desirable. No-allocator generally means you need to provide the memory for everything, which, while typically a feature for embedded, is also very nice when working on a large server where you want to have a few hundred million coroutines floating around without getting smacked by time complexity.