Zero-copy host buffers for Metal kernels on Apple Silicon

I ran into this while trying a GPU implementation of 1BRC on an M3. My input is already memory mapped, but the documented Mojo flow appears to require allocating a DeviceBuffer and copying the input into it:

var input = ctx.enqueue_create_buffer[DType.uint8](size)
ctx.enqueue_copy(input, mmap_ptr)

Metal can wrap a suitable existing allocation without copying via makeBuffer(bytesNoCopy:...). Is there a supported Mojo/AsyncRT way to do the equivalent for host memory?

If not, could DeviceContext expose something roughly like this?

var input = ctx.wrap_external_buffer[DType.uint8](
    mmap_ptr, size, ownership=.borrowed
)

I found #6145, now tracked under the Apple Silicon GPU epic #5468, but that issue is about passing an UnsafePointer directly to a kernel. I am asking whether Mojo could provide an explicit, lifetime-safe way to register existing host memory as a device buffer.

It would need to be existing host memory that satisfies all of Apple’s mostly undocumented requirements for what metal buffers need to be.

6145 will require this feature as part of the work, but there may be some reverse engineering work as part of that.