# Apple Silicon GPU support in Mojo

**URL:** <https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295>\
**Category:** GPU Programming\
**Created:** [September 21, 2025, 3:57pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295 "2025-09-21T15:57:27Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![BradLarson](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/bradlarson/32/62_2.png) [@BradLarson](https://forum.modular.com/u/BradLarson)\
**Post date:** [September 21, 2025, 3:57pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/1 "2025-09-21T15:57:27Z")

</div>

The latest nightly releases of Mojo (and our next stable release) include initial support for a new accelerator architecture: Apple Silicon GPUs!

We know that one of the biggest barriers to programming GPUs is access to hardware. It’s our hope that by making it possible to use Mojo to develop for a GPU present in every modern Mac, we can further democratize developing GPU-accelerated algorithms and AI models. This should also enable new paths of local-to-cloud development for AI models and more.

To get started, you need to have an Apple Silicon Mac (we support all M1 - M4 series chips) running macOS 15 or newer, with Xcode 16 or newer installed. The version of the Metal Shading Language we use (3.2, AIR bitcode version 2.7.0) needs the macOS 15 SDK, and you’ll get an error about incompatible bitcode versions if you run on an older macOS or use an older version of Xcode that doesn’t have the macOS 15 SDK.

You can clone our [`modular`](https://github.com/modular/modular) repository and try out one of our GPU function examples in the `examples/mojo/gpu-functions` directory. All but the `reduction.mojo` example should work on Apple Silicon GPUs today in the latest nightlies. Additionally, puzzles 1-15 of the [Mojo GPU puzzles](https://puzzles.modular.com/introduction.html) should now work on Apple Silicon GPUs with the latest nightly, and we’ve added `osx-arm64` as a supported architecture there.

# Current capabilities

This is just the beginning of our support for Apple Silicon GPUs, and many pieces of functionality still need to be built out. Known features that don’t work today include:

- Intrinsics for many hardware capabilities
  - Not all Mojo GPU examples work, such as `reduction.mojo` and the more complex matrix multiplication examples
  - GPU puzzles 16 and above need more advanced hardware features

- Basic MAX graphs
- MAX custom ops
- PyTorch interoperability
- Running AI models
- Serving AI models

I’ll emphasize that even simple MAX graphs, and by extension AI models, don’t yet run on Apple Silicon GPUs. In our Python APIs, `accelerator_count()` will still return 0 until we have basic MAX graph support enabled. Hopefully, that won’t be long.

# Next steps

We’ve identified many of the technical blockers to progressively enable the above. The current list of what we plan to work on includes:

- ~~Handle `MAX_THREADS_PER_BLOCK_METADATA` and similar aliases~~ (commit)
- ~~Support GridDim, lane\_id~~ ([commit](https://github.com/modular/modular/commit/ce7706c9ad81d9d8394cd4d01e3c125677689a14), [commit](https://github.com/modular/modular/commit/b22fc93094bf0c9f5ad458ff54ffd7727c2cec52))
- Enable `async_copy_*`
- ~~Convert arguments of an array type to a pointer type~~ (internal)
- ~~Support `bfloat16` on ARM devices~~ ([commit](https://github.com/modular/modular/commit/de8f2301ad5d9bf258e04962009377941cad5c36))
- Support `SubBuffer`
- Enable atomic operations
- Complete implementation of `MetalDeviceContext::synchronize`
- Enable captured arguments
- Support `print` and `debug_assert`

I apologize for some of the cryptic error messages you may get when hitting a piece of missing functionality, or encountering a system configuration we aren’t yet compatible with. We hope to improve the messaging over time, and to provide better guides for debugging failures.

# How this works

To learn more about how Mojo code is compiled to target Apple Silicon GPUs, check out [Amir Nassereldine’s detailed technical presentation from our recent Modular Community Meeting](https://youtu.be/t2hAfgDYHoc?si=5X7E70xAi_eMjCMN&t=1557). He did amazing work in establishing the fundamentals during his summer internship, and we are now building on that to advance Mojo on this new architecture.

In brief, a multi-step process is used to compile and run Mojo code on an Apple Silicon GPU. First, we compile Mojo GPU functions to Apple Intermediate Representation (AIR) bitcode. This is done through lowering to LLVM IR, and then specifically converting to Metal-compatible AIR.

Mojo handles interactions with an accelerator through the `DeviceContext` type. In the case of Apple Silicon GPUs, we’ve specialized this into a`MetalDeviceContext` that handles the next stages in compilation and execution.

The `MetalDeviceContext` uses the [Metal-cpp API](https://developer.apple.com/metal/cpp/) to compile the AIR representation into a `.metallib` for execution on device. Once the `.metallib` is ready, the `MetalDeviceContext` manages a Metal `CommandQueue`, and buffers operations for moving data, running a GPU function, and more. All of this happens behind-the-scenes and a Mojo developer doesn’t need to worry about any of it.

Code that you’ve written to run on an NVIDIA or AMD GPUs should mostly just work on an Apple Silicon GPU, assuming no device-specific features were being used. Obviously, different patterns will be required to get the most performance out of each GPU, and we’re excited to explore this new optimization space on Apple Silicon GPUs with you.

# Just the beginning

While we’d love help in bringing up Apple Silicon GPU support, some of the infrastructure for introducing support for new AIR intrinsics and compiling them to a `.metallib` currently requires Modular developers for implementation. We’ll get more of the basics in place before work moves primarily to the open-source standard library and kernels, at which point community members will be able to do a lot more to advance compatibility. Contributions are always welcome, but we don’t want you to hit missing non-public components and get frustrated by being unable to move forward.

We’ll share much more documentation and content on how to work with and optimize for this new hardware family, but we’re extremely excited about even these first few steps onto Apple Silicon GPUs. I’ll to try to keep this post up to date as we expand functionality.

---

<div class="post-metadata">

**Author:** ![HowardChu](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/howardchu/32/765_2.png) [@HowardChu](https://forum.modular.com/u/HowardChu)\
**Post date:** [September 21, 2025, 5:13pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/2 "2025-09-21T17:13:49Z")

</div>

Excited to try this!

---

<div class="post-metadata">

**Author:** ![jklaivins](https://avatars.discourse-cdn.com/v4/letter/j/87869e/32.png) [@jklaivins](https://forum.modular.com/u/jklaivins)\
**Post date:** [September 21, 2025, 6:15pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/3 "2025-09-21T18:15:30Z")

</div>

Is there a docker image / container that configures this? I work on a m4 mac air, would be fun to test that out, but I work in docker containers as a main driver..

---

<div class="post-metadata">

**Author:** ![BradLarson](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/bradlarson/32/62_2.png) [@BradLarson](https://forum.modular.com/u/BradLarson)\
**Post date:** [September 21, 2025, 7:11pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/4 "2025-09-21T19:11:38Z")

</div>

Our Docker containers largely use Linux within them. Due to the requirement for Xcode tooling to build the `.metallib` from Mojo code, you’d need to have a virtualized macOS environment in the container. We haven’t yet built a container like that.

---

<div class="post-metadata">

**Author:** ![jklaivins](https://avatars.discourse-cdn.com/v4/letter/j/87869e/32.png) [@jklaivins](https://forum.modular.com/u/jklaivins)\
**Post date:** [September 21, 2025, 7:13pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/5 "2025-09-21T19:13:53Z")

</div>

Googling / AI Overlording around it looks like you can’t even mount the gpu from mac into a container anyways. Unless someone knows something I don’t, I’m guessing the guidance on mojo + metal is via host.

---

<div class="post-metadata">

**Author:** ![vguerra](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/vguerra/32/459_2.png) [@vguerra](https://forum.modular.com/u/vguerra)\
**Post date:** [September 21, 2025, 8:12pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/6 "2025-09-21T20:12:11Z")

</div>

I went ahead an tried to run `vector_addition.mojo` within `examples/mojo/gpu-functions` but faced the following issue:

```bash
xcrun: error: sh -c ‘/Applications/Xcode.app/Contents/Developer/usr/bin/xcodebuild -sdk /Applications/Xcode.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX26.0.sdk -find metallib 2> /dev/null’ failed with exit code 17664: (null) (errno=No such file or directory)

```

If you are facing this, you have to make sure of 2 things:

### 1. Right developer directory is selected

Make sure that you have selected the right developer directory:

```bash
$> xcode-select -p
/Applications/Xcode.app/Contents/Developer

```

if the output is different, fix it by running

```bash
sudo xcode-select --switch /Applications/Xcode.app/Contents/Developer

```

### 2. Metal toolchain is available

Make sure that the Metal toolchain is available.

```bash
$> xcrun -sdk macosx metal
metal: error: no input files

```

if the output is instead:

```bash
error: error: cannot execute tool 'metal' due to missing Metal Toolchain; use: xcodebuild -downloadComponent MetalToolchain

```

proceed to install the Metal toolchain:

```bash
xcodebuild -downloadComponent MetalToolchain

```

if you run `xcrun -sdk macosx metal`, it should give you the `no input files` error.

### You should be able to run mojo code on the GPU

If everything is well set, you can now run the example code.

```bash
$> mojo vector_addition.mojo
Resulting vector: 3.75 3.75 3.75 3.75 3.75 3.75 3.75 3.75 3.75 3.75

```

---

<div class="post-metadata">

**Author:** ![Dafmdev](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/dafmdev/32/713_2.png) [@Dafmdev](https://forum.modular.com/u/Dafmdev)\
**Post date:** [October 4, 2025, 2:11am UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/8 "2025-10-04T02:11:36Z")

</div>

> [@HowardChu](#):
>
> Excited to try this!

Excited to try this! x2

---

<div class="post-metadata">

**Author:** ![Dafmdev](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/dafmdev/32/713_2.png) [@Dafmdev](https://forum.modular.com/u/Dafmdev)\
**Post date:** [October 4, 2025, 3:37am UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/9 "2025-10-04T03:37:50Z")

</div>

I don’t know how to express the happiness I feel right now.

```auto
Found GPU: Apple M1 Pro
LHS buffer: HostBuffer([0.0, 1.0, 2.0, ..., 997.0, 998.0, 999.0])
RHS buffer: HostBuffer([0.0, 0.5, 1.0, ..., 498.5, 499.0, 499.5])
Result vector: HostBuffer([0.0, 1.5, 3.0, ..., 1495.5, 1497.0, 1498.5])

```

---

<div class="post-metadata">

**Author:** ![clattner](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/clattner/32/201_2.png) [@clattner](https://forum.modular.com/u/clattner)\
**Post date:** [October 5, 2025, 3:24am UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/10 "2025-10-05T03:24:43Z")

</div>

nice!

---

<div class="post-metadata">

**Author:** ![yiakwy-ml-core-team](https://avatars.discourse-cdn.com/v4/letter/y/ba8739/32.png) [@yiakwy-ml-core-team](https://forum.modular.com/u/yiakwy-ml-core-team)\
**Post date:** [November 12, 2025, 7:07am UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/11 "2025-11-12T07:07:45Z")

</div>

Recently we are working on M3 Ultra, and I found its extremely useful in deploying agentic LLM in one stop with 512 GB memory.

I observed great performance for serving gpt-oss-120b . As for kernel DSA, I quickly realized clatter used to work in Apple, why should I try mojo ?

When one finds it hard to find equivalent of triton in Apple silicon, Clattner and his team is the answer !

---

<div class="post-metadata">

**Author:** ![fzngagan](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/fzngagan/32/802_2.png) [@fzngagan](https://forum.modular.com/u/fzngagan)\
**Post date:** [November 28, 2025, 9:26am UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/12 "2025-11-28T09:26:00Z")

</div>

Ah, I just realized that Mac support is a recent thing. It’d be lovely to be able to run mac gpus on jupyter lab too. Posted about it:

> [@\`has\_apple\_gpu\_accelerator()\` is False on jupyter-lab on my Macbook](https://forum.modular.com/t/has-apple-gpu-accelerator-is-false-on-jupyter-lab-on-my-macbook/2488):
>
> from sys import has\_apple\_gpu\_accelerator print(has\_apple\_gpu\_accelerator()) gives False on jupyter-lab on my Macbook AIr M4. Is there a way to get it to work in jupyter? Also, can I get syntax highlighting in jupyter for mojo code?

---

<div class="post-metadata">

**Author:** ![BradLarson](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/bradlarson/32/62_2.png) [@BradLarson](https://forum.modular.com/u/BradLarson)\
**Post date:** [December 19, 2025, 9:57pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/13 "2025-12-19T21:57:34Z")

</div>

Realized I hadn’t updated this thread in a while, despite some significant advancements in Apple silicon GPU support.

In the time since this was originally written, many intrinsics and other capabilities have been implemented, greatly expanding the surface of GPU programming compatible with these devices. That is now reflected in [all but three of the non-NVIDIA-specific Mojo GPU puzzles](https://puzzles.modular.com/howto.html#gpu-support-matrix) now being compatible with Apple silicon GPUs! Those last three all require additions to our PyTorch interoperability to support the differences between Apple devices and CUDA / HIP ones, and we’re working on that.

Special thanks to @Ethan for landing several PRs to help us expand Apple silicon support!

Some basic one-operation MAX graphs using Mojo custom operations are functional today, but we’re still working to enable most basic graphs on the path to getting models working on this hardware.

We’ve also slightly broadened the range of Apple silicon devices compatible with Mojo, adding the M5 series devices in the nightly that went out yesterday (2025121805). Macs with M1-M5 systems-on-chip should be usable with Mojo now, although we are investigating some differences we’ve observed on older M1 devices.

As always, keep your eyes on the nightlies for the latest capabilities as we roll them out.

---

<div class="post-metadata">

**Author:** ![koliyo](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/koliyo/32/65_2.png) [@koliyo](https://forum.modular.com/u/koliyo)\
**Post date:** [January 30, 2026, 10:00am UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/14 "2026-01-30T10:00:45Z")

</div>

Any comments on timeline for debugger for Apple Silicon, similar to `mojo-cuda-gdb`?

> **[Debugging | Modular](https://docs.modular.com/mojo/tools/debugging/)**
>
> Debugging Mojo programs.

---

<div class="post-metadata">

**Author:** ![BradLarson](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/bradlarson/32/62_2.png) [@BradLarson](https://forum.modular.com/u/BradLarson)\
**Post date:** [January 30, 2026, 4:25pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/15 "2026-01-30T16:25:24Z")

</div>

I’ll be honest, that may take a little bit. We’re still working on some more basic debugging functionality like being able to `print()` or use `debug_assert()` within GPU functions on Apple silicon GPUs. Adding print support itself [can be a journey](https://github.com/modular/modular/blob/main/docs/eng-design/docs/amd-printf-lessons-learned.md).

---

<div class="post-metadata">

**Author:** ![bwibking](https://avatars.discourse-cdn.com/v4/letter/b/dc4da7/32.png) [@bwibking](https://forum.modular.com/u/bwibking)\
**Post date:** [March 11, 2026, 5:25pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/16 "2026-03-11T17:25:08Z")

</div>

Are device kernels allowed to use host buffers on Apple Silicon?

I have tried this, and the compiler allows it, but the kernels give incorrect results. I can only get it working correctly if I do stage-in/stage-out copies from “host” buffers to “device” buffers as if they memory were actually physically separate.

I’ve created an issue with a minimal reproducer here: [[BUG] Zero-copy write in device kernels gives incorrect results on Apple Silicon GPUs · Issue #6145 · modular/modular · GitHub](https://github.com/modular/modular/issues/6145)

---

<div class="post-metadata">

**Author:** ![BradLarson](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/bradlarson/32/62_2.png) [@BradLarson](https://forum.modular.com/u/BradLarson)\
**Post date:** [March 11, 2026, 8:19pm UTC](https://forum.modular.com/t/apple-silicon-gpu-support-in-mojo/2295/17 "2026-03-11T20:19:23Z")

</div>

We probably still have some baked-in assumptions about device and host memory being separate, shared-memory systems are still fairly new for the platform. We can definitely look into it.
