# Multi GPU support for Gemma 3

**URL:** <https://forum.modular.com/t/multi-gpu-support-for-gemma-3/1213>\
**Category:** MAX\
**Created:** [April 7, 2025, 3:35am UTC](https://forum.modular.com/t/multi-gpu-support-for-gemma-3/1213 "2025-04-07T03:35:26Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![zacksiri](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/zacksiri/32/396_2.png) [@zacksiri](https://forum.modular.com/u/zacksiri)\
**Post date:** [April 7, 2025, 3:35am UTC](https://forum.modular.com/t/multi-gpu-support-for-gemma-3/1213/1 "2025-04-07T03:35:26Z")

</div>

Hi

I just started experimenting with modular’s ecosystem. I managed to get up and running with llama 3 family of models however my daily driver for my main work is Gemma 3 class of models.

When I tried loading the model by setting the `--model-path` and `--weight-path` I get the error that says multi gpu support only works with llama models.

Would it be possible to get multi gpu support with non llama models? Also I use q8\_0 or q6\_k\_l would these quants be supported?

I have 2 A4500 GPU with Nvlink.

---

<div class="post-metadata">

**Author:** ![BradLarson](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/bradlarson/32/62_2.png) [@BradLarson](https://forum.modular.com/u/BradLarson)\
**Post date:** [April 7, 2025, 4:47pm UTC](https://forum.modular.com/t/multi-gpu-support-for-gemma-3/1213/2 "2025-04-07T16:47:29Z")

</div>

Sorry about MAX not having a drop-in solution today for your preferred Gemma 3 model architecture in a multi-GPU configuration. There are two things we need to expand support in MAX for: a native MAX Graph implementation of the `Gemma3ForConditionalGeneration` architecture family, and broadening multi-GPU support beyond the Llama-family models. Both are being tracked internally, although I can’t promise when they’ll be available.

As I mention in [this post](https://forum.modular.com/t/looking-for-examples-of-mulit-gpu-usage-with-mojo/1201/3), the full Python source code for the multi-GPU DistributedLlama3 model architecture [is available](https://github.com/modular/max/blob/main/src/max/pipelines/architectures/llama3/distributed_llama.py). If you really wanted to hack on it yourself, it is possible to extend that architecture to possibly cover the Gemma 3 models, but we are only starting to pull together tutorials and documentation around building your own models in MAX or extending existing ones.

For quantization, we currently support `Q6_K_M` quantization on CPU only, and have received requests to look into `Q8_0` quantization, so those capabilities are also tracked internally. QPTQ quantization is what we’ve favored so far on GPU for MAX models.

Thanks for the requests, that helps us determine priorities for bringup.

---

<div class="post-metadata">

**Author:** ![nlaanait](https://sea1.discourse-cdn.com/flex001/user_avatar/forum.modular.com/nlaanait/32/56_2.png) [@nlaanait](https://forum.modular.com/u/nlaanait)\
**Post date:** [April 7, 2025, 7:28pm UTC](https://forum.modular.com/t/multi-gpu-support-for-gemma-3/1213/3 "2025-04-07T19:28:26Z")

</div>

@BradLarson Great to hear that model building docs are in the works! Happy to give feedback on the docs once you have a draft.

btw I started the implementation of `mamba-2` architecture in `MAX`: [max-mamba](https://github.com/nlaanait/max-mamba).

---

<div class="post-metadata">

**Author:** ![warshanks](https://avatars.discourse-cdn.com/v4/letter/w/f08c70/32.png) [@warshanks](https://forum.modular.com/u/warshanks)\
**Post date:** [June 24, 2025, 4:57am UTC](https://forum.modular.com/t/multi-gpu-support-for-gemma-3/1213/4 "2025-06-24T04:57:43Z")

</div>

Just chiming in to say that I would love multi GPU capability for the Gemma 3 family. I’d love to be able to replace vLLM with Max for my Medgemma 4b deployment.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex001/uploads/modular/original/1X/2751e0fbdc595a99718b216730957e9db4448cfd.jpeg) [@system](https://forum.modular.com/u/system)\
**Post date:** [December 21, 2025, 4:57am UTC](https://forum.modular.com/t/multi-gpu-support-for-gemma-3/1213/5 "2025-12-21T04:57:54Z")

</div>

This topic was automatically closed 180 days after the last reply. New replies are no longer allowed.
