awni/mlx_distributed_deepseek.md

Last active February 7, 2026 10:32

Star (143) You must be signed in to star a gist
Fork (18) You must be signed in to fork a gist

Select an option

Learn more about clone URLs
Clone this repository at <script src="https://gist.github.com/awni/ec071fd27940698edd14a4191855bba6.js"></script>
Save awni/ec071fd27940698edd14a4191855bba6 to your computer and use it in GitHub Desktop.

Download ZIP

Run DeepSeek R1 or V3 with MLX Distributed

Raw

mlx_distributed_deepseek.md

Setup

On every machine in the cluster install openmpi and mlx-lm:

conda install conda-forge::openmpi
pip install -U mlx-lm

Next download the pipeline parallel run script. Download it to the same path on every machine:

curl -O https://raw.githubusercontent.com/ml-explore/mlx-examples/refs/heads/main/llms/mlx_lm/examples/pipeline_generate.py

Make a hosts.json file on the machine you plan to launch the generation. For two machines it should look like this:

[
  {"ssh": "hostname1"},
  {"ssh": "hostname2"}
]

Also make sure you can ssh hostname from every machine to every other machine. Check-out the MLX documentation for more information on setting up and testing MPI.

Set the wired limit on the machines to use more memory. For example on a 192GB M2 Ultra set this:

sudo sysctl iogpu.wired_limit_mb=180000

Run

Run the generation with a command like the following:

mlx.launch \
  --hostfile path/to/hosts.json \
  --backend mpi \
  path/to/pipeline_generate.py \ 
  --prompt "What number is larger 6.9 or 6.11?" \
  --max-tokens 128 \
  --model mlx-community/DeepSeek-R1-4bit

For DeepSeek R1 quantized in 3-bit you need in aggregate 350GB of RAM accross the cluster of machines, e.g. two 192 GB M2 Ultras. To run the model quantized to 4-bit you need 450GB in aggregate RAM or three 192 GB M2 Ultras.

Tmoss11 commented Feb 11, 2025

Thank you for the interesting and helpful writeup! I'm excited to run this on a large number of hosts if possible.

I've got 31x M1 hosts provisioned with 16 GB RAM configured with iogpu.wired_limit_mb=12000, running macOS 15.2, openmpi 5.0.6, mlx 0.22.1, mlx-lm 0.21.4 and validated all hosts can reach one another via SSH key.

I've tested both DeepSeek R1 3-bit and 2-Bit but end up with the same MLX error + stack trace from multiple hosts after the shards are downloaded. I was just wondering if you had any suggestions for troubleshooting here?

Traceback (most recent call last):
  File "/Users/administrator/mlx-distributed/pipeline_generate.py", line 112, in <module>
    for response in stream_generate(
  File "/opt/homebrew/Caskroom/miniconda/base/lib/python3.12/site-packages/mlx_lm/utils.py", line 529, in stream_generate
    for n, (token, logprobs) in enumerate(token_generator):
  File "/opt/homebrew/Caskroom/miniconda/base/lib/python3.12/site-packages/mlx_lm/utils.py", line 307, in generate_step
    y, logprobs = _step(y)
                  ^^^^^^^^
  File "/opt/homebrew/Caskroom/miniconda/base/lib/python3.12/site-packages/mlx_lm/utils.py", line 279, in _step
    logits = model(y[None], cache=prompt_cache)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/homebrew/Caskroom/miniconda/base/lib/python3.12/site-packages/mlx_lm/models/deepseek_v3.py", line 474, in __call__
    out = self.model(inputs, cache, mask)
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/homebrew/Caskroom/miniconda/base/lib/python3.12/site-packages/mlx_lm/models/deepseek_v3.py", line 435, in __call__
    mask = create_attention_mask(h, cache)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/homebrew/Caskroom/miniconda/base/lib/python3.12/site-packages/mlx_lm/models/base.py", line 50, in create_attention_mask
    if cache is not None and cache[0] is not None:
                             ~~~~~^^^
IndexError: list index out of range
    mask = create_attention_mask(h, cache)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/homebrew/Caskroom/miniconda/base/lib/python3.12/site-packages/mlx_lm/models/base.py", line 50, in create_attention_mask
    if cache is not None and cache[0] is not None:
                             ~~~~~^^^
IndexError: list index out of range

leozusa commented Mar 5, 2025

Can two of the new Mac Studios with M3 Ultra and max 512Gb of unified memory, and networked using Thunderbolt 5, run the non-quantized R1 version? (saw the news and got curious)

Author

awni commented Mar 5, 2025

You could run the 8 bit model with 1T RAM. That's quantized but perf should be about the same as the original fp8.

drixs2050 commented Mar 8, 2025

Is there a limit of setting the available ram for GPU? Just wondering for the coming 512gb mac studio how much I can squeeze out for GPU alone, I assume that if I can leave only something like 16gb for os on 2 machines and get 496x2 vram for deepseek r1 I can run the full version with fp16 on core attention and fp8 on the rest of the params?
Also can mlx utilize multiple tb5 connect bandwidth? Since the mac studio comes with multiple tb5 port it would be nice if we can use all of them.

fengyy0111 commented Mar 10, 2025 •

edited

Loading

I used two 192G Mac Studio to run the DeepSeeker R1-3bit model, using the following command: "mpirun-np 2-- hostfile hosts. txt python3 pipine_generation. py -- prompt" What number is larger 6.9 or 6.11? "-- model mlx community/DeepSeeker R1-3bit".
I have limited the memory usage limit on both devices, but the remote device crashed due to high memory usage. How can I solve this problem?

Author

awni commented Mar 10, 2025 •

edited

Loading

but the remote device crashed due to high memory usage

Could you share the error message? I don't see it in your post.

Did you set the sysctl like so for both machines?

sudo sysctl iogpu.wired_limit_mb=180000

fengyy0111 commented Mar 11, 2025

Yes, I have set up sysctl for both devices, which allows me to control remote device downloads of models. However, downloads usually encounter errors and display network issues. I am in China, is it due to regional restrictions?
`mpirun -np 3 --hostfile hosts.txt python3 pipeline_generate.py --prompt "What number is larger 6.9 or 6.11?" --model mlx-community/DeepSeek-R1-3bit
/Users/zhangchi/Library/Python/3.9/lib/python/site-packages/urllib3/init.py:35: NotOpenSSLWarning: urllib3 v2 only supports OpenSSL 1.1.1+, currently the 'ssl' module is compiled with 'LibreSSL 2.8.3'. See: urllib3/urllib3#3020
warnings.warn(
/Users/zhangchi/Library/Python/3.9/lib/python/site-packages/urllib3/init.py:35: NotOpenSSLWarning: urllib3 v2 only supports OpenSSL 1.1.1+, currently the 'ssl' module is compiled with 'LibreSSL 2.8.3'. See: urllib3/urllib3#3020
warnings.warn(
Fetching 5 files: 100%|██████████| 5/5 [00:00<00:00, 22525.80it/s]
Fetching 5 files: 100%|██████████| 5/5 [00:00<00:00, 46707.17it/s]
Fetching 70 files: 17%|█▋ | 12/70 [06:54<37:14, 38.52s/it]--------------------------------------------------------------------------
PRTE has lost communication with a remote daemon.

HNP daemon : [prterun-Mac-Studio-6-41295@0,0] on node Mac-Studio-6
Remote daemon: [prterun-Mac-Studio-6-41295@0,1] on node Mac-Studio-8

This is usually due to either a failure of the TCP network
connection to the node, or possibly an internal failure of
the daemon itself. We cannot recover from this failure, and
therefore will terminate the job.
--------------------------------------------------------------------------`

zengqingfu1442 commented Mar 17, 2025 •

edited

Loading

Does the new Mac Studios with M3 Ultra and max 512Gb of unified memory can run native FP8 deepseek-r1 model? Does the new Mac Studios with M3 Ultra and max 512Gb of unified memory support FP8?

Author

awni commented Mar 17, 2025

You can run a 4-bit quantized model on the 512GB machine. 8-bit is too big.

zengqingfu1442 commented Mar 18, 2025

You can run a 4-bit quantized model on the 512GB machine. 8-bit is too big.

If i hvae 4 such Mac Studios, each with 512GB, how can i run the 8-bit model distributed on the 4 machines?

Author

awni commented Mar 18, 2025

If each machine has 512 you only need 2 of them (and I wouldn't recommend using more because it will be slower)
The above setup should work.. though we'd need to make an 8-bit quant for that.

Author

awni commented Mar 18, 2025

You can make the 8-bit quant like so once ml-explore/mlx-lm#32 lands.

mlx_lm.convert --hf-path deepseek-ai/DeepSeek-R1 -q --q-bits 8 --upload-repo mlx-community/DeepSeek-R1-8bit

zengqingfu1442 commented Mar 18, 2025 •

edited

Loading

If each machine has 512 you only need 2 of them (and I wouldn't recommend using more because it will be slower)

The above setup should work.. though we'd need to make an 8-bit quant for that.

Why does 4 machines slower than 2 machines? Because the connection speed of Thunderbolt 5 is too slow? I thought that more machines mean more kvcache space and it would be faster.

Author

awni commented Mar 18, 2025

With pipeline parallelism (which is used here).. if you have enough RAM to fit the model then you only are adding communication latency as you add more machines.

zengqingfu1442 commented Mar 18, 2025

With pipeline parallelism (which is used here).. if you have enough RAM to fit the model then you only are adding communication latency as you add more machines.

Does mlx-lm support tensor parallelism?

Author

awni commented Mar 18, 2025

Sort of. There is a PR for it.. but it's too slow right now to be practical. Something we are working on / hoping to improve for the future.

jiyzhang commented Mar 31, 2025

404 for the url
https://raw.githubusercontent.com/ml-explore/mlx-examples/refs/heads/main/llms/mlx_lm/examples/pipeline_generate.py

The run script can be downloaded at
https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/examples/pipeline_generate.py

Basten7 commented Apr 11, 2025 •

edited

Loading

Good News

pipeline_generate.py work very well with other DeepSeek model "DeepSeek-V2.5-1210-3bit "

mlx.launch --hosts mac1,mac2 --backend mpi "pipeline_generate.py" --max-tokens 12800 --model mlx-community/DeepSeek-V2.5-1210-3bit --prompt "Generate a python script"

==========
Prompt: 21 tokens, 85.378 tokens-per-sec
Generation: 776 tokens, 17.794 tokens-per-sec
Peak memory: 55.234 GB

mlx.launch --hosts mac1,mac2 --backend mpi "pipeline_generate.py" --max-tokens 12800 --model mlx-community/DeepSeek-V2.5-1210-4bit --prompt "Generate a python script"

==========
Prompt: 21 tokens, 80.473 tokens-per-sec
Generation: 901 tokens, 17.410 tokens-per-sec
Peak memory: 70.257 GB

Less good News

1°) When I run mlx_distributed_deepseek.py
error message :

except statement is broken in "distributed_run.py"

Edit around line 175. Find:
in the file "except e:"
replace with
"except Exception as e:"

2°) And when I run this command: mlx.distributed_config --verbose --hosts
error message :

/miniconda3/envs/mlxmpi/lib/python3.11/site-packages/mlx/distributed_run.py", line 507, in prepare_tb_ring
connected_to = items[0]["domain_uuid_key"]
~~~~~~~~^^^^^^^^^^^^^^^^^^^
KeyError: 'domain_uuid_key'

zengqingfu1442 commented Apr 15, 2025

Does mlx support gguf format?

RaylenFarnor commented May 5, 2025 •

edited

Loading

To set up a machine cluster for MLX generation, install OpenMPI and mlx-lm on each machine using conda and pip. Download the pipeline script and place it in the same path on all machines. Create a hosts.json file listing the machines, ensuring SSH access between them. Set memory limits with sudo sysctl. If you’ve ever had to finish an essay overnight, you know the pressure. I’ve relied on UKWritings in such cases their write my essay service at https://ukwritings.com/write-my-essay is perfect for fast, high-quality academic writing that meets urgent deadlines.

georgiedekker commented Aug 8, 2025

Also works for DeepSeek v2. But not for other models.

would it be possible, to make a smaller model have this option of distributed, just to be able to get it to work on less expensive hardware haha. I have a few base model m4 mac minis and before i fork out for the m3 ultras, I'd love to see it actually worked. tried many ways to achieve it with grpc, tensor parallelism, etc. most of the time things got corrupted in the kv cache synchronization.

Author

awni commented Aug 8, 2025

You can try one of the smaller DSV2 models like mlx-community/DeepSeek-Coder-V2-Lite-Instruct-4bit

georgiedekker commented Aug 9, 2025

You can try one of the smaller DSV2 models like mlx-community/DeepSeek-Coder-V2-Lite-Instruct-4bit

Thank you so much Awni, I was trying to convince Claude Code to have it work with Qwen models but it just would not get it to work properly. Then I guess it had too much memory of that approach to accept your guidance. Finally got it to work and I asked Claude to do a beginner friendly writeup, please review/comment as i think this would be a nice example that many people would be able to run.
https://github.com/georgiedekker/mlx_distributed_ring_inference

georgiedekker commented Aug 15, 2025

@awni made another attempt at speeding things up. seems to work please review, feel free to use any as example in the mlx repo. https://github.com/georgiedekker/mlx_distributed_ring_inference_v2

maxims-eject commented Sep 14, 2025

Hi @awni,

I have been using the pipeline generate with mpi distributed (ring, two/three nodes, TB network) since may but starting from version "mlx-lm==0.27.0", I am seeing this crash (gpu timeout)

libc++abi: terminating due to uncaught exception of type std::runtime_error: [METAL] Command buffer execution failed: Caused GPU Timeout Error (00000002:kIOGPUCommandBufferCallbackErrorTimeout)

This is a very deterministic issue on my system in multiple scenarios (servers mpi, single text generation...), basically all the pipeline distributed generation seems to stop working from mlx-lm v0.27.0, while I have still no issue if I remain to previous versions, e.g. python -m pip install -U "mlx-lm==0.26.4". The nodes I am running are M4, M2 Pro and a mix of them.

For debugging, I was able to reproduce the issue using the pipeline_generate.py from the example directory and the following command - note I am using the smaller model due to hw ram sizes, but this used to be supported by pipeline_generate as part of the deepseek_v2 family:

clear && mlx.launch \
    --hostfile hosts_2.json \
    --backend ring \
    pipeline_generate.py \
    --prompt "What number is larger 6.9 or 6.11?" \
    --max-tokens 512 \
    --model mlx-community/DeepSeek-Coder-V2-Lite-Instruct-4bit

Below the two logs produced on my conda env with the old version of mlx-lm (up to 0.26.4) and the new versions (either 0.27.0 or 0.27.1):

python -m pip install -U "mlx-lm==0.26.4":

python -m pip install -U "mlx-lm==0.27.0":

Author

awni commented Sep 18, 2025

Fix is incoming here: ml-explore/mlx-lm#483

georgiedekker commented Nov 5, 2025

@awni made another attempt at speeding things up. seems to work please review, feel free to use any as example in the mlx repo. https://github.com/georgiedekker/mlx_distributed_ring_inference_v2

has anyone tried this code? Please let me know if it works for you as it does for me.

georgiedekker commented Nov 5, 2025

@awni what has to be done to make this pipeline sharding work on other model provider's models?

Author

awni commented Nov 5, 2025

Which models did you have in mind? If you file an issue in mlx-lm we can look into adding it.

georgiedekker commented Nov 5, 2025

Not one particular model, just would be great to know how to use this example you provided with any model from any provider. I'm currently running a small model mlx-community/Qwen3-1.7B-8bit on a single m4 mac mini with 16gb, would be great to test out many different models (from different providers) with simple cheap hardware.

awni/mlx_distributed_deepseek.md

Setup

Run

Tmoss11 commented Feb 11, 2025

Uh oh!

leozusa commented Mar 5, 2025

Uh oh!

awni commented Mar 5, 2025

Uh oh!

drixs2050 commented Mar 8, 2025

Uh oh!

fengyy0111 commented Mar 10, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

awni commented Mar 10, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

fengyy0111 commented Mar 11, 2025

Uh oh!

zengqingfu1442 commented Mar 17, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

awni commented Mar 17, 2025

Uh oh!

zengqingfu1442 commented Mar 18, 2025

Uh oh!

awni commented Mar 18, 2025

Uh oh!

awni commented Mar 18, 2025

Uh oh!

zengqingfu1442 commented Mar 18, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

awni commented Mar 18, 2025

Uh oh!

zengqingfu1442 commented Mar 18, 2025

Uh oh!

awni commented Mar 18, 2025

Uh oh!

jiyzhang commented Mar 31, 2025

Uh oh!

Basten7 commented Apr 11, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

zengqingfu1442 commented Apr 15, 2025

Uh oh!

RaylenFarnor commented May 5, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

georgiedekker commented Aug 8, 2025

Uh oh!

awni commented Aug 8, 2025

Uh oh!

georgiedekker commented Aug 9, 2025

Uh oh!

georgiedekker commented Aug 15, 2025

Uh oh!

maxims-eject commented Sep 14, 2025

Uh oh!

awni commented Sep 18, 2025

Uh oh!

georgiedekker commented Nov 5, 2025

Uh oh!

georgiedekker commented Nov 5, 2025

Uh oh!

awni commented Nov 5, 2025

Uh oh!

georgiedekker commented Nov 5, 2025

Uh oh!

fengyy0111 commented Mar 10, 2025 •

edited

Loading

awni commented Mar 10, 2025 •

edited

Loading

zengqingfu1442 commented Mar 17, 2025 •

edited

Loading

zengqingfu1442 commented Mar 18, 2025 •

edited

Loading

Basten7 commented Apr 11, 2025 •

edited

Loading

RaylenFarnor commented May 5, 2025 •

edited

Loading