# Apple M4 Max AMX Linear Algebra performance versus CPU and GPU

**URL:** https://discourse.julialang.org/t/apple-m4-max-amx-linear-algebra-performance-versus-cpu-and-gpu/132537
**Category:** Performance
**Tags:** linearalgebra, apple
**Created:** [September 21, 2025, 1:41pm UTC](https://discourse.julialang.org/t/apple-m4-max-amx-linear-algebra-performance-versus-cpu-and-gpu/132537 "2025-09-21T13:41:54Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![PetarM](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petarm/32/218700_2.png) [@PetarM](https://discourse.julialang.org/u/PetarM)
#### Post date: [September 21, 2025, 1:41pm UTC](https://discourse.julialang.org/t/apple-m4-max-amx-linear-algebra-performance-versus-cpu-and-gpu/132537/1 "2025-09-21T13:41:55Z")

</div>

Hello everybody!

I made a video investigation of Apple’s AMX accelerator Linear Algebra performance in Julia: [https://www.youtube.com/watch?v=TjfA9LVgHXk](https://www.youtube.com/watch?v=TjfA9LVgHXk)

According to my findings, the 2 AMX cores achieve almost 3 times the peak performance of the 12 P-cores in FP32 and close to same performance as the Ryzen 9950X in dense matrix-matrix multiplication. At the same time they are 10 times more power efficient than the P-cores and about 6 times more power efficient than the Zen 5 cores.

Somewhat disappointingly, the AMX cores have the same throughput in FP16 as in FP32, even though according to Apple’s patents it seemed reasonable to expect a 4-fold increase

In FP64, the AMX core throughput drops to 1/4 that of FP32, as expected. However, this is still superior to the 12 P-core performance.

The other problem I examined was matrix-vector multiplication and there the performance was memory bound and roughly matched that of the P-cores.

When comparing against the GPU, it becomes quite apparent why Apple introduced the AMX: it uplifts performance exactly where the GPU is weak: small problem sizes and it does so with significantly higher power efficiency.

Curiously, I achieved better peak GPU performance in Julia than using Apple’s MLX! Kudos to everyone who made it possible to so easily leverage the GPU in Julia.

 ![AMX vs GPU](https://global.discourse-cdn.com/julialang/original/3X/e/a/eab8744ffd3691f6ce0ccbb7a0c285486f8c0d11.jpeg)

 ![Matrix-vector](https://global.discourse-cdn.com/julialang/original/3X/8/f/8f30c7194679fda2218bb73da0d3c29397ffc95a.jpeg)

---

<div class="post-metadata">

### Author: ![rveltz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rveltz/32/2707_2.png) [@rveltz](https://discourse.julialang.org/u/rveltz)
#### Post date: [September 21, 2025, 9:32pm UTC](https://discourse.julialang.org/t/apple-m4-max-amx-linear-algebra-performance-versus-cpu-and-gpu/132537/2 "2025-09-21T21:32:05Z")

</div>

Very nice!

Is the code available?

---

<div class="post-metadata">

### Author: ![Norman](https://avatars.discourse-cdn.com/v4/letter/n/97f17d/32.png) [@Norman](https://discourse.julialang.org/u/Norman)
#### Post date: [September 22, 2025, 12:17am UTC](https://discourse.julialang.org/t/apple-m4-max-amx-linear-algebra-performance-versus-cpu-and-gpu/132537/3 "2025-09-22T00:17:46Z")

</div>

Great video! Just bumped into your video on Youtube and then saw the post here.
