# Parallel computing and GPU support in neuralPDE.jl package

**URL:** <https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976>\
**Category:** New to Julia\
**Tags:** package, performance, neural-network\
**Created:** [October 15, 2023, 12:13pm UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976 "2023-10-15T12:13:08Z")\
**Posts on this page:** 10\
**Page:** 3

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [November 17, 2023, 11:36am UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976/41 "2023-11-17T11:36:08Z")

</div>

> [@Aakhash\_Sundaresan](#):
>
> ┌ Warning: Using `gpu` inside performance critical code will cause massive slowdowns due to type inference failure. Please update your code to use `gpu_device` API.  
> └ @ Lux C:\Users\Sunda.julia\packages\Lux\hlo4t\src\deprecated.jl:32  
> ┌ Warning: No functional GPU backend found! Defaulting to CPU.  
> │  
> │ 1. If no GPU is available, nothing needs to be done.  
> │ 2. If GPU is available, load the corresponding trigger package.  
> └ @ LuxDeviceUtils C:\Users\Sunda.julia\packages\LuxDeviceUtils\rMeCf\src\LuxDeviceUtils.jl:158
> 
> Here is the issue when I use the exact code in the documentation and run it on my NVIDIA GPU, and its massively slow, slower than CPU for the case of a example problem outlined in the documentaiton for “Training NeuralPDE on the GPU”.

This is not a question about NeuralPDE.jl. You will get better answers if you ask your questions better. If you instead ask for help with your CUDA installation, you will get everyone who knows CUDA, a lot larger setup than the number of people who know physics-informed neural networks with CUDA, to help you.

You have not given enough information to solve this question anyways. Did you `]add CUDA`? When you did that, what CUDA drivers do you have? What GPU do you have?

---

<div class="post-metadata">

**Author:** ![Aakhash\_Sundaresan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aakhash_sundaresan/32/202478_2.png) [@Aakhash\_Sundaresan](https://discourse.julialang.org/u/Aakhash_Sundaresan)\
**Post date:** [November 19, 2023, 11:56am UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976/42 "2023-11-19T11:56:56Z")

</div>

Hi @ChrisRackauckas, extremely sorry for posting the irrelevant question here…apologies…I’m having the exact same problem as outlined in [NeuralPDE features and GPU compatibility](https://discourse.julialang.org/t/neuralpde-features-and-gpu-compatibility/104870). No wthat I have shifted to Flux and the problem seems to be solved… However, I ran a matmul smoke test with x=rand(100000,100000), y = rand(100000,100000) and I got 11.2 s of computation time on cpu and “Out of Memory error” on the GPU (Nvidia RTX 3090 24 GB VRAM). With the matrix size reduced by an order the time taken with CPU was lowest when compared to GPU. So that means GPU doesn’t perform well right?. However I have 9 neural networks to predict the 9 variables in my governing equation. I’m unable to imrpove the performance of the NeuralPDE for my specific problem. Is it possible to have a single neural network with 9 output neurons, instead of creating separate networks for learning each of the variables in the neuralPDE framework? As adviced by you, I had visited the Chromatography repo…but, the problem is the formulation of their problem is entirely different from mine and I’m not knowing how to adapt to my specific case…What to do? Please help me out/…Thanks a lot in advance Chris…

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [November 19, 2023, 12:32pm UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976/43 "2023-11-19T12:32:49Z")

</div>

> [@Aakhash\_Sundaresan](#):
>
> However, I ran a matmul smoke test with x=rand(100000,100000), y = rand(100000,100000) and I got 11.2 s of computation time on cpu and “Out of Memory error” on the GPU (Nvidia RTX 3090 24 GB VRAM).

Yes, that is too big to fit onto any GPU that exists today.

> [@Aakhash\_Sundaresan](#):
>
> With the matrix size reduced by an order the time taken with CPU was lowest when compared to GPU. So that means GPU doesn’t perform well right?

If your matrices are actually that size. Are you using neural networks with layers of size `100000`? The documentation doesn’t.

---

<div class="post-metadata">

**Author:** ![Aakhash\_Sundaresan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aakhash_sundaresan/32/202478_2.png) [@Aakhash\_Sundaresan](https://discourse.julialang.org/u/Aakhash_Sundaresan)\
**Post date:** [November 19, 2023, 11:13pm UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976/44 "2023-11-19T23:13:32Z")

</div>

Nope, I’m using 9 neural networks with 3 Hidden layers and 40 neurons in each layer. So how would I do the matmul smoke test for this case? What is the size of MxN that I have to use and how do you calculate that ?  
What are the factors that affect the performance of the NeuralPDE framework? DOes it have any limit on the number of equations and BC’s that can be used, or the maximum number of layers and neurons that one can use? Or the maximum order of the differential equation that can be used?..

Also, is it that GridTraining is faster on the GPU’s and QuasiRandom sampling isn’t? On the documentation page, it says to use GridTraining only for testing purposes. That means GridTraining strategu cannot be used for final training or what is the exact reason @ChrisRackauckas

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [November 21, 2023, 10:42am UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976/45 "2023-11-21T10:42:24Z")

</div>

> [@Aakhash\_Sundaresan](#):
>
> ’m using 9 neural networks with 3 Hidden layers and 40 neurons in each layer. So how would I do the matmul smoke test for this case?

Test 40xN matmuls. N would be the batch size in the sampling, or number of samples. So like, 40x100.

> [@Aakhash\_Sundaresan](#):
>
> What are the factors that affect the performance of the NeuralPDE framework? DOes it have any limit on the number of equations and BC’s that can be used, or the maximum number of layers and neurons that one can use? Or the maximum order of the differential equation that can be used?..

There is no maximum order, though physical equations don’t tend to have above 4 and I think that’s where most of the hard-coded extra optimizations would stop. There is not a limit on the number of equations, and you generally scale linearly with that. You just scale quadratically with larger size (or think cubically as the samples grow), which is a fundamental limitation because of the matmuls in the neural networks.

> [@Aakhash\_Sundaresan](#):
>
> Also, is it that GridTraining is faster on the GPU’s and QuasiRandom sampling isn’t? On the documentation page, it says to use GridTraining only for testing purposes. That means GridTraining strategu cannot be used for final training or what is the exact reason

GridTraining does not overcome curse of dimensionality, does not hit random points so it tends to not give great results between grid points, and has slower convergence than a quasi-random low discrepancy sampler.

---

<div class="post-metadata">

**Author:** ![Aakhash\_Sundaresan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aakhash_sundaresan/32/202478_2.png) [@Aakhash\_Sundaresan](https://discourse.julialang.org/u/Aakhash_Sundaresan)\
**Post date:** [November 21, 2023, 12:46pm UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976/46 "2023-11-21T12:46:42Z")

</div>

Thanks for your guidance, but I had tested the matmuls on the GPU both using Flux and Lux, they seem to perform better than the CPU case, but when it comes to the NeuralPDE framework, its just very slow for some reason and I’m not able to figure out the problem…I’m just using 40 Neurons 3 Hidden layers, QuasiRandomTraining with 1000 points and LatinHypercubesampling() stratgey with LBFGS for the RANS equations with energy (so total of 4 equations) and 9 boundary conditions. I have 9 variables to be predicted, so there are 9 neural networks. @ChrisRackauckas Now you can maybe calculate the size of the parameters and estimate the performance now…Its like suddenly the iterations are faster and then it hangs for about maybe say 5 minutes or so…Is this due to the stiffness in the Loss landscape?.. If it is so, how can i overcome this issue?

---

<div class="post-metadata">

**Author:** ![Aakhash\_Sundaresan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aakhash_sundaresan/32/202478_2.png) [@Aakhash\_Sundaresan](https://discourse.julialang.org/u/Aakhash_Sundaresan)\
**Post date:** [November 22, 2023, 4:12am UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976/47 "2023-11-22T04:12:50Z")

</div>

@ChrisRackauckas Can you comment on the above issue please?

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [November 22, 2023, 6:36am UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976/48 "2023-11-22T06:36:45Z")

</div>

That’s just due to line search failures. Indeed the stiffness gives PINNs a problem, so for RANS equations PINNs aren’t great.

---

<div class="post-metadata">

**Author:** ![Aakhash\_Sundaresan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aakhash_sundaresan/32/202478_2.png) [@Aakhash\_Sundaresan](https://discourse.julialang.org/u/Aakhash_Sundaresan)\
**Post date:** [November 22, 2023, 6:40am UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976/49 "2023-11-22T06:40:59Z")

</div>

@ChrisRackauckas So what’s the alternative to doing this when I have the experimental data that is not close to the wall and I want to accurately extrapolate to the wall using ML ?

---

<div class="post-metadata">

**Author:** ![Aakhash\_Sundaresan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aakhash_sundaresan/32/202478_2.png) [@Aakhash\_Sundaresan](https://discourse.julialang.org/u/Aakhash_Sundaresan)\
**Post date:** [November 22, 2023, 11:39am UTC](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976/50 "2023-11-22T11:39:30Z")

</div>

@ChrisRackauckas Can you please comment on the above please?

[Previous page](https://discourse.julialang.org/t/parallel-computing-and-gpu-support-in-neuralpde-jl-package/104976.md?page=2)
