# GPU compute & high precision general questions

**URL:** <https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527>\
**Category:** New to Julia\
**Tags:** gpu, cuda, opencl\
**Created:** [May 20, 2021, 4:06pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527 "2021-05-20T16:06:54Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![HexSpin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexspin/32/19319_2.png) [@HexSpin](https://discourse.julialang.org/u/HexSpin)\
**Post date:** [May 20, 2021, 4:06pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/1 "2021-05-20T16:06:54Z")

</div>

Yesterday, for my birthday, my wife took me to MicroCenter and gave me her blessing to build a replacement for my 2010 Mac Pro. Soon I had a cart loaded with a 16 core Ryzen 9, case, memory …and in a fit of irrational optimism, I tossed in a 1200 watt power supply and a motherboard that could easily accomodate 2 large graphics cards.  
I’ve awakened with a MicroCenter hangover … I’ve never programmed a GPU before! What was I thinking? I can’t even buy a GPU because of the shortage! Assuming that one day GPUs return to the marketplace, like toilette paper did this year, I wonder if you might offer advice for someone interested in GPU compute…  
— My computation is primarily FP64. AMD appears to offer more cores for that sort of computation over NVIDIA. However, CUDA seems like the more useful language - and the Julia CUDA tool appears better developed. Thoughts on OpenCL vs CUDA?  
— I also need 128 bit or better precision. Is it naive to think I might use GPUs for such work?  
— I’m also leaving OSX … I grabbed a copy of Windows 10 … but I recall(?) reading a thread suggesting Linux would be better in GPU applications…?

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [May 20, 2021, 4:24pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/2 "2021-05-20T16:24:41Z")

</div>

Most consumer-grade GPUs will have a limited number of `Float64` compute units, and I wouldn’t be surprised if your Ryzen 9 outstrips the `Float64` performance of any reasonably-priced GPU.

To my knowledge, there’s only been one effort to provide extended-precision support for GPUs, and it’s experimental and may be out of date: [https://github.com/lumianph/gpuprec](https://github.com/lumianph/gpuprec)

If you’ve never done GPU computing before, Nvidia’s ecosystem has much better documentation and will be easier to pick up, but given your goals, I’m not sure that a consumer-grade GPU will do much for you.

Regarding OS choice, Linux almost always has the most robust developer tooling support, but is not the most ergonomic for day-to-day usability. I’ve been happy running Windows 10 with [Windows Subsystem for Linux](https://docs.microsoft.com/en-us/windows/wsl/install-win10), which lets you easily switch between the two operating systems as needed.

---

<div class="post-metadata">

**Author:** ![HexSpin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexspin/32/19319_2.png) [@HexSpin](https://discourse.julialang.org/u/HexSpin)\
**Post date:** [May 20, 2021, 5:08pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/3 "2021-05-20T17:08:35Z")

</div>

Thank you. I agree about the Ryzen, that’s why I decided on a new machine rather than just an external GPU (and my dear Apple makes the Pro harder to run every year). I offer for your consideration the Radeon Pro VII being introduced in June with 6.5 TFLOPS in FP64 … I’m not sure how translates to my application, but it sure seems promising. At $2K, it’s not exactly reasonably priced; however, the older Radeon VII, $700, if to be found, has FP64 of 3.5 TFLOPS. I’d like to find at least one of those…

---

<div class="post-metadata">

**Author:** ![HexSpin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexspin/32/19319_2.png) [@HexSpin](https://discourse.julialang.org/u/HexSpin)\
**Post date:** [May 20, 2021, 5:16pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/4 "2021-05-20T17:16:05Z")

</div>

Oh - and Windows Subsystem for Linux looks perfect!

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [May 20, 2021, 5:27pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/5 "2021-05-20T17:27:01Z")

</div>

You may be able to find a used Nvidia Tesla K80 for $300-400. It won’t deliver cutting-edge performance (~1.9 TFLOPS F64), but CUDA has better learning materials, and the cost is more reasonable for a component where you’re not yet sure whether it’ll be useful for your workloads.

---

<div class="post-metadata">

**Author:** ![HexSpin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexspin/32/19319_2.png) [@HexSpin](https://discourse.julialang.org/u/HexSpin)\
**Post date:** [May 21, 2021, 2:29pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/6 "2021-05-21T14:29:00Z")

</div>

FYI - K80s can be found on Amazon for $299, which is great, but apparently they have no love for non-workstation motherboards and have other cooling and compatibility issues. I’ll just get a basic NVIDIA that’s available for proof of concept.

---

<div class="post-metadata">

**Author:** ![maxfreu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maxfreu/32/17468_2.png) [@maxfreu](https://discourse.julialang.org/u/maxfreu)\
**Post date:** [May 21, 2021, 3:02pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/7 "2021-05-21T15:02:06Z")

</div>

Nvidia just [announced](https://blogs.nvidia.com/blog/2021/05/18/lhr/?ncid=afm-chs-44270&ranMID=44270&ranEAID=kXQk6*ivFEQ&ranSiteID=kXQk6.ivFEQ-JdchVbxuILMfToYqsNZ3zA) to release consumer GPUs with „low hashrate“ where they block crypto mining. So hopefully prices will drop soon and supply gets better.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [May 21, 2021, 3:16pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/8 "2021-05-21T15:16:30Z")

</div>

> I also need 128 bit or better precision. Is it naive to think I might use GPUs for such work?

Yes I think so. Despite seeing (I think this is misleading/misunderstanding): [https://www.nvidia.com/en-us/geforce/forums/off-topic/25/89829/128-bit-supercomputing-using-gpus/](https://www.nvidia.com/en-us/geforce/forums/off-topic/25/89829/128-bit-supercomputing-using-gpus/)

> Eventually, Nvidia’s research into 128-bit supercomputing will trickle back down to PCs in the form of real-time ray-tracing and 128-bit computing on the personal computer.

64-bit is clearly second-class for CUDA GPUs, and more so over time, with 16-bit 128/32=4 times faster than 6-bit in CUDA 7.0, but 256/2= 128x faster in CUDA 8.6 (see table 3).

I.e 16-bit floats have best performance (256 units of Throughput).

[https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html)

> Global memory instructions support reading or writing words of size equal to 1, 2, 4, 8, or 16 bytes.

16 bytes is 128-bit, but it probably refers to load/store pair of 64-bit, as in ARMv8.

> Sub-byte WMMA operations provide a way to access the low-precision capabilities of Tensor Cores. They are considered a preview feature i.e. the data structures and APIs for them are subject to change and may not be compatible with future releases

No GPU or other hardware, that I know of (except for a few CPUs) support 128-bit floats (and no hardware supports 256-bit floats):

> **[Quadruple-precision floating-point format](https://en.wikipedia.org/wiki/Quadruple-precision_floating-point_format)**
>
> In computing, quadruple precision (or quad precision) is a binary floating-point–based computer number format that occupies 16 bytes (128 bits) with precision at least twice the 53-bit double precision.
> This 128-bit quadruple precision is designed not only for applications requiring results in higher than double precision, but also, as a primary function, to allow the computation of double precision results more reliably and accurately by minimising overflow and round-off errors in intermediate ...

> Quadruple-precision (128-bit) hardware implementation should not be confused with “128-bit FPUs” that implement [SIMD](https://en.wikipedia.org/wiki/SIMD) instructions

The trend for GPUs and CPUs is to smaller:

> **[IBM Reveals Next-Generation IBM POWER10 Processor](https://newsroom.ibm.com/2020-08-17-IBM-Reveals-Next-Generation-IBM-POWER10-Processor)**
>
> IBM (NYSE: IBM) today revealed the next generation of its IBM POWER central processing unit (CPU) family: IBM POWER10. Designed to offer a platform to meet the unique needs of enterprise hybrid...

> faster AI inference for FP32, BFloat16 and INT8

I’ve misplaced their long documentation, but they actually defined 4-bit matrix multiply too (I believe a first for a CPU). And the also add 128-bit integers (e.g. compare, also a first I think, i.e. not just SIMD registers), and are now variable-length (the 32-bit-length, plus now 64-bit instructions), not like RISC applied too well before (or any modern chip).

Intriguingly 36-bit computers, at least Honeywell from the 1960s, supported more _precision_ (and if I recall also larger integers, i.e. 72-bit) than e.g. ARM CPUs and it seems even x86’s (by now lesser performant) 80-bit float format, and all GPUs (Honeywell’s double add instruction takes twice the time of its single-precision add):

> **[GE-600 series](https://en.wikipedia.org/wiki/GE-600_series)**
>
> The GE-600 series was a family of 36-bit mainframe computers originating in the 1960s, built by General Electric (GE). When GE left the mainframe business the line was sold to Honeywell, which built similar systems into the 1990s as the division moved to Groupe Bull and then NEC.
> The system is perhaps best known as the hardware used by the Dartmouth Time Sharing System (DTSS) and the Multics operating system. Multics was supported by virtual memory additions made in the GE 645.
> The 600-series C...

> supported [floating point](https://en.wikipedia.org/wiki/Floating_point) in both 36-bit single-precision and 2 x 36-bit double precision, the exponent being stored separately, allowing up to 71 bits of precision (one bit being used for the sign).

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [May 21, 2021, 4:01pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/9 "2021-05-21T16:01:10Z")

</div>

I believe MultiFloats.jl should work on GPUs.

---

<div class="post-metadata">

**Author:** ![HexSpin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexspin/32/19319_2.png) [@HexSpin](https://discourse.julialang.org/u/HexSpin)\
**Post date:** [May 21, 2021, 4:14pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/10 "2021-05-21T16:14:35Z")

</div>

Thank you for your thoughtful response. I guess I’ll just have to rely on my CPU for the higher precision. As an aside about CPUs, I sure wish 80 bit precision were allowed in Julia.

---

<div class="post-metadata">

**Author:** ![HexSpin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexspin/32/19319_2.png) [@HexSpin](https://discourse.julialang.org/u/HexSpin)\
**Post date:** [May 21, 2021, 4:19pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/11 "2021-05-21T16:19:16Z")

</div>

Interesting… though : [https://github.com/dzhang314/MultiFloats.jl/issues/23](https://github.com/dzhang314/MultiFloats.jl/issues/23)  
Perhaps Mr. Zhang will hear my plea…

---

<div class="post-metadata">

**Author:** ![HexSpin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexspin/32/19319_2.png) [@HexSpin](https://discourse.julialang.org/u/HexSpin)\
**Post date:** [May 21, 2021, 4:24pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/12 "2021-05-21T16:24:47Z")

</div>

And thanks for the heads up about MultiFloats.jl… I’ve been happily using BigFloat, which is said to be 56x slower. Even if I get a fraction of this speed up, I’ll be thrilled!

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [May 21, 2021, 4:25pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/13 "2021-05-21T16:25:20Z")

</div>

Probably (since GPUs are also Turing-complete, all [Julia] code should work _in theory_?).

While claimed “fast”: [GitHub - dzhang314/MultiFloats.jl: Fast extended-precision floating-point arithmetic for Julia](https://github.com/dzhang314/MultiFloats.jl)

> At 100-bit precision, **MultiFloats.jl** is roughly **40x faster than [`BigFloat`](https://docs.julialang.org/en/v1/manual/integers-and-floating-point-numbers/#Arbitrary-Precision-Arithmetic)** and **2x faster than [DoubleFloats.jl](https://github.com/JuliaMath/DoubleFloats.jl)**.

These claims are for CPUs, and I very much doubt you get the same speedup on GPUs. Besides, do Julia’s BigFloats (or any other implementation) work on GPUs? Seems not according to the closed issue link below (could only in theory work).

I would look at all the competing options, possibly even posits, for GPUs, to see which package is fastest (and without rounding issue) if you can’t do away with lesser precision. See e.g. here:

> <https://github.com/JuliaGPU/GPUArrays.jl/issues/189>
>
> I would like advice on how to proceed with an enhancement to Julia and GPU array…s. I'd like to figure out how to speed up (if possible) BigFloat performance - perhaps at first for a subset of BigFloat like Float128. My first thought is to try rewriting MPFR routines directly in Julia to avoid the ccall overhead. It would be even better if the Julia implementation would then also "just work" with GPUArrays. Any advice on whether this is a worthwhile endeavor would be greatly appreciated.

---

<div class="post-metadata">

**Author:** ![HexSpin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexspin/32/19319_2.png) [@HexSpin](https://discourse.julialang.org/u/HexSpin)\
**Post date:** [May 21, 2021, 4:32pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/14 "2021-05-21T16:32:01Z")

</div>

Agreed. To date all of my work is CPU, and my week long runs would be much more fun if they were hours long. I’ll take your advice about the various options. It’s a very obvious and simple optimization that I just didn’t think to investigate. If I could do that a couple times, maybe the GPU effort would become moot.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [May 21, 2021, 4:40pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/15 "2021-05-21T16:40:14Z")

</div>

Note, even with GPUs not having the 128-bit or higher data types, it’s interesting what’s already done:

> **[Mixed-Precision Programming with CUDA 8 | NVIDIA Technical Blog](https://developer.nvidia.com/blog/mixed-precision-programming-cuda-8/)**
>
> Update, March 25, 2019: The latest Volta and Turing GPUs now incoporate Tensor Cores, which accelerate certain types of FP16 matrix math. This enables faster and easier mixed-precision computation…

> Many technical and HPC applications require high precision computation with 32-bit (single float, or FP32) or 64-bit (double float, or FP64) floating point, and [there are even GPU-accelerated applications that rely on even higher precision](http://link.springer.com/article/10.1007%2Fs10915-012-9679-3) (128- or 256-bit floating point!). But there are many applications for which much lower precision arithmetic suffices.

> **[High-Precision Numerical Simulations on a CUDA GPU: Kerr Black Hole Tails -...](https://link.springer.com/article/10.1007/s10915-012-9679-3)**
>
> Computational science has advanced significantly over the past decade and has impacted almost every area of science and engineering. Most numerical scientific computation today is performed with double-precision floating-point accuracy (64-bit or...

> Kerr black hole space-time. More specifically, this application involves a hyperbolic partial-differential-equation solver that uses high-order finite-differencing and quadruple (128-bit or ∼30 decimal digits) or octal (256-bit or ∼60 decimal digits) floating-point precision. Given the computational demands of this high-order and high-precision solver, in addition to the rather long evolutions required for these studies, we accelerate the solver using a many-core Nvidia graphics-processing-unit and obtain an order-of-magnitude speed-up over a high-end multi-core processor.

So it’s at least possible and faster (in some cases, we don’ know how much they tried to optimize the CPU code, or if recent large CPUs could compete better). I don’t know what you’re doing, if “differential-equation solver”, I would look into reusing that one, or whatever is out there or possibly lower-level libraries they may build on. See what’s out there, that’s probably not already wrapped by Julia, while should be possible. It’s likely faster than Julia code made and optimized for CPUs.

---

<div class="post-metadata">

**Author:** ![HexSpin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hexspin/32/19319_2.png) [@HexSpin](https://discourse.julialang.org/u/HexSpin)\
**Post date:** [May 21, 2021, 5:14pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/16 "2021-05-21T17:14:44Z")

</div>

Thank you …looks like I have some reading to do. The math I do is pretty simple - adding sequential fractions raised to small powers. So, along the lines of,  
1/(10^13+0)^0.83+ 1/(10^13+1)^0.83+1/(10^13+2)^0.83+… for real numbers and complex. Brute force math, not elegant differentials.

---

<div class="post-metadata">

**Author:** ![rguerrab](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rguerrab/32/22134_2.png) [@rguerrab](https://discourse.julialang.org/u/rguerrab)\
**Post date:** [December 29, 2021, 8:06pm UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/17 "2021-12-29T20:06:05Z")

</div>

Sorry to necropost like this, but if Multifloats.jl works well on GPUs as is, would it not make sense to copy-paste their Float64xN types to make Float32xN that make up for the egregious crippling of FP64 performance in commercial GPUs? Of course, this won’t get past the lack of ECC (which AMD is surprisingly good about) whenever that may be necessary.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [December 30, 2021, 12:22am UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/18 "2021-12-30T00:22:13Z")

</div>

if you want more than 60 bits of precision you still probably want to use float64 based multifloats since multiplication is quadratic in number of elements.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [December 30, 2021, 12:26am UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/19 "2021-12-30T00:26:12Z")

</div>

also float32 gives you a hard floor of about 2^-150 which makes them less than great when chained together.

---

<div class="post-metadata">

**Author:** ![rguerrab](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rguerrab/32/22134_2.png) [@rguerrab](https://discourse.julialang.org/u/rguerrab)\
**Post date:** [December 30, 2021, 2:11am UTC](https://discourse.julialang.org/t/gpu-compute-high-precision-general-questions/61527/20 "2021-12-30T02:11:47Z")

</div>

Sure, but float64 FLOPS on any of Nvidia’s consumer or even workstation cards are 1/64 or 1/32x that of float32 operations. So, unless everyone who wants to do FP64 computing jumps over to Radeon Pro VII’s or forks over $5-10K for a Tesla card, access to a simple single-single GPU library could help people who can’t afford supercomputer hardware.
