# BLAS vs CUBLAS benchmark

**URL:** <https://discourse.julialang.org/t/blas-vs-cublas-benchmark/46205>\
**Category:** Performance\
**Tags:** question, blas, cuda\
**Created:** [September 7, 2020, 3:17pm UTC](https://discourse.julialang.org/t/blas-vs-cublas-benchmark/46205 "2020-09-07T15:17:44Z")\
**Posts on this page:** 1\
**Showing post:** 12

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [September 10, 2020, 5:21pm UTC](https://discourse.julialang.org/t/blas-vs-cublas-benchmark/46205/12 "2020-09-10T17:21:15Z")

</div>

I don’t suppose `rockblas_get_stream(handle())` the correct one to synchronize? I.e., that `hipStreamSynchronize(rockblas_get_stream(handle()))` would be correct?  
I’m also new to GPUs and don’t actually know what a stream is.

So for now, I used

```julia
gmul!(C,A,B) = (mul!(C,A,B); AMDGPU.HIP.hipDeviceSynchronize())

```

New results:

 ![gemmFloat64_100_4000_skylake-avx512_AVX512](https://global.discourse-cdn.com/julialang/original/3X/9/c/9cabbd12097c5180a422ed2c4ff6d67bea068e4e.png)  
`>10` TFLOPS is pretty good.

I’ll test your PR with build system updates.

---

_[View the full topic](https://discourse.julialang.org/t/blas-vs-cublas-benchmark/46205)._
