# \[YouTube/GitHub\] What is the FASTEST Computer Language? 45 Languages Tested

**URL:** <https://discourse.julialang.org/t/youtube-github-what-is-the-fastest-computer-language-45-languages-tested/64506>\
**Category:** Performance\
**Created:** [July 12, 2021, 1:37pm UTC](https://discourse.julialang.org/t/youtube-github-what-is-the-fastest-computer-language-45-languages-tested/64506 "2021-07-12T13:37:34Z")\
**Posts on this page:** 1\
**Showing post:** 10

<div class="post-metadata">

**Author:** ![louie4825](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/louie4825/32/27359_2.png) [@louie4825](https://discourse.julialang.org/u/louie4825)\
**Post date:** [August 2, 2021, 5:53am UTC](https://discourse.julialang.org/t/youtube-github-what-is-the-fastest-computer-language-45-languages-tested/64506/10 "2021-08-02T05:53:20Z")

</div>

> [@Elrod](#):
>
> Given that the actual amount of computation is very small, my guess is that not being SIMD is fvaster, and that `clear_factors!` isn’t SIMD.  
> But I’d have to look at the code, and having something I can copy/paste would make that a lot easier for me.

I didn’t actually expect `LoopVectorization` to make the code any faster since it isn’t very conducive to SIMD, as you said. It has a large stride between array elements. Here’s the code I used to test it, similar to what I [posted before](https://discourse.julialang.org/t/why-is-this-simd-loop-faster-than-a-while-loop-even-if-it-has-longer-assembly/65585):

```julia
using BenchmarkTools
using LoopVectorization

# Auxillary functions
begin
const _uint_bit_length = sizeof(UInt) * 8
const _div_uint_size_shift = Int(log2(_uint_bit_length))
@inline _mul2(i::Integer) = i << 1
@inline _div2(i::Integer) = i >> 1
@inline _map_to_index(i::Integer) = _div2(i - 1)
@inline _map_to_factor(i::Integer) = _mul2(i) + 1
@inline _mod_uint_size(i::Integer) = i & (_uint_bit_length - 1)
@inline _div_uint_size(i::Integer) = i >> _div_uint_size_shift
@inline _get_chunk_index(i::Integer) = _div_uint_size(i + (_uint_bit_length - 1))
@inline _get_bit_index_mask(i::Integer) = UInt(1) << _mod_uint_size(i - 1)
end

# Main code
function clear_factors_while!(arr::Vector{UInt}, factor_index::Integer, max_index::Integer)
    factor = _map_to_factor(factor_index)
    index = _div2(factor * factor)
    while index <= max_index
        @inbounds arr[_get_chunk_index(index)] |= _get_bit_index_mask(index)
        index += factor
    end
    return arr
end

function clear_factors_simd!(arr::Vector{UInt}, factor_index::Integer, max_index::Integer)
    factor = _map_to_factor(factor_index)
    @simd for index in _div2(factor * factor):factor:max_index
        @inbounds arr[_get_chunk_index(index)] |= _get_bit_index_mask(index)
    end
    return arr
end

function clear_factors_turbo!(arr::Vector{UInt}, factor_index::Integer, max_index::Integer)
    factor = _map_to_factor(factor_index)
    factor < _uint_bit_length && error("Factor must be greater than UInt bit length to avoid memory dependencies")
    start_index = _div2(factor * factor)
    iterations = cld(max_index - start_index, factor) - 1
    @turbo for i in 0:iterations
        index = start_index + (factor * i)
        @inbounds arr[(index + 63) >> _div_uint_size_shift] |= 1 << ((index - 1) & 63)
    end
    return arr
end

println(
    clear_factors_while!(zeros(UInt, cld(500_000, _uint_bit_length)), 202, 500_000) ==
    clear_factors_simd!(zeros(UInt, cld(500_000, _uint_bit_length)), 202, 500_000) ==
    clear_factors_turbo!(zeros(UInt, cld(500_000, _uint_bit_length)), 202, 500_000) ==
)

const x = zeros(UInt, cld(500_000, sizeof(UInt) * 8))
@benchmark clear_factors_while!(x, 202, 500_000)
@benchmark clear_factors_simd!(x, 202, 500_000)
@benchmark clear_factors_turbo!(x, 202, 500_000)

```

---

_[View the full topic](https://discourse.julialang.org/t/youtube-github-what-is-the-fastest-computer-language-45-languages-tested/64506)._
