# knucleotide benchmark improvement for Julia and hashing

**URL:** <https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286>\
**Category:** Community\
**Tags:** announcement\
**Created:** [January 30, 2019, 5:47pm UTC](https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286 "2019-01-30T17:47:37Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![jean-pierre\_both](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jean-pierre_both/32/12636_2.png) [@jean-pierre\_both](https://discourse.julialang.org/u/jean-pierre_both)\
**Post date:** [January 30, 2019, 5:47pm UTC](https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286/1 "2019-01-30T17:47:37Z")

</div>

Just to announce a net improvement in the knucleotide benchmark described in :

[[k-nucleotide (Benchmarks Game)](https://benchmarksgame-team.pages.debian.net/benchmarksgame/performance/knucleotide.html)]

The julia#2 solution is found at:  
[[https://benchmarksgame-team.pages.debian.net/benchmarksgame/program/knucleotide-julia-2.html](https://benchmarksgame-team.pages.debian.net/benchmarksgame/program/knucleotide-julia-2.html)]

The improvement is at :  
[https://gitlab.com/jpboth/knucleotide-benchmark](https://gitlab.com/jpboth/knucleotide-benchmark)

The results shows an improvement by factor 5 by encoding bases and so hashing Int64 instead of String in Dict.

One question is that by defining a hash function doing nothing as described Base.hash documentation I did not get further improvement.

---

<div class="post-metadata">

**Author:** ![datnamer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datnamer/32/3471_2.png) [@datnamer](https://discourse.julialang.org/u/datnamer)\
**Post date:** [January 30, 2019, 6:00pm UTC](https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286/2 "2019-01-30T18:00:51Z")

</div>

That’s a great improvement. Why is Julia’s speed only roughly equiv to python 3, with a more complex implementation? I guess string and dict ops aren’t as optimized as python’s c implementation.

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [January 30, 2019, 6:03pm UTC](https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286/3 "2019-01-30T18:03:08Z")

</div>

> **[GitHub - JuliaPerf/BenchmarksGame.jl](https://github.com/JuliaPerf/BenchmarksGame.jl)**
>
> Contribute to JuliaPerf/BenchmarksGame.jl development by creating an account on GitHub.

Has a 15x improvement implementation.

---

<div class="post-metadata">

**Author:** ![viralbshah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/viralbshah/32/54_2.png) [@viralbshah](https://discourse.julialang.org/u/viralbshah)\
**Post date:** [January 31, 2019, 6:26am UTC](https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286/4 "2019-01-31T06:26:29Z")

</div>

Have you already submitted these upstream? Or should we request others to help submit the benchmarks upstream?

---

<div class="post-metadata">

**Author:** ![Diego\_Javier\_Zea](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/diego_javier_zea/32/1858_2.png) [@Diego\_Javier\_Zea](https://discourse.julialang.org/u/Diego_Javier_Zea)\
**Post date:** [January 31, 2019, 7:40am UTC](https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286/5 "2019-01-31T07:40:41Z")

</div>

The Python 3 and Python 3 # 3 versions are running each sequence in parallel in a quad core machine.

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [January 31, 2019, 7:48am UTC](https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286/6 "2019-01-31T07:48:03Z")

</div>

I have started with one or two of them.

---

<div class="post-metadata">

**Author:** ![jean-pierre\_both](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jean-pierre_both/32/12636_2.png) [@jean-pierre\_both](https://discourse.julialang.org/u/jean-pierre_both)\
**Post date:** [February 1, 2019, 8:15pm UTC](https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286/7 "2019-02-01T20:15:33Z")

</div>

I submitted the base encoded implementation, still waiting…

If you can help, it will be fine.

---

<div class="post-metadata">

**Author:** ![jean-pierre\_both](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jean-pierre_both/32/12636_2.png) [@jean-pierre\_both](https://discourse.julialang.org/u/jean-pierre_both)\
**Post date:** [February 1, 2019, 8:15pm UTC](https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286/8 "2019-02-01T20:15:39Z")

</div>

Rust (great language too) people use in their implementation use Indexmap (instead of HashMap) which is inspired  
by python 3.6 dictionnary, which they say is efficient and preserve insertion order see their doc:

[[GitHub - bluss/indexmap: A hash table with consistent order and fast iteration; access items by key or sequence index](https://github.com/bluss/indexmap)]([GitHub - bluss/indexmap: A hash table with consistent order and fast iteration; access items by key or sequence index](https://github.com/bluss/indexmap)]

As in their implementation I tried use a very simple hash for use in Dict (even identity function)  
but no success, I did not understand why.

---

<div class="post-metadata">

**Author:** ![jean-pierre\_both](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jean-pierre_both/32/12636_2.png) [@jean-pierre\_both](https://discourse.julialang.org/u/jean-pierre_both)\
**Post date:** [February 1, 2019, 9:19pm UTC](https://discourse.julialang.org/t/knucleotide-benchmark-improvement-for-julia-and-hashing/20286/9 "2019-02-01T21:19:41Z")

</div>

In fact i had difficulty submitting to the benchmark site with the first threaded version and the second paused a problem beccause of the Match package so some help would be welcome.

For knucleotide, Rust people used indexmap which is inspired by python 3.6 dict. Does julia has plans for sthing like that?

Julia is very fine language, thanks
