# Funny Benchmark with Julia (no longer) at the bottom

**URL:** <https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611>\
**Category:** Performance\
**Tags:** benchmark\
**Created:** [October 5, 2023, 1:12pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611 "2023-10-05T13:12:30Z")\
**Posts on this page:** 20\
**Page:** 5

<div class="post-metadata">

**Author:** ![Syx\_Pek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/syx_pek/32/6364_2.png) [@Syx\_Pek](https://discourse.julialang.org/u/Syx_Pek)\
**Post date:** [October 11, 2023, 5:10am UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/81 "2023-10-11T05:10:59Z")

</div>

My last (faster) algorithm completely changes how the method works. As such I don’t know if it will satisy the constraints as is.

I’m also unsure if my method is faster in their environment ☹

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 11, 2023, 5:25am UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/82 "2023-10-11T05:25:09Z")

</div>

I think other languages implementation should pick up your sorting algorithms soon

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 11, 2023, 11:28am UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/83 "2023-10-11T11:28:07Z")

</div>

If someone gives me a color map between language names and colors, I will happily use that

 ![image](https://global.discourse-cdn.com/julialang/original/3X/3/5/351eea91d64928435758d2af5c198b20a0686d96.png)

You need `tokei` installed

> <https://gist.github.com/Moelf/21fcb48fa1543b4c68f3235fe1411809>

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [October 11, 2023, 11:51am UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/84 "2023-10-11T11:51:25Z")

</div>

> [@jling](#):
>
> You need `tokei` installed

Or you can use `Tokei_jll`

---

<div class="post-metadata">

**Author:** ![xgdgsc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xgdgsc/32/608_2.png) [@xgdgsc](https://discourse.julialang.org/u/xgdgsc)\
**Post date:** [October 11, 2023, 12:43pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/85 "2023-10-11T12:43:56Z")

</div>

> <https://github.com/markembling/github-languages-palette/blob/master/palettes/githublangs.json>

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 11, 2023, 1:11pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/86 "2023-10-11T13:11:06Z")

</div>

ah man that thing is missing Odin… anyway, I have updated the Gist

 ![image](https://global.discourse-cdn.com/julialang/original/3X/f/a/fa14be907cb3175a037dd4a1c4e85896c230e554.png)

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 11, 2023, 1:29pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/87 "2023-10-11T13:29:18Z")

</div>

now I really want to also shrink code size …

```diff
diff --git a/julia/related.jl b/julia/related.jl
index e21c8a2..b79b866 100644
--- a/julia/related.jl
+++ b/julia/related.jl
@@ -1,9 +1,4 @@
-using JSON3
-using StructTypes
-using Dates
-using StaticArrays
-
-# warmup is done by hyperfine
+using JSON3, Dates, StaticArrays
 
 function relatedIO()
     json_string = read("../posts.json", String)
@@ -33,12 +28,9 @@ struct RelatedPost
     related::SVector{5,PostData}
 end
 
-StructTypes.StructType(::Type{PostData}) = StructTypes.Struct()
-
-function fastmaxindex!(xs::Vector, topn, maxn, maxv)
- maxn .= 1
- maxv .= 0
- top = maxv[1]
+function fastmaxindex!(xs, topn, maxn::AbstractVector{T}, maxv) where T
+ maxn .= one(T)
+ maxv .= top = zero(T)
     for (i, x) in enumerate(xs)
         if x > top
             maxv[1] = x
@@ -59,13 +51,7 @@ function fastmaxindex!(xs::Vector, topn, maxn, maxv)
 end
 
 function related(posts)
- for T in (UInt8, UInt16, UInt32, UInt64)
- if length(posts) < typemax(T)
- return related(T, posts)
- end
- end
-end
-function related(::Type{T}, posts) where {T}
+ T = UInt32
     topn = 5
     # key is every possible "tag" used in all posts
     # value is indicies of all "post"s that used this tag

```

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 11, 2023, 1:38pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/88 "2023-10-11T13:38:54Z")

</div>

Code size is typically better off measured in terms of the compressed file size instead of the number of lines of code.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 11, 2023, 1:46pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/89 "2023-10-11T13:46:29Z")

</div>

if you have a way to extract that information that would be great.

there are folders with configuration files for Swift and Java.

You need a way to extract only source code files, and exclude comments, and compress it.

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 11, 2023, 1:49pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/90 "2023-10-11T13:49:21Z")

</div>

just `gzip` the source file. Comment filtering would be a bit annoying though yeah

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 11, 2023, 1:50pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/91 "2023-10-11T13:50:40Z")

</div>

I can’t. because I don’t know what’s source file (i.e. I’m not going to code that up manually). I rely on `Tokei` to only count source files + only line with code

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [October 11, 2023, 2:12pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/92 "2023-10-11T14:12:17Z")

</div>

The size of the source is kind of a murky metric for a piece of code, compressed or uncompressed. It is tempting to have a 2-axis chart, but either a regular 1-axis bar-chart is enough or a different 2nd axis can be chosen. For example:

- x-axis = time/post
- y-axis = # of posts  
and see how the programs scale with # posts.

---

<div class="post-metadata">

**Author:** ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)\
**Post date:** [October 11, 2023, 5:42pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/93 "2023-10-11T17:42:24Z")

</div>

Code size [in that benchmark](https://benchmarksgame-team.pages.debian.net/benchmarksgame/how-programs-are-measured.html) is

> ## How source code size is measured
> 
> We start with the source-code markup you can see, remove comments, remove duplicate whitespace characters, and then apply minimum GZip compression. The measurement is the size in bytes of that GZip compressed source-code file.
> 
> Thanks to Brian Hurt for the idea of using **size of compressed source code** instead of lines of code.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 11, 2023, 6:38pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/94 "2023-10-11T18:38:05Z")

</div>

> [@jar1](#):
>
> Thanks to Brian Hurt for the idea of using **size of compressed source code** instead of lines of code.

size of compressed code isn’t that great of a metric if you’re concerned about human readability and writeability. If you’ve got a lot of boilerplate that you have to repeat often, it’s easy to compress this away, but you can’t not read it or not write it… it’s annoying crap in the code file.

So from a human readability perspective you’d want also something like compression ratio. A language with a very high compression ratio is basically a language that sucks to write or read.

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 11, 2023, 6:41pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/95 "2023-10-11T18:41:13Z")

</div>

I mean, nobody _likes_ reading repeated boilerplate, but also compressing it seems pretty fair. It doesn’t completely ignore the boilerplate, but it reduces it’s cost to the code length metric, which is also how I tend to interact with boilerplate when I’m reading it.

I glance at it, skim it, and if I see something repeated a bunch, I look at the repeated code less and less each time it’s repeated.

* * *

As a concrete example, consider this part of the diff @jling posted:

```diff
-using JSON3
-using StructTypes
-using Dates
-using StaticArrays

+using JSON3, Dates, StaticArrays

```

When the word `using` is repeated a bunch of times in a very predictable way like this, it’s very easy for my eye to just completely skip it and only look at the words to the right of the `using` on each line. I don’t need to figure out what is going on each time, I immediately see that all of those are `using` statements.

In fact, if there were a few more `using`s in there, I might actually find the verbose version more pleasant to read than the one-liner version.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 11, 2023, 6:57pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/96 "2023-10-11T18:57:42Z")

</div>

Yes, but consider this from the [Wikipedia on boilerplate](https://en.wikipedia.org/wiki/Boilerplate_code), and then multiply by 25 for some code with 25 different subclasses of Pet.

```julia
public class Pet {
    private String name;
    private Person owner;

    public Pet(String name, Person owner) {
        this.name = name;
        this.owner = owner;
    }

    public String getName() {
        return name;
    }

    public void setName(String name) {
        this.name = name;
    }

    public Person getOwner() {
        return owner;
    }

    public void setOwner(Person owner) {
        this.owner = owner;
    }
}

```

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 11, 2023, 7:03pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/97 "2023-10-11T19:03:30Z")

</div>

Okay, but is the experience of reading that really so significantly different from the experience of reading this hypothetical syntax?

```plaintext
public class Pet {
    private String name
    private Person owner
    Pet(String name, Person owner) = {this.name = name; this.owner = owner}
    String getName() = return name
    void setName(String name) = this.name = name
    Person getOwner() = return owner
    void setOwner(Person owner) = this.owner = owner}

```

I’m not really convinced it’s such a big difference.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 11, 2023, 7:06pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/98 "2023-10-11T19:06:23Z")

</div>

I am not a Java programmer, but my impression was Java allowed inheriting interface but not implementation. So if you have 25 pets, say Dog, Cat, Bird, Fish etc… then you have this boilerplate 25 times even though it does absolutely nothing different each time. Then when you gzip it it will compress about 50 to 1. Which means when you’re reading through this code only 2% of it is even worth looking at in some sense.

So it seems to me that if you want a “human complexity” metric it should be something that takes into account when the compression ratio is very high that reading or writing that code really sucks.

EDIT: I looked this up, and apparently maybe I’m wrong, Java does allow some implementation inheritance, but I do think there are languages that don’t and you wind up with lots of heavily repeated boilerplate.

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [October 11, 2023, 7:15pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/99 "2023-10-11T19:15:39Z")

</div>

I think non-compressed metrics reward golfing too much (eg get everything onto one line if it is per line, or use all 1-character variable names, if it is by character). Compressing devalues these approaches, which I think is appropriate.

(Also the data compression limit is the Shannon entropy, so there’s a nice information theoretic flavor to compression-based metrics).

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 11, 2023, 7:26pm UTC](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611/100 "2023-10-11T19:26:41Z")

</div>

To be clear, what I’m actually saying is not anything directly against compression as a metric, I like using compression, its more that it’s insufficient by itself, the compression ratio is also an additional important dimensionless ratio in the question.

[Previous page](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611.md?page=4)

[Next page](https://discourse.julialang.org/t/funny-benchmark-with-julia-no-longer-at-the-bottom/104611.md?page=6)
