# Interesting observations about Julia language google trends time series

**URL:** <https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519>\
**Category:** Community\
**Tags:** time-series\
**Created:** [March 5, 2018, 9:02pm UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519 "2018-03-05T21:02:49Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [March 5, 2018, 9:02pm UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/1 "2018-03-05T21:02:49Z")

</div>

**Update: removed the peak in a followup in the thread**

I have downloaded some google trends data about Julia and I have I did some basic anlaysis, and it looks like Julia is on a slightly downward trend.

![julia_ts_decomposed](https://global.discourse-cdn.com/julialang/original/3X/8/7/8744b8ba7f64c828a3addc18cb41c80d403dbcb9.png)

Also there is typically a big dip on gogole trends in December for other languages. December is usually time off for people, so there is usually a big dropoff which you can quantify by estimating the seasonality for each month, e.g. for R the December seasonality is 89% i.e. December search volumes is usually 11% lower than what the trend predicts.

If I don’t moreve the spike in Feb 2015, then for Julia the December seasonality is 1.00036, so there is no drop off. This could mean that Julia isn’t used so much at work and so when workers take time off it’s not as affected or it could mean that people are interested in Julia and want to use their free time to learn more 🙂. Interestingly, Julia’s Jun, Jul, Aug, Sep, Oct search volumes seasonality factor are all below 1.

If I remove the spike then the December seasonality is 0.9936326.

Here is my code. I plan to expand this into a look at all popular languages

```julia
using RCall
using GLM, DataFrames, DataFramesMeta, Lazy, Plots

langs = DataFrame(
    name = ["julia", "r"], 
    lang = ["/m/0j3djl7", "/m/0212jm"]
)

# jld = gtrends(keyword = "/m/0j3djl7") #julia
# jld = gtrends(keyword = "/m/0212jm") # r
# jld = gtrends(keyword = "/m/09gbxjr") # go
# jld = gtrends(keyword = "python")
# jld = gtrends(keyword = "/m/0n50hxv") #typescript
# jld = gtrends("/m/02js86") # groovy
# jld = gtrends("/m/0ncc1sv") #elm
# jld = gtrends(keyword = "/m/06ff5")
# jld = gtrends(keyword = "/m/02l0yf8", time="all") # sas

function analysis_trend(geo = "", lang = "/m/0j3djl7", title = "")
    res = R"""
    library(gtrendsR)
    gtrends(keyword = $lang, geo = $geo)[[1]] #julia
    """

    # trends dataset
    df = DataFrame(res)

    # create a daily average
    df1 = DataFrame(
        day = reduce(vcat, (a .- Dates.Day.(0:6) for a in Date.(df[:date]))),
        hits = repeat(df[:hits], inner = 7)
    )

    sort!(df1, cols = :day)

    df1[:year] = Dates.year(df1[:day])
    df1[:month] = Dates.month(df1[:day])

    df2 = @> df1 begin
        @by([:year, :month], meanh = mean(:hits))
    end

    sort!(df2, cols=[:year, :month])

    df3 = deepcopy(df2)
    if lang == "/m/0j3djl7" # if julia then clean up
        # there is a big spike in Feb 2015 so smooth that out
        mm = maximum(df2[:meanh])
        @> df2 begin
            @where(:meanh .== mm)
        end

        mdf2 = @> df2 begin
            @where((:year .== 2015) .& (:month .== 1) .| (:year .== 2015) .& (:month .== 3))
        end

        mmm = mean(mdf2[:meanh])

        idx = find(
            (df2[:year] .== 2015) .& (df2[:month] .== 2)
        )

        df3 = deepcopy(df2)
        df3[idx,:meanh] = mmm
        plot(df2[:meanh], label="original")
        plot!(df3[:meanh], label="removed spike")
        savefig("julia_trend.png")
    end

    @rput df3

    dt = R"""
    png("julia_ts_decomposed.png")
    dt = decompose(ts(df3$meanh, deltat=1/12, start=c(2013, 3)), type="m")
    plot(dt)
    dev.off()
    dt
    """
end

analysis_trend("US") # Julia in the US
analysis_trend() # Julia in the world
analysis_trend("", "/m/0212jm") # R in the world

```

---

<div class="post-metadata">

**Author:** ![Mattriks](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mattriks/32/351_2.png) [@Mattriks](https://discourse.julialang.org/u/Mattriks)\
**Post date:** [March 5, 2018, 9:28pm UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/2 "2018-03-05T21:28:10Z")

</div>

Did you do the comparison with the first 5 years of a ‘new’ language? That would be statistically interesting.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [March 5, 2018, 9:55pm UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/3 "2018-03-05T21:55:08Z")

</div>

> [@xiaodai](#):
>
> For Julia the December seasonality is 1.00036, so there is no drop off. This could mean that Julia isn’t used so much at work and so when workers take time off it’s not as affected or it could mean that people are interested in Julia and want to sue their free time to learn more 🙂.

Academics don’t take December off. Scratch that, they don’t take time off. You just found proof 😉. It’s `1.00036`, not `1.0`, because there’s a few people who don’t publish in December (they indeed perished).

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [March 6, 2018, 7:23am UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/4 "2018-03-06T07:23:36Z")

</div>

Note that two-sided filters are notorious for introducing spurious patterns (eg [Hamilton (2017)](http://econweb.ucsd.edu/~jhamilto/hp.pdf) has some nice examples). There was some kind of a blip in 2015, which is generating the whole hump. Otherwise, there is no apparent trend either way since 2014 (not that one can draw strong conclusions from Google trends anyway).

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [March 6, 2018, 10:36am UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/5 "2018-03-06T10:36:56Z")

</div>

> [@ChrisRackauckas](#):
>
> It’s 1.00036, not 1.0,

I updated the originally analysis by first removing the peak in Feb 2015 (what event happened then?). Now the seasonality in Dec is 0.9936983 so there is a bit of slacking off there 😜. I repeated the anlaysis for US only and I found the same pattern, and I assume December is big thing in the US due to Christmas.

 ![julia_trend](https://global.discourse-cdn.com/julialang/original/3X/4/0/40d9d2cff6abdc0aaf63d0614b71393f346ad62d.png)

![julia_ts_decomposed](https://global.discourse-cdn.com/julialang/original/3X/7/a/7a381d96c1b511693f2a987336060be22ed08c76.png)

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [March 6, 2018, 10:43am UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/6 "2018-03-06T10:43:07Z")

</div>

> [@xiaodai](#):
>
> Now the seasonality in Dec is 0.9936983 so there is a bit of slacking off there 😜.

It’s those gosh darn Millennials.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [March 6, 2018, 10:43am UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/7 "2018-03-06T10:43:25Z")

</div>

> [@xiaodai](#):
>
> I assume December is big thing in the US due to Christmas.

I guess Julia 1.0 should be released around Christmas then.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [March 6, 2018, 10:51am UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/8 "2018-03-06T10:51:48Z")

</div>

Now I just need a Google Trends Julia Package and a time series package in Julia to complete the analysis in pure Julia

---

<div class="post-metadata">

**Author:** ![y4lu](https://avatars.discourse-cdn.com/v4/letter/y/47e85d/32.png) [@y4lu](https://discourse.julialang.org/u/y4lu)\
**Post date:** [March 14, 2018, 5:37am UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/9 "2018-03-14T05:37:10Z")

</div>

A little attempt at an overall season-length trend  
Shown are 8 roughly balanced centroids, and the mean with ± 1 standard deviation  
There appears to be about a 5% of St. dev. increase over 13 weeks as the current long-term average  
 ![season-trend](https://global.discourse-cdn.com/julialang/original/3X/b/d/bd7ef501f1978fcd1f4be1777cf7b25ceaca1053.png)

mash-up of sub-components  
 ![season-trendb](https://global.discourse-cdn.com/julialang/original/3X/7/6/76fb2bd508e1641e708988cb383a49700bbb5252.png)

---

<div class="post-metadata">

**Author:** ![Fred](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fred/32/14175_2.png) [@Fred](https://discourse.julialang.org/u/Fred)\
**Post date:** [March 14, 2018, 7:49am UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/10 "2018-03-14T07:49:04Z")

</div>

@xiaodai please could you explain the “@\>” in your code ? 🤨

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [March 14, 2018, 7:57am UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/11 "2018-03-14T07:57:39Z")

</div>

It’s from Lazy.jl it means pipe the results of the last line into the first argument of macro/function.

So

```julia
@> df begin
  FN(2)
  GN(7)
end

```

Is the same as

```julia
GN(FN(df,2),7)

```

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [March 14, 2018, 7:58am UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/12 "2018-03-14T07:58:20Z")

</div>

> [@y4lu](#):
>
> little attempt at an overall season-length trend
> 
> Shown are 8 roughly balanced centroids, and the mean with ± 1 standard deviation
> 
> There appears to be about a 5% of St. dev. increase over 13 weeks as the current long-term average

In layman’s terms?

---

<div class="post-metadata">

**Author:** ![y4lu](https://avatars.discourse-cdn.com/v4/letter/y/47e85d/32.png) [@y4lu](https://discourse.julialang.org/u/y4lu)\
**Post date:** [March 14, 2018, 9:17am UTC](https://discourse.julialang.org/t/interesting-observations-about-julia-language-google-trends-time-series/9519/13 "2018-03-14T09:17:40Z")

</div>

I might conclude that it isn’t on too much of a downward trend just yet, but the tools used were a bit crusty  
Or, on average over a given three months, approx 5% of the amount it jumps around translates to an increase, if that’s a sensible metric

- balanced as in the centroids are roughly proportionate, or have a similar weight / number of units per group
- season-length would be like seasonal data augmentation, so allowing overlapped offsets of 13 week sequences (~5x13 \> 52x13)

[Code for a couple of k-means variants](https://nofile.io/f/YuvEJkTfeHm/clustersB.jl), balanced (as above, prefix zl\_) and balanced + floating (prefix zf\_) which was the alg used
