# Julia cookbook available

**URL:** <https://discourse.julialang.org/t/julia-cookbook-available/23358>\
**Category:** New to Julia\
**Created:** [April 20, 2019, 8:40pm UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358 "2019-04-20T20:40:21Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![iwelch](https://avatars.discourse-cdn.com/v4/letter/i/8c91f0/32.png) [@iwelch](https://discourse.julialang.org/u/iwelch)\
**Post date:** [April 20, 2019, 8:40pm UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/1 "2019-04-20T20:40:22Z")

</div>

I am pleased to announce the availability of my julia cookbook at

\*\*\* [http://julia.cookbook.tips](http://julia.cookbook.tips)\*\*

most of the early chapters are in pretty good shape and updated to julia 1.0. the latter and more complex subject chapters are a mix of good and bad shapes. this is partly because julia is itself still shifting.

* * *

PS: as to myself, I will come back to julia when it will have acquired [a] superior data frame handling, ideally language-integrated; [b] superior fast (gzipped) csv IO, and [c] parallel processing. until then, for the kind of data analysis tasks that _I_ am involved in, julia remains much slower than R. but julia has many other excellent use cases.

---

<div class="post-metadata">

**Author:** ![StevenSiew](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevensiew/32/218393_2.png) [@StevenSiew](https://discourse.julialang.org/u/StevenSiew)\
**Post date:** [April 20, 2019, 9:00pm UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/2 "2019-04-20T21:00:08Z")

</div>

Is it vegetarian or non-vegetarian?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [April 21, 2019, 12:09am UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/3 "2019-04-21T00:09:53Z")

</div>

Glad someone has similar experience

---

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [April 21, 2019, 2:27pm UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/4 "2019-04-21T14:27:18Z")

</div>

> [@iwelch](#):
>
> [a] superior data frame handling, ideally language-integrated; [b] superior fast (gzipped) csv IO, and [c] parallel processing

Do you mind elaborating on [a] and [c]? What are the specific issues that bother you the most?

---

<div class="post-metadata">

**Author:** ![js135005](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/js135005/32/8219_2.png) [@js135005](https://discourse.julialang.org/u/js135005)\
**Post date:** [April 21, 2019, 3:08pm UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/5 "2019-04-21T15:08:29Z")

</div>

Also have you looked at TableReader.jl for b? It processes gzip files directly and, for many CSV applications, it seems to be quite fast and competitive with the R readers.

---

<div class="post-metadata">

**Author:** ![rdeits](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rdeits/32/286_2.png) [@rdeits](https://discourse.julialang.org/u/rdeits)\
**Post date:** [April 21, 2019, 3:54pm UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/6 "2019-04-21T15:54:12Z")

</div>

Are you interested in corrections to this, and/or do you have a preferred mechanism for submitting them? I’ve found a few factual errors, particularly relating to the way the cookbook talks about the performance of “machine native” types.

---

<div class="post-metadata">

**Author:** ![iwelch](https://avatars.discourse-cdn.com/v4/letter/i/8c91f0/32.png) [@iwelch](https://discourse.julialang.org/u/iwelch)\
**Post date:** [April 21, 2019, 6:36pm UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/7 "2019-04-21T18:36:40Z")

</div>

hi tk—I thought for a while how best to communicate this and how to figure out in the future whether julia has become mature enough for _our_ basic data needs. I decided that I may as well write some simple R code that demonstrates the need. the first program is not the test, but just writes a typpical 1.8GB data set to disk:

```julia-auto
library(data.table)
set.seed(0)

NF <- 1000000

permno <- 1:NF
startdt <- as.integer(runif( NF )*500)
enddt <- as.integer(startdt+1+runif( NF )*500)

all.nrows <- sum(enddt)-sum(startdt)+NF
all.p <- rep(NA, all.nrows)
all.t <- rep(NA, all.nrows)

cm <- 1
for (permno in 1:NF) {
    all.p[cm:(cm+enddt[permno]-startdt[permno]) ] <- permno
    all.t[cm:(cm+enddt[permno]-startdt[permno]) ] <- startdt[permno]:enddt[permno]
}

d <- data.frame( permno= all.p, t=all.t )
d <- within(d, prc <- rnorm( nrow(d), 100, 1 ))

d[["prc"]][sample(1:nrow(d), nrow(d)/30 )] <- NA

fwrite(d, file="test.csv")
system("gzip test.csv")

cat("Test Data Set Created\n")
system("ls -lh test.csv.gz")

```

so, please use the above csv file as input, both into R and julia code. this kind of csv coding is standard in my field.

## Test Code

now, let’s get to the benchmark. the following test program does what my students and I need to do most of the time: read an irregular data set, create some time-series and cross-sectional variables (here, returns and market returns), run regressions by and for many firms, and then finally save the results in a csv file.

```julia-auto
print( system.time( {
    lagseries <- function(x) x[c(NA, 1:(length(x) - 1))]

    d <- fread("test.csv.gz")

    ## calculate rates of return in a panel
    d <- within(d, ret <- prc / lagseries(prc)-1)
    d <- within(d, ret <- ifelse( permno != lagseries(permno), NA, ret ))

    ## create the market rate of return
    d <- within(d, mktret <- ave( ret, t, FUN=function(x) mean(x,na.rm=TRUE) ))

    ## a market-model creates an alpha and a beta
    marketmodel <- function(d) {
        if (nrow(d) < 5) return(NULL) ## minimum of 5 observations
        coef(lm( ret ~ mktret, data=d )) ## the actual regression
    }

    indexes <- split( 1:nrow(d), d$permno )
    betas <- mclapply( indexes, FUN=function(.index) marketmodel(d[.index, , drop=FALSE]) )

    ## transform it into a nicer version
    betas <- do.call("rbind", betas)
    names(betas) <- c("alpha", "beta")

    ## and write it
    write.csv(betas, file="betas.csv")
} ))

```

This can be further speeded up by using the R `compiler` package, wrapping a `compfun` around the `marketmodel` function. But let’s just leave it this way.

I have not managed to write a julia implementation that can compete with R. by this, I mean julia no more than 30% slower than R on a 6 to 8-core machine. ideally, julia would be as fast and have nicer and cleaner code.

if someone can demonstrate competitive julia code for this task, I would be thrilled and will reconsider using julia here at UCLA for teaching quant finance.

---

<div class="post-metadata">

**Author:** ![iwelch](https://avatars.discourse-cdn.com/v4/letter/i/8c91f0/32.png) [@iwelch](https://discourse.julialang.org/u/iwelch)\
**Post date:** [April 21, 2019, 6:40pm UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/8 "2019-04-21T18:40:21Z")

</div>

hi robin—I would indeed be interested.

if you send me an email with a username and passcode, I can give you wiki editing privileges. (same holds for anyone else who wants to tinker with it.)

(I also wrote some code to test automatically that everything is up to date and still gives the very same output [when Julia or packages update], but this takes too much maintenance if I don’t end up using julia myself any longer.)

/iaw

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [April 21, 2019, 8:27pm UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/9 "2019-04-21T20:27:28Z")

</div>

It would be quite challenging for any Julia package to beat data.table in a short time. There have been so much work put into pandas, but it is still much slower than data.table in many tasks.

Since data.table is written in C, and Julia can call C code easily. Is it possible to use data.table code to create a Julia package? Python already has it.

> **[datatable](https://pypi.org/project/datatable/)**
>
> Python library for fast multi-threaded data manipulation and munging.

---

<div class="post-metadata">

**Author:** ![iwelch](https://avatars.discourse-cdn.com/v4/letter/i/8c91f0/32.png) [@iwelch](https://discourse.julialang.org/u/iwelch)\
**Post date:** [April 21, 2019, 9:33pm UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/10 "2019-04-21T21:33:35Z")

</div>

yes, R is heavily optimized for data analysis. I don’t think julia will be able to beat it.

if julia is half the speed for applied data analysis, then few data analysts will want to switch from R to julia, whether we like it or not.

---

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [April 22, 2019, 12:33am UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/11 "2019-04-22T00:33:19Z")

</div>

I think quite a few people, including myself, are waiting for the new multithreading feature so that we can speed up our code more easily. Reading CSV file can be parallelized but the current story is not great because IO operations are not thread safe at the moment.

I’m glad that this PR has been merged though.

[https://github.com/JuliaLang/julia/pull/22631](https://github.com/JuliaLang/julia/pull/22631)

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [April 22, 2019, 1:08am UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/12 "2019-04-22T01:08:43Z")

</div>

I am happy happy with the current status of data wrangling tools in Julia. Importing data quickly is important, but being able to import out-of-memory data is more important to me.

Are there are any working examples of JuliaDB working with larger-than-memory data?

---

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [April 22, 2019, 1:15am UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/13 "2019-04-22T01:15:39Z")

</div>

Which packages do you use regularly?

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [April 22, 2019, 1:21am UTC](https://discourse.julialang.org/t/julia-cookbook-available/23358/14 "2019-04-22T01:21:03Z")

</div>

I only use DataFramesMeta. If I can use JuliaDB to work on out-of-core large data, that would be great.
