# Reading in Excel Data from GitHub

**URL:** https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055
**Category:** New to Julia
**Tags:** github, xlsx, io
**Created:** [March 26, 2021, 10:59pm UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055 "2021-03-26T22:59:04Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![Michael\_Barmann](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/michael_barmann/32/20389_2.png) [@Michael\_Barmann](https://discourse.julialang.org/u/Michael_Barmann)
#### Post date: [March 26, 2021, 10:59pm UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055/1 "2021-03-26T22:59:04Z")

</div>

I am preparing a Minimal Reproducible Example which involves reading in some data contained in an Excel file. I’m trying to figure out the best way to make the data easily accessible for anyone who wants to run my code. I’ve posted the data here: [sample\_data](https://github.com/MichaelBarmann/ParamEst_MRE/blob/main/sample_data.xlsx)

In Python for example, I can simply do

```julia
import pandas as pd
df = pd.read_excel("https://github.com/MichaelBarmann/ParamEst_MRE/blob/main/sample_data.xlsx")

```

and the data is placed into the DataFrame df.

In Julia I’ve tried `xl_file = XLSX.readxlsx("https://github.com/MichaelBarmann/ParamEst_MRE/blob/main/sample_data.xlsx")`, but unfortunately this doesn’t work.

Is there an analogous way of doing this in Julia? (using XLSX, or something similar?)

---

<div class="post-metadata">

### Author: ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)
#### Post date: [March 27, 2021, 12:33am UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055/2 "2021-03-27T00:33:08Z")

</div>

If you are willing to abandon excel in favor of CSV, CSV.jl can do that

See [here](https://csv.juliadata.org/stable/#CSV.jl-Documentation).

> To read a csv file from a url, use the HTTP.jl package, where the `HTTP.Response` body can be passed like:
> 
> ```julia
> f = CSV.File(HTTP.get(url).body)
> 
> ```

Maybe the same strategy works for excel actually?

---

<div class="post-metadata">

### Author: ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)
#### Post date: [March 27, 2021, 6:59am UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055/3 "2021-03-27T06:59:32Z")

</div>

I don’t think this works with `XSLX` out of the box, as the read methods expect a filepath (string), and `HTTP.get` returns a vector of UInts. The easiest thing would probably to just download the file to a temp dir and then read it in from there, but if it’s just some standard tabular data I agree with Peter that it would be better to move to a saner format than xlsx like csv, or maybe arrow.

---

<div class="post-metadata">

### Author: ![Michael\_Barmann](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/michael_barmann/32/20389_2.png) [@Michael\_Barmann](https://discourse.julialang.org/u/Michael_Barmann)
#### Post date: [March 27, 2021, 5:02pm UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055/4 "2021-03-27T17:02:29Z")

</div>

Thanks, this is helpful! I had 3 spreadsheets in my Excel file, so I had to put each into an individual CSV file before attempting to implement your suggestion. Unfortunately, I get a weird error:

```julia
using HTTP, CSV
data_url = "https://github.com/MichaelBarmann/ParamEst_MRE/blob/main/sample_data.csv"
my_file = CSV.File(HTTP.get(data_url).body; header = 1)

```

```julia
┌ Warning: thread = 1 warning: parsed expected 1 columns, but didn't reach end of line around data row: 53. Ignoring any extra columns on this row
└ @ CSV C:\Users\Michael\.julia\packages\CSV\CJfFO\src\file.jl:606
┌ Warning: thread = 1 warning: parsed expected 1 columns, but didn't reach end of line around data row: 57. Ignoring any extra columns on this row
└ @ CSV C:\Users\Michael\.julia\packages\CSV\CJfFO\src\file.jl:606
┌ Warning: thread = 1 warning: parsed expected 1 columns, but didn't reach end of line around data row: 149. Ignoring any extra columns on this row
└ @ CSV C:\Users\Michael\.julia\packages\CSV\CJfFO\src\file.jl:606
┌ Warning: thread = 1 warning: parsed expected 1 columns, but didn't reach end of line around data row: 163. Ignoring any extra columns on this row
└ @ CSV C:\Users\Michael\.julia\packages\CSV\CJfFO\src\file.jl:606
┌ Warning: thread = 1 warning: parsed expected 1 columns, but didn't reach end of line around data row: 178. Ignoring any extra columns on this row
└ @ CSV C:\Users\Michael\.julia\packages\CSV\CJfFO\src\file.jl:606
┌ Warning: thread = 1 warning: parsed expected 1 columns, but didn't reach end of line around data row: 179. Ignoring any extra columns on this row
└ @ CSV C:\Users\Michael\.julia\packages\CSV\CJfFO\src\file.jl:606
178-element CSV.File{false}:
 CSV.Row: (<!DOCTYPE html> = "<html lang=\"en\" >",)
 CSV.Row: (<!DOCTYPE html> = " <head>",)
 CSV.Row: (<!DOCTYPE html> = " <meta charset=\"utf-8\">",)
 CSV.Row: (<!DOCTYPE html> = " <link rel=\"dns-prefetch\" href=\"https://github.githubassets.com\">",)
 CSV.Row: (<!DOCTYPE html> = " <link rel=\"dns-prefetch\" href=\"https://avatars.githubusercontent.com\">",)
 CSV.Row: (<!DOCTYPE html> = " <link rel=\"dns-prefetch\" href=\"https://github-cloud.s3.amazonaws.com\">",)
 CSV.Row: (<!DOCTYPE html> = " <link rel=\"dns-prefetch\" href=\"https://user-images.githubusercontent.com/\">",)
 CSV.Row: (<!DOCTYPE html> = " <link crossorigin=\"anonymous\" media=\"all\" integrity=\"sha512-+D6Fy+EweT791upBwWDwZQxdKdpY58/tEdY1xnOvq7urHPMZe2ZtMhRmzWUL6+Dknj2BwHdlVvxPHDB6ep7wQQ==\" rel=\"stylesheet\" href=\"https://github.githubassets.com/assets/frameworks-f83e85cbe130793efdd6ea41c160f065.css\" />",)
 CSV.Row: (<!DOCTYPE html> = " <link crossorigin=\"anonymous\" media=\"all\" integrity=\"sha512-B+uagPXLI8XoStGPc9DrQ9xVWfBS0VvAKvYTQopaizdxo8T46TCtAVE9sY15FQ0kseA1LB3V5GKklDGuwjqCHg==\" rel=\"stylesheet\" href=\"https://github.githubassets.com/assets/site-07eb9a80f5cb23c5e84ad18f73d0eb43.css\" />",)
 CSV.Row: (<!DOCTYPE html> = " <link crossorigin=\"anonymous\" media=\"all\" integrity=\"sha512-BgutnhwyiKQLFDdsLnN+C/pdoX66e9VdVPnjGR1KkI6lMDTbqn8SWQLjp3dqzfGblsErstNfjUq94M/kZ8/1qA==\" rel=\"stylesheet\" href=\"https://github.githubassets.com/assets/behaviors-060bad9e1c3288a40b14376c2e737e0b.css\" />",)
 CSV.Row: (<!DOCTYPE html> = " ",)
 CSV.Row: (<!DOCTYPE html> = " ",)
 CSV.Row: (<!DOCTYPE html> = " ",)
 ⋮
 CSV.Row: (<!DOCTYPE html> = "</div>",)
 CSV.Row: (<!DOCTYPE html> = " </div>",)
 CSV.Row: (<!DOCTYPE html> = "</div>",)
 CSV.Row: (<!DOCTYPE html> = " </main>",)
 CSV.Row: (<!DOCTYPE html> = " </div>",)
 CSV.Row: (<!DOCTYPE html> = " </div>",)
 CSV.Row: (<!DOCTYPE html> = " ",)
 CSV.Row: (<!DOCTYPE html> = "<div class=\"footer container-xl width-full p-responsive\" role=\"contentinfo\">",)
 CSV.Row: (<!DOCTYPE html> = " <div class=\"position-relative d-flex flex-row-reverse flex-lg-row flex-wrap flex-lg-nowrap flex-justify-center flex-lg-justify-between pt-6 pb-2 mt-6 f6 color-text-secondary border-top color-border-secondary \">",)
 CSV.Row: (<!DOCTYPE html> = " <ul class=\"list-style-none d-flex flex-wrap col-12 col-lg-5 flex-justify-center flex-lg-justify-between mb-2 mb-lg-0\">",)
 CSV.Row: (<!DOCTYPE html> = " <li class=\"mr-3 mr-lg-0\">&copy; 2021 GitHub",)
 CSV.Row: (<!DOCTYPE html> = " <li class=\"mr-3 mr-lg-0\"><a href=\"https://docs.github.com/en/github/site-policy/github-terms-of-service\" data-ga-click=\"Footer",)

```

Any idea what’s going wrong? Do I need to specify some additional keywords for it to property parse the data? (I tried including `header = 1`, and specifying the column types, but it didn’t change anything…)

---

<div class="post-metadata">

### Author: ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)
#### Post date: [March 27, 2021, 5:25pm UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055/5 "2021-03-27T17:25:10Z")

</div>

@Michael_Barmann, using the github’s link to the raw version of your csv file, it works:

```julia
using HTTP, CSV
data_url = "https://raw.githubusercontent.com/MichaelBarmann/ParamEst_MRE/main/sample_data.csv"
my_file = CSV.File(HTTP.get(data_url).body)
df = DataFrame(my_file)

103×3 DataFrame
 Row │ Date t Ic    
     │ String Int64 Int64 
─────┼──────────────────────────
   1 │ 8/24/2020 0 79
   2 │ 8/25/2020 1 168
   3 │ 8/26/2020 2 222
   4 │ 8/27/2020 3 282
   5 │ 8/28/2020 4 347
   6 │ 8/29/2020 5 397
   7 │ 8/30/2020 6 501
   8 │ 8/31/2020 7 731
   9 │ 9/1/2020 8 930
  10 │ 9/2/2020 9 1050
  11 │ 9/3/2020 10 1138
  12 │ 9/4/2020 11 1242
  13 │ 9/5/2020 12 1279
  14 │ 9/6/2020 13 1316
  ⋮ │ ⋮ ⋮ ⋮
  90 │ 11/21/2020 89 3998
  91 │ 11/22/2020 90 4017
  92 │ 11/23/2020 91 4043
  93 │ 11/24/2020 92 4056
  94 │ 11/25/2020 93 4070
  95 │ 11/26/2020 94 4103
  96 │ 11/27/2020 95 4124
  97 │ 11/28/2020 96 4195
  98 │ 11/29/2020 97 4221
  99 │ 11/30/2020 98 4253
 100 │ 12/1/2020 99 4276
 101 │ 12/2/2020 100 4301
 102 │ 12/3/2020 101 4321
 103 │ 12/4/2020 102 4346
                 75 rows omitted

```

---

<div class="post-metadata">

### Author: ![Michael\_Barmann](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/michael_barmann/32/20389_2.png) [@Michael\_Barmann](https://discourse.julialang.org/u/Michael_Barmann)
#### Post date: [March 27, 2021, 5:43pm UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055/6 "2021-03-27T17:43:44Z")

</div>

Awesome, thank you! Where can I find the url of the “raw version”? Or is the raw version url always just obtained by replacing '“github” with “raw.githubusercontent”?

---

<div class="post-metadata">

### Author: ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)
#### Post date: [March 27, 2021, 5:46pm UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055/7 "2021-03-27T17:46:37Z")

</div>

Look here:  
 ![raw_github](https://global.discourse-cdn.com/julialang/original/3X/8/7/87aa14e94db89a6691493f472bb08ec5bef4730b.png)

---

<div class="post-metadata">

### Author: ![Michael\_Barmann](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/michael_barmann/32/20389_2.png) [@Michael\_Barmann](https://discourse.julialang.org/u/Michael_Barmann)
#### Post date: [March 27, 2021, 5:47pm UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055/8 "2021-03-27T17:47:27Z")

</div>

Ah, thanks!!

---

<div class="post-metadata">

### Author: ![rikh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rikh/32/204104_2.png) [@rikh](https://discourse.julialang.org/u/rikh)
#### Post date: [September 17, 2021, 3:15pm UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055/9 "2021-09-17T15:15:23Z")

</div>

You can also use Artifacts to manage downloading a GitHub repository at a specific version for you. That way, you can avoid downloading the file every time that you or an user restarts Julia

Docs: [Artifacts · The Julia Language](https://docs.julialang.org/en/v1/stdlib/Artifacts/)

---

<div class="post-metadata">

### Author: ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)
#### Post date: [August 30, 2026, 8:01am UTC](https://discourse.julialang.org/t/reading-in-excel-data-from-github/58055/10 "2026-08-30T08:01:14Z")

</div>

XLSX.jl can read from an IO using either `readxlsx`, `readto` or `openxlsx`. Something like:

```julia-auto
buf = IOBuffer(HTTP.get(url).body)

df = XLSX.readto(buf, DataFrame)

```
