# Handling pre-parsed XLSX file

**URL:** <https://discourse.julialang.org/t/handling-pre-parsed-xlsx-file/76455>\
**Category:** Data\
**Tags:** xlsx\
**Created:** [February 14, 2022, 7:11pm UTC](https://discourse.julialang.org/t/handling-pre-parsed-xlsx-file/76455 "2022-02-14T19:11:28Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jeremiedb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeremiedb/32/29150_2.png) [@jeremiedb](https://discourse.julialang.org/u/jeremiedb)\
**Post date:** [February 14, 2022, 7:11pm UTC](https://discourse.julialang.org/t/handling-pre-parsed-xlsx-file/76455/1 "2022-02-14T19:11:29Z")

</div>

Package [XLSX.jl](https://github.com/felipenoris/XLSX.jl) provides tool to import XLSX file from path name: `readxlsx`.

However, I struggle handle an XLSX that is first parsed by the file upload functionality in Dash.jl. The resulting parsing from the upload is a string that looks like:

```julia
data:application/vnd.openxmlformats-officedocument.spreadsheetml.sheet;base64,UEsDBBQABgAIAAAAIQBi7p1oXg...

```

What I was first hoping to do was to directly save that parsed content into a `.xlsx` that could then be read by `readxlsx`. However, resulting file content is no longer in a valid format. The valid `xlsx` format should have looked like:

```julia
PK ! bîh^  [Content_Types].xml ¢( ¬”ËNÃ0E÷HüCä-JÜ²@5í‚Ç*Q>ÀÄ“Æªc[žiiÿž‰ûB¡j7±ÏÜ{2ñÍh²nm¶‚ˆÆ»R‹ÈÀU^e7/ÅÇì%¿’rZYï e@1__f› ˜q·ÃR4DáAJ¬h>€ãÚÇVßÆ¹ªZ¨9ÈÛÁàNVÞ8Ê©ÓãÑÔji){^óã-I‹"{Üv^¥P!XS)bR¹rú—K¾s(¸3Õ`cÞ0†½ÝÎß»¾7M4²©ŠôªZÆk+¿|\|z¿(Ž‹ôPúº6h_-[ž@!‚ÒØ Pk‹´­2nÏ}Ä?£LËðÂ Ýû%áÄßedºždN"m,à¥ÇžDO97*‚~§ÈÉ¸8ÀOíc|n¦ÑaäEøÿöéºóÀBÉÀ!$}‡íàÈé;{ìÐå[ƒîñ–é2þ ÿÿ PK ! µU0#ô L _rels/.rels ¢( ¬’MOÃ0†ïHü‡È÷ÕÝBKwAH»!T~€IÜµ£$eÝ¿'TƒG½~üÊÛÝ<êÈ!öâ4¬‹;#¶w­†—úqua*&r–Fq¬áÄvÕõÕö™GJy(v½*«¸¨¡KÉß#FÓñD±Ï.W	¥†=™ZÆMYÞbø®ÕBSí­†°·7 ê“Ï›×–¦é
?ˆ9LìÒ™ÈsbgÙ®|

```

Is there a way to get that first parsed string into a valid format, compatbiel with `XLSX.readxlsx`?

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [February 14, 2022, 7:31pm UTC](https://discourse.julialang.org/t/handling-pre-parsed-xlsx-file/76455/2 "2022-02-14T19:31:48Z")

</div>

I think you just need to `base64decode` everything that follows `base64,`:

```nohighlight
julia> using Base64

julia> String(base64decode("UEsDBBQABgAIAAAAIQBi7p1o"))
"PK\x03\x04\x14\0\x06\0\b\0\0\0!\0b\xee\x9dh"

```

---

<div class="post-metadata">

**Author:** ![jeremiedb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeremiedb/32/29150_2.png) [@jeremiedb](https://discourse.julialang.org/u/jeremiedb)\
**Post date:** [February 14, 2022, 9:22pm UTC](https://discourse.julialang.org/t/handling-pre-parsed-xlsx-file/76455/3 "2022-02-14T21:22:48Z")

</div>

Thanks! By writing a file that only contains the decoded base64 from the parsed string, it results in a valid xslx file:

```julia
    # assuming the file was parsed in the variable `input`
    content_start = findfirst(";base64,", input)[end] + 1 # finds where the base 64 content starts
    raw = base64decode(input[content_start:end])
    open(file.xlsx, "w") do io
        write(io, raw)
    end

```
