# Downloading website content (HTTP, BufferedStreams)

**URL:** <https://discourse.julialang.org/t/downloading-website-content-http-bufferedstreams/112595>\
**Category:** General Usage\
**Tags:** question, strings, unicode, streaming\
**Created:** [April 6, 2024, 1:14am UTC](https://discourse.julialang.org/t/downloading-website-content-http-bufferedstreams/112595 "2024-04-06T01:14:43Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Anton\_Hinneck](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/anton_hinneck/32/22716_2.png) [@Anton\_Hinneck](https://discourse.julialang.org/u/Anton_Hinneck)\
**Post date:** [April 6, 2024, 1:14am UTC](https://discourse.julialang.org/t/downloading-website-content-http-bufferedstreams/112595/1 "2024-04-06T01:14:43Z")

</div>

Hi,  
I would like to download HTML from urls. For this purpose I use

- HTTP.jl and
- BufferedStreams.jl.

```julia

using HTTP
using BufferedStreams

url = "https://discourse.julialang.org/"

htmlData = String[]

io = nothing
HTTP.open("GET", url) do http
    io = BufferedInputStream(http)
    while !eof(http)
        push!(htmlData, read(http, String))
    end
end

htmlData[1]

```

With this code I get the following data:  
"\x1f\x8b\b\0\0\0\0\0\0\x03\xcd\\xedr … "

This looks like there is some encoding problem.

I tinkered around with StringEncodings.jl but could not get it to work.

How can I retrieve clean HTML?

---

<div class="post-metadata">

**Author:** ![jules](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@jules](https://discourse.julialang.org/u/jules)\
**Post date:** [April 6, 2024, 6:54am UTC](https://discourse.julialang.org/t/downloading-website-content-http-bufferedstreams/112595/2 "2024-04-06T06:54:54Z")

</div>

Maybe try [Downloads · The Julia Language](https://docs.julialang.org/en/v1/stdlib/Downloads/#Downloads.download)

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [April 6, 2024, 9:12am UTC](https://discourse.julialang.org/t/downloading-website-content-http-bufferedstreams/112595/3 "2024-04-06T09:12:25Z")

</div>

My guess is, you get a gzip stream.  
There are some more flaws in your code, e.g. you don’t use your `io`. Probably because of some experiments during your error finding process.  
Last but not least, HTTP typically return the body as a Vector{UInt8}.

Try this code:

```julia
using CodecZlib, HTTP, BufferedStreams

url = "https://discourse.julialang.org/"

htmlData = UInt8[]
HTTP.open("GET", url) do http
    io = GzipDecompressorStream(BufferedInputStream(http))
    while !eof(io)
        push!(htmlData, read(io, UInt8))
    end
end
body=String(htmlData)

```

If there is no special reason for using a `BufferedInputStream` you can have it much easier:

```julia
using HTTP

url = "https://discourse.julialang.org/"

res=HTTP.request("GET",url)
body=String(res.body)

```

---

<div class="post-metadata">

**Author:** ![abraemer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/abraemer/32/51403_2.png) [@abraemer](https://discourse.julialang.org/u/abraemer)\
**Post date:** [April 6, 2024, 8:09pm UTC](https://discourse.julialang.org/t/downloading-website-content-http-bufferedstreams/112595/4 "2024-04-06T20:09:47Z")

</div>

> [@oheil](#):
>
> My guess is, you get a gzip stream.

Yep that’s right. The first 4 bytes are the [magic bytes of gzip](https://en.m.wikipedia.org/wiki/Gzip#File_format)
