# How to make it be a memory leak program(without lower function such as GC.xxx used)

**URL:** <https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347>\
**Category:** General Usage\
**Tags:** question\
**Created:** [October 27, 2022, 7:11am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347 "2022-10-27T07:11:06Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![HANHAOHAN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hanhaohan/32/36366_2.png) [@HANHAOHAN](https://discourse.julialang.org/u/HANHAOHAN)\
**Post date:** [October 27, 2022, 7:11am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/1 "2022-10-27T07:11:07Z")

</div>

Again and again, the memory occupation grows to 80%(about 12G) slowly as my program running.

My program is simple, as follows:

```julia
using Images
using HTTP
using libwebp_jll
using ImageView
using CSV
using ProgressMeter

function get_html(url)::Vector{UInt8}
    html = HTTP.get(url, 
        [("User-Agent", "Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.36 SE 2.X MetaSr 1.0")]; 
        connect_timeout = 30,
        connection_limit = 20,
        readtimeout = 3,
        ).body
end

function get_max_img(url)
    html = get_html(url) |> String
    imgs_t = [@async read_webimg(x.match) for x in eachmatch(r"(https?:[\S]*?.(jpg|jpeg|png|gif|bmp|bytes)(?=\ ))", html)]
    append!(imgs_t, [@async read_webimg(complete(url, x.captures[1])) for x in eachmatch(r"""<img [^>]*?src *?= *?"(.*?)(?=")""", html)])
    # @show typeof(imgs_t[1])
    imgs = map(fetch, imgs_t)
    if length(imgs)>0
        sort!(imgs; by=x->prod(size(x)), rev=true)
        imgs[1]
    else
        zeros(RGB{N0f8}, (0,0))
    end
end

function main(i, url)
    try
        img = get_max_img(String(url))

        if prod(size(img)) > 100
            save("target_max_img/$(i).jpg", img)
        end
    catch e
        @show e
    end
    i
end
     
data = [...]

task_running = Dict{Int, Task}()

for (i, d) in data
    task_running[i] = @async main(i, d.url)
    next!(p)
    while length(task_running) > 100
        yield()
        filter!(x->!(istaskdone(x[2])||istaskfailed(x[2])), task_running)
        println(length(task_running))
    end
end

```

Can anyone tell me where is the bug?

Is there a way to create a memory-leak program, with language with GC?

---

<div class="post-metadata">

**Author:** ![digital\_carver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/digital_carver/32/33818_2.png) [@digital\_carver](https://discourse.julialang.org/u/digital_carver)\
**Post date:** [October 27, 2022, 9:09am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/2 "2022-10-27T09:09:19Z")

</div>

It looks like you’re downloading lots of images from lots of pages, and have them in memory concurrently. That’s most likely the cause for high memory usage.  
You can try printing the size of `imgs` after the `map` in `get_max_img` - like `Base.summarysize(imgs)` - and see if those add up to the gigabytes of memory usage you see.

Instead of getting all the images and sorting them, you can fetch them one at a time, compare them to the current `max_img`’s size (initializing `max_img` to `zeros(RGB{N0f8}, (0,0))` first), and set `max_img = current_img` if the current image has a larger size. This way, you don’t need to have more than one image (per `main` call) in memory at a time. That should help reduce the memory usage a lot.

---

<div class="post-metadata">

**Author:** ![digital\_carver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/digital_carver/32/33818_2.png) [@digital\_carver](https://discourse.julialang.org/u/digital_carver)\
**Post date:** [October 27, 2022, 9:16am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/3 "2022-10-27T09:16:56Z")

</div>

Something like

```julia
function get_max_img(url)
    html = get_html(url) |> String
    imgs_t = [@async read_webimg(x.match) for x in eachmatch(r"(https?:[\S]*?.(jpg|jpeg|png|gif|bmp|bytes)(?=\ ))", html)]
    append!(imgs_t, [@async read_webimg(complete(url, x.captures[1])) for x in eachmatch(r"""<img [^>]*?src *?= *?"(.*?)(?=")""", html)])
    max_img = zeros(RGB{N0f8}, (0,0))
    for img in Iterators.map(fetch, imgs_t)
         if length(img) > length(max_img)
             max_img = img
         end
    end
    max_img
end

```

Note: untested code.

`length(x)` is equivalent here to `size(prod(x))` btw.

---

<div class="post-metadata">

**Author:** ![HANHAOHAN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hanhaohan/32/36366_2.png) [@HANHAOHAN](https://discourse.julialang.org/u/HANHAOHAN)\
**Post date:** [October 27, 2022, 9:21am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/4 "2022-10-27T09:21:09Z")

</div>

Thanks a lot!  
Seems correct. There are too many images opened in memory. And they are in memory concurrently!

Thanks again.

I should get the size of them one by one to reduce memory consume.

---

<div class="post-metadata">

**Author:** ![HANHAOHAN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hanhaohan/32/36366_2.png) [@HANHAOHAN](https://discourse.julialang.org/u/HANHAOHAN)\
**Post date:** [October 28, 2022, 9:37am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/5 "2022-10-28T09:37:15Z")

</div>

I removed the code decoding images in memory, and replace `prod(size(img))` with `parse(Int64, Dict(reply.headers)["Content-Length"])`, and replace `sort!(...)` with `findmax(...)`, and replace each `@async` with `Threads.@spawn` too.

It works at the beginning, about 3G memory occupied, but, 11G memory occupied again by now.

Notes: I reduced the argument 100 in `if prod(size(img)) > 100` to 30.

Let every page have 100 images, each image with size 1M, the memory size should not be larger than 4G, right?

---

<div class="post-metadata">

**Author:** ![HANHAOHAN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hanhaohan/32/36366_2.png) [@HANHAOHAN](https://discourse.julialang.org/u/HANHAOHAN)\
**Post date:** [October 28, 2022, 10:22am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/6 "2022-10-28T10:22:28Z")

</div>

![截图 2022-10-28 18-21-47](https://global.discourse-cdn.com/julialang/original/3X/c/9/c962757b3ba3495362a54efc067058023819066c.png)

---

<div class="post-metadata">

**Author:** ![digital\_carver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/digital_carver/32/33818_2.png) [@digital\_carver](https://discourse.julialang.org/u/digital_carver)\
**Post date:** [October 28, 2022, 11:15am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/7 "2022-10-28T11:15:22Z")

</div>

Could you show the current version of your code? Including that of `read_webimg`.

---

<div class="post-metadata">

**Author:** ![HANHAOHAN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hanhaohan/32/36366_2.png) [@HANHAOHAN](https://discourse.julialang.org/u/HANHAOHAN)\
**Post date:** [October 31, 2022, 1:23am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/8 "2022-10-31T01:23:23Z")

</div>

```julia
# using Images
using HTTP
# using libwebp_jll
# using ImageView
using CSV
using ProgressMeter

function get_html(url)
    r = HTTP.get(url, 
        [("User-Agent", "Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.36 SE 2.X MetaSr 1.0")]; 
        connect_timeout = 5,
        connection_limit = 1000,
        readtimeout = 3,
        status_exception = false
        )
    @assert isa(r.body, Vector{UInt8}) "body is not Vector{UInt8}!!!\n"

    r
end

function webp_decode(bytes::Vector{UInt8})::Matrix{RGB{N0f8}}
    height = [0]
    width = [0]
    ccall((:WebPGetInfo, "libwebp"), Int, (Ptr{UInt8}, Csize_t, Ptr{Int}, Ptr{Int}), Base.unsafe_convert(Ptr{UInt8},bytes), length(bytes), Ptr{Int}(Base.unsafe_convert(Ptr{Int64}, width)), Ptr{Int}(Base.unsafe_convert(Ptr{Int64}, height)))
    x = ccall((:WebPDecodeRGB, "libwebp"), Ptr{UInt8}, (Ptr{UInt8}, Csize_t, Ptr{Int}, Ptr{Int}), Base.unsafe_convert(Ptr{UInt8},bytes), length(bytes), Ptr{Int}(Base.unsafe_convert(Ptr{Int64}, width)), Ptr{Int}(Base.unsafe_convert(Ptr{Int64}, height)))
    img = [RGB(reinterpret(N0f8,unsafe_load(x, 3*(width[1]*i+j)+1)),reinterpret(N0f8,unsafe_load(x, 3*(width[1]*i+j)+2)),reinterpret(N0f8,unsafe_load(x, 3*(width[1]*i+j)+3)),) for (i,j) in Iterators.product(0:height[1]-1, 0:width[1]-1)];
end

# function open_webp(fpath::String)::Matrix{RGB{N0f8}}
# bytes = read(fpath);
# webp_decode(bytes)
# end

function read_webimg(url::AbstractString, io::IO, iolock) #::Matrix{RGB{N0f8}}
    global lk
    r = get_html(url)
    headers = Dict(r.headers)
    type = headers["Content-Type"]
    content_len = haskey(headers, "Content-Length") ? parse(Int64, headers["Content-Length"]) : 0
    # @assert r.status == 200 "r.status != 200 $(url)"
    # @assert startswith(type, "image/") """startswith(type, "image/") is false"""

    if r.status == 200 && startswith(type, "image/")
        type = split(type, "/")[end]
        bts = r.body
        # img = type == "webp" ? webp_decode(bts) : load(bts |> IOBuffer)
        # @assert size(img)[1] in 200:1600 && size(img)[2] in 200:1600 && max(size(img)...) /min(size(img)...) < 3 "图片大小不符合要求 $(size(img))"
        # if size(img)[1] in 200:1600 && size(img)[2] in 200:1600 && max(size(img)...) /min(size(img)...) < 3
            return content_len, type, bts
        # end
    end
    while !trylock(iolock)
        sleep(0.001)
    end
    println(io, "\t", url)
    unlock(iolock)
    0, "", Vector{UInt8}()
end

complete(url_main, url) = startswith(url, "http") ? url : (startswith(url, "//") ? "https:"*url : match(r"https?:\/\/.*?(?=\/)", url_main).match * url)

function main(i, url)
    global lk
    global io

    buflock = ReentrantLock()

    buf = IOBuffer()
    println(buf, url)
    try
        html = get_html(String(url))
        if html.status == 200
            html = html.body |> String

            imgs_t = [Threads.@spawn read_webimg(x.match, buf, buflock) for x in eachmatch(r"(https?:[\S]*?.(jpg|jpeg|png|gif|bmp|bytes)(?=\ ))", html)]
            append!(imgs_t, [Threads.@spawn read_webimg(complete(url, x.captures[1]), buf, buflock) for x in eachmatch(r"""<img [^>]*?src *?= *?"(.*?)(?=")""", html)])
            
            imgs = map(fetch, imgs_t)
            
            (sz, index) = length(imgs) > 0 ? findmax(x->x[1], imgs) : (0, 0)
        
            if sz>2000
                res = imgs[index]
                open("""target_max_img/$(i).$(res[2])""", "w") do fo
                    write(fo, res[3])
                end
            else

            end
        end
    catch e
        println(buf, "\t", e)
    end
    while !trylock(lk)
        sleep(0.01)
    end
    println(io, buf.data |> String)
    flush(io)
    unlock(lk)
end

io = open("log_down_webp", "w")

lk = ReentrantLock()
# save("test.jpg", read_webimg("https://www.thisisnotdietfood.com/wp-content/uploads/2019/06/baconcheeseburgergrilledcheesecasserole-13-min-min.jpg"))
csv_data = CSV.Rows("origin_data.csv")
urls_todo = enumerate(csv_data)

# main(1, urls_todo[1][2].targetUrl)
len = 0
for (i,row) in urls_todo 
    global len = i 
end

p = Progress(len)

task_running = Dict{Int, Task}()

for (i, row) in urls_todo
    # main(i, row.targetUrl)
    task_running[i] = Threads.@spawn main(i, row.targetUrl)
    next!(p)
    while length(task_running) > 30
        for (k, t) in task_running
            if istaskfailed(t)
                @show fetch(t)
            end
        end
        # yield()
        sleep(1)
        filter!(x->!(istaskdone(x[2])||istaskfailed(x[2])), task_running)
        # println(length(task_running))
    end
end

close(io)

```

Copy.

Terminate, by os, during running.

---

<div class="post-metadata">

**Author:** ![HANHAOHAN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hanhaohan/32/36366_2.png) [@HANHAOHAN](https://discourse.julialang.org/u/HANHAOHAN)\
**Post date:** [November 2, 2022, 1:52am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/9 "2022-11-02T01:52:18Z")

</div>

The memory leak still!  
I tried with Channel instead with `@async` for each, and got rid of download images, just get from url and judge the response status compare with 200.

So, the bug generate from `HTTP`.

```julia
using HTTP
using CSV
using ProgressMeter
using DataFrames

function can_open(cin::Channel, outputfile::String, df::DataFrame)
    while isopen(cin)
        id, d = take!(cin)
        try
            if length(df[!,"output"]) > 10
                CSV.write(outputfile, df; append=true)
                empty!(df)
            end

            r = HTTP.get(String(d.targetUrl), 
                [("User-Agent", "Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.36 SE 2.X MetaSr 1.0")]; 
                connect_timeout = 10,
                connection_limit = 100,
                readtimeout = 5,
                status_exception = false
                )
            if r.status == 200
                push!(df, [id, d.title1, d.title2, d.productName, d.targetUrl, d.output])
            end
            
        catch e
            @show e
        end
    end
end

ordered_data = enumerate(CSV.File("origin_data.csv"))
len = 0
for (i,row) in ordered_data 
    global len = i 
end
@show len

cin = Channel(1000)
df = DataFrame("idex"=>[], "title1"=>[], "title2"=>[], "productName"=>[], "targetUrl"=>[],"output"=>[])
CSV.write("available.csv", df; append=false)

for i in 1:100
    @async can_open(cin, "available.csv", df)
end

p = Progress(len)
for (i, d) in ordered_data
    put!(cin, (i,d))
    next!(p)
end

while !isempty(cin)
    sleep(1)
end

close(cin)

```

 ![截图 2022-11-02 09-53-50](https://global.discourse-cdn.com/julialang/original/3X/b/5/b5f218d3c6bf30c6df50f51adb2cf93e68bf6bb8.png)

---

<div class="post-metadata">

**Author:** ![HANHAOHAN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hanhaohan/32/36366_2.png) [@HANHAOHAN](https://discourse.julialang.org/u/HANHAOHAN)\
**Post date:** [November 2, 2022, 2:08am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/10 "2022-11-02T02:08:54Z")

</div>

Memory occupy(the program) is about 1.6G, after replace `r = HTTP.get(...); if r.status == 200` with `if true`.

So, it is caused by `HTTP`? Other wise, `try catch`?

---

<div class="post-metadata">

**Author:** ![HANHAOHAN](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hanhaohan/32/36366_2.png) [@HANHAOHAN](https://discourse.julialang.org/u/HANHAOHAN)\
**Post date:** [November 4, 2022, 7:53am UTC](https://discourse.julialang.org/t/how-to-make-it-be-a-memory-leak-program-without-lower-function-such-as-gc-xxx-used/89347/11 "2022-11-04T07:53:07Z")

</div>

I think I’v found the problem.  
Version 1: create Task for each.  
Version 2: not lock the ops on global var and io.
