# Limited parallel downloads

**URL:** <https://discourse.julialang.org/t/limited-parallel-downloads/71146>\
**Category:** Web Stack\
**Created:** [November 8, 2021, 12:52pm UTC](https://discourse.julialang.org/t/limited-parallel-downloads/71146 "2021-11-08T12:52:27Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![robertdj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robertdj/32/103_2.png) [@robertdj](https://discourse.julialang.org/u/robertdj)\
**Post date:** [November 8, 2021, 12:52pm UTC](https://discourse.julialang.org/t/limited-parallel-downloads/71146/1 "2021-11-08T12:52:27Z")

</div>

I want to make a large number of REST calls to download files. Each call is slow and I can only make a certain number of concurrent downloads (say `N`).

I would like to max out on the number of concurrent calls. I imagine making a new call as soon as one of the existing calls finishes.

From my limited understanding of Channels they sound like a way to go, but cannot figure out how to implement my idea. I have looked at [ThreadPools](https://github.com/tro3/ThreadPools.jl) and topics [like this one](https://discourse.julialang.org/t/basic-examples-of-tasks-channels-or-more-verbose-documentation/6797).  
Currently it is only downloading and not parsing, so I think coroutines are sufficient.

Any input is appreciated. Let me know if I should provide more info.

Thanks!

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [November 8, 2021, 2:21pm UTC](https://discourse.julialang.org/t/limited-parallel-downloads/71146/2 "2021-11-08T14:21:12Z")

</div>

What is a ‘large number’ ? Which OS Are you using - if Linux/Unix check on your limits.

There are limits on the max open files (connections) on a system. I learnt this the hard way (long story)  
You can tune these limits higher, and also alter the values of tcp\_tw\_recycle and tcp\_tw\_reuse which mean a closed connection can be recycled faster.

---

<div class="post-metadata">

**Author:** ![fredrikekre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fredrikekre/32/1688_2.png) [@fredrikekre](https://discourse.julialang.org/u/fredrikekre)\
**Post date:** [November 8, 2021, 2:49pm UTC](https://discourse.julialang.org/t/limited-parallel-downloads/71146/3 "2021-11-08T14:49:54Z")

</div>

Here is one example:

```julia
import Downloads

urls = ["https://julialang.org" for _ in 1:100]
max_concurrent = 4
jobs = Channel(max_concurrent)

@sync begin
    @async begin
        # create the jobs
        foreach(url -> put!(jobs, url), urls)
        close(jobs)
    end
    for _ in 1:max_concurrent
        # spawn max_concurrent tasks
        @async for url in jobs
            r = Downloads.request(url)
            @show r.status
        end
    end
end

```

---

<div class="post-metadata">

**Author:** ![lbilli](https://avatars.discourse-cdn.com/v4/letter/l/59ef9b/32.png) [@lbilli](https://discourse.julialang.org/u/lbilli)\
**Post date:** [November 8, 2021, 3:57pm UTC](https://discourse.julialang.org/t/limited-parallel-downloads/71146/4 "2021-11-08T15:57:11Z")

</div>

Doesn’t `asyncmap` accomplish something similar?

```julia
asyncmap(urls, ntasks=max_concurrent) do u
                                        r = Downloads.request(u)
                                        @show r.status
                                        r
                                      end

```

---

<div class="post-metadata">

**Author:** ![robertdj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robertdj/32/103_2.png) [@robertdj](https://discourse.julialang.org/u/robertdj)\
**Post date:** [November 8, 2021, 3:59pm UTC](https://discourse.julialang.org/t/limited-parallel-downloads/71146/5 "2021-11-08T15:59:03Z")

</div>

Thanks for you answer – it looks like it sort of does what I want 🙂

But it appears that it makes chunks of `max_concurrent` async calls and then waits until they are all done before moving on to the next chunk, at least if I add a `sleep`:

```julia
        @async for url in jobs
            sleep(1)
            r = Downloads.request(url)
            @show r.status
        end

```

Is it possible to pick up entry `max_concurrent + 1` from `jobs` as soon as the first download succeeds?

Also, I don’t understand why the second loop

```julia
    for _ in 1:max_concurrent
        # spawn max_concurrent tasks
        @async for url in jobs
            r = Downloads.request(url)
            @show r.status
        end
    end

```

actually behaves the way it does. Omitting the outer loop and only keeping

```julia
        @async for url in jobs
            r = Downloads.request(url)
            @show r.status
        end

```

it appears to download all the `url`s in one go. How is the outer loop alleviating this? It would be great if you can elaborate on this 🙂

**Edit** : Even if `asyncmap` can solve the problem at hand I would still love to learn more about tasks/channels if you have time to explain 🙂

---

<div class="post-metadata">

**Author:** ![robertdj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robertdj/32/103_2.png) [@robertdj](https://discourse.julialang.org/u/robertdj)\
**Post date:** [November 8, 2021, 4:00pm UTC](https://discourse.julialang.org/t/limited-parallel-downloads/71146/6 "2021-11-08T16:00:08Z")

</div>

Probably 10\_000 calls and my number of max concurrent calls are in the 10s. I don’t know if this gives problems with the OS (Win/Linux).

---

<div class="post-metadata">

**Author:** ![robertdj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robertdj/32/103_2.png) [@robertdj](https://discourse.julialang.org/u/robertdj)\
**Post date:** [November 8, 2021, 4:01pm UTC](https://discourse.julialang.org/t/limited-parallel-downloads/71146/7 "2021-11-08T16:01:10Z")

</div>

From the docs it appears to do exactly what I want – thanks!

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [November 8, 2021, 4:09pm UTC](https://discourse.julialang.org/t/limited-parallel-downloads/71146/8 "2021-11-08T16:09:09Z")

</div>

10 concurrent calls will be fine.
