# Parallel bzip2 ang gzip: pbzip2 and pigz

**URL:** <https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096>\
**Category:** Optimization (Mathematical)\
**Tags:** jump\
**Created:** [September 11, 2022, 7:53pm UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096 "2022-09-11T19:53:44Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Stuart\_Rogers](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stuart_rogers/32/10694_2.png) [@Stuart\_Rogers](https://discourse.julialang.org/u/Stuart_Rogers)\
**Post date:** [September 11, 2022, 7:53pm UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/1 "2022-09-11T19:53:44Z")

</div>

Do JuMP’s write\_to\_file and read\_from\_file use pbzip2 and pigz, parallel implementations of bzip2 and gzip, when writing or reading models stored in .bz2 and .gz compressed files?

[https://jump.dev/JuMP.jl/stable/reference/models/#JuMP.write\_to\_file](https://jump.dev/JuMP.jl/stable/reference/models/#JuMP.write_to_file)

> **[Parallel BZIP2 (PBZIP2)](http://compression.ca/pbzip2/)**
>
> Parallel BZIP2 (pbzip2), a parallel implementation of the bzip2 file compressor

[https://zlib.net/pigz/](https://zlib.net/pigz/)

---

<div class="post-metadata">

**Author:** ![miles.lubin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/miles.lubin/32/279_2.png) [@miles.lubin](https://discourse.julialang.org/u/miles.lubin)\
**Post date:** [September 11, 2022, 8:35pm UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/2 "2022-09-11T20:35:52Z")

</div>

No, JuMP uses [GitHub - JuliaIO/CodecZlib.jl: zlib codecs for TranscodingStreams.jl.](https://github.com/JuliaIO/CodecZlib.jl) and [GitHub - JuliaIO/CodecBzip2.jl: A bzip2 codec for TranscodingStreams.jl.](https://github.com/JuliaIO/CodecBzip2.jl), which in turn use the standard zlib and libbzip2.

---

<div class="post-metadata">

**Author:** ![Stuart\_Rogers](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stuart_rogers/32/10694_2.png) [@Stuart\_Rogers](https://discourse.julialang.org/u/Stuart_Rogers)\
**Post date:** [September 11, 2022, 9:01pm UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/3 "2022-09-11T21:01:28Z")

</div>

pbzip2 and pigz ought to be much faster than bzip2 and gzip since they utilize multiple cores through multithreading. Are there plans for TranscodingStreams.jl to use pbzip2 and pigz instead?

---

<div class="post-metadata">

**Author:** ![odow](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/odow/32/28685_2.png) [@odow](https://discourse.julialang.org/u/odow)\
**Post date:** [September 11, 2022, 9:12pm UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/4 "2022-09-11T21:12:49Z")

</div>

No plans. Is the speed of writing our your file a bottleneck?

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [September 12, 2022, 8:44am UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/5 "2022-09-12T08:44:08Z")

</div>

For what it’s worth, gzip decompression is not very amenable to parallellization. (Compression is a different story.)

---

<div class="post-metadata">

**Author:** ![Stuart\_Rogers](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stuart_rogers/32/10694_2.png) [@Stuart\_Rogers](https://discourse.julialang.org/u/Stuart_Rogers)\
**Post date:** [September 12, 2022, 6:54pm UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/6 "2022-09-12T18:54:02Z")

</div>

My workflow is:

1. Construct the LP model in JuMP: 37 minutes.
2. Write the LP model to a MPS file using JuMP.write\_to\_file: 100 minutes.
3. Using 24 threads, presolve with PaPILO, solve with PDLP (1e-4 relative tolerance), postsolve with PaPILO: 145 minutes (PDLP takes 140 minutes to solve the presolved model).

Writing the MPS file takes 35% of the overall time. Writing a compressed MPS.GZ file takes about the same amount of time, but the compressed MPS.GZ file (1.2 GB) is 11% the size of the uncompressed MPS file (11 GB).

For another similarly-sized LP model, step 3 takes 50 minutes, so that writing the MPS file takes 53% of the overall time.

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [September 12, 2022, 8:14pm UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/7 "2022-09-12T20:14:11Z")

</div>

Have you profiled where the time is spent in step 2?

The only way to spend 100 minutes writing 11 GB (that’s 1.8 MB per second) is to have a very slow network disk but in that case the compressed writing should only require 11 minutes and spending 89 minutes compressing 11 GB of data sounds like entirely the wrong order of magnitude, even if it’s single-threaded.

My gut feeling is that writing or compressing+writing takes up a few minutes and the rest is spent on something else but I have no insights in the code so I can’t even guess what that might be. Profiling is the only way to find out where the time truly is spent.

An easy experiment is to write the uncompressed file to disk and then compress it with command line gzip. How much time does the latter step require? It should be in the same ballpark as writing the compressed file from Julia.

---

<div class="post-metadata">

**Author:** ![Stuart\_Rogers](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stuart_rogers/32/10694_2.png) [@Stuart\_Rogers](https://discourse.julialang.org/u/Stuart_Rogers)\
**Post date:** [September 12, 2022, 8:39pm UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/8 "2022-09-12T20:39:52Z")

</div>

I am using a fairly new Lambda workstation. This is how I measure the time of write\_to\_file:  
MPS\_fn = “/data/my.mps” # MPS filename. or MPS GZ filename: MPS\_fn = “/data/my.mps.gz”  
MPS\_time = @elapsed begin  
write\_to\_file(m,MPS\_fn)  
end

---

<div class="post-metadata">

**Author:** ![odow](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/odow/32/28685_2.png) [@odow](https://discourse.julialang.org/u/odow)\
**Post date:** [September 12, 2022, 8:40pm UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/9 "2022-09-12T20:40:34Z")

</div>

> 1. Construct the LP model in JuMP: 37 minutes.
> 2. Write the LP model to a MPS file using JuMP.write\_to\_file: 100 minutes.

I think we’ve had this conversation a couple of times, but JuMP might to be the best tool for the job. We don’t optimize for writing to a file. Part of the “write” is actually a “copy the entire model in memory at least once” which is probably part of the issue. I’ll have a think to see if there’s a way we could improve things.

---

<div class="post-metadata">

**Author:** ![odow](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/odow/32/28685_2.png) [@odow](https://discourse.julialang.org/u/odow)\
**Post date:** [September 12, 2022, 8:42pm UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/10 "2022-09-12T20:42:10Z")

</div>

> [@Stuart\_Rogers](#):
>
> Writing a compressed MPS.GZ file takes about the same amount of time

Yeah, that’s because the “write” isn’t timing only the write to file. It also has a bunch of overhead on the JuMP side to turn the problem into something that can be written to an MPS file (which involves a copy of the entire model), to make sure every variable and constraint has a unique name, to order the columns, etc. The issue isn’t the compression.

---

<div class="post-metadata">

**Author:** ![Stuart\_Rogers](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stuart_rogers/32/10694_2.png) [@Stuart\_Rogers](https://discourse.julialang.org/u/Stuart_Rogers)\
**Post date:** [September 20, 2022, 3:07am UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/11 "2022-09-20T03:07:17Z")

</div>

Where is the source code for JuMP.write\_to\_file?

Why doesn’t JuMP.write\_to\_file support xz compression, which is included in TranscodingStreams.jl?  
[https://jump.dev/JuMP.jl/stable/reference/models/#I/O](https://jump.dev/JuMP.jl/stable/reference/models/#I/O)

> **[GitHub - JuliaIO/TranscodingStreams.jl: Simple, consistent interfaces for any...](https://github.com/JuliaIO/TranscodingStreams.jl)**
>
> Simple, consistent interfaces for any codec. Contribute to JuliaIO/TranscodingStreams.jl development by creating an account on GitHub.

---

<div class="post-metadata">

**Author:** ![odow](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/odow/32/28685_2.png) [@odow](https://discourse.julialang.org/u/odow)\
**Post date:** [September 20, 2022, 3:35am UTC](https://discourse.julialang.org/t/parallel-bzip2-ang-gzip-pbzip2-and-pigz/87096/12 "2022-09-20T03:35:22Z")

</div>

`JuMP.write_to_file` is a thin wrapper:

> <https://github.com/jump-dev/JuMP.jl/blob/master/src/file_formats.jl>

around `MOI.write_to_file`:

> <https://github.com/jump-dev/MathOptInterface.jl/blob/be1c8c3b30cc3a27aa50c285b9b38940716d47ca/src/FileFormats/FileFormats.jl#L113-L117>

other extensions are possible, but need implementing:

> <https://github.com/jump-dev/MathOptInterface.jl/blob/be1c8c3b30cc3a27aa50c285b9b38940716d47ca/src/FileFormats/utils.jl#L190-L252>

PRs to improve things are welcome.

For your case, the better long-term outcome is probably to write a C interface to PDLP. Then you could go straight to the C library without having to read and write files.
