# Reading jpg images for Kaggle challenge

**URL:** <https://discourse.julialang.org/t/reading-jpg-images-for-kaggle-challenge/114772>\
**Category:** Tooling\
**Tags:** question, package\
**Created:** [May 27, 2024, 9:40am UTC](https://discourse.julialang.org/t/reading-jpg-images-for-kaggle-challenge/114772 "2024-05-27T09:40:18Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Smara\_Kazenango](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/smara_kazenango/32/130333_2.png) [@Smara\_Kazenango](https://discourse.julialang.org/u/Smara_Kazenango)\
**Post date:** [May 27, 2024, 9:40am UTC](https://discourse.julialang.org/t/reading-jpg-images-for-kaggle-challenge/114772/1 "2024-05-27T09:40:18Z")

</div>

How can I read jpg images in Julia.

Has anyone performed any work on the following dataset: ([Chest X-Ray (Pneumonia,Covid-19,Tuberculosis) | Kaggle](https://www.kaggle.com/datasets/jtiptj/chest-xray-pneumoniacovid19tuberculosis))

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [May 27, 2024, 11:13am UTC](https://discourse.julialang.org/t/reading-jpg-images-for-kaggle-challenge/114772/2 "2024-05-27T11:13:58Z")

</div>

I suppose this is it, with docs how: [GitHub - JuliaIO/JpegTurbo.jl: Julia interface to libjpeg-turbo](https://github.com/JuliaIO/JpegTurbo.jl)

It’s the fastest option/backend (also faster then MATLAB and Python/OpenCV): [benchmark results against other backends · Issue #15 · JuliaIO/JpegTurbo.jl · GitHub](https://github.com/JuliaIO/JpegTurbo.jl/issues/15#issuecomment-1732963159)

In case you looked no further than Julia’s docs, then that’s usually a mistake, it provides e.g. LinearAlgebra and some other math, but look at [juliahub.com](http://juliahub.com) for most other stuff, also googling e.g. JPEG Julia would have worked.

FYI: This is probably the first package for JPEG/image decoding/encoding, and supports all kinds of formats including PNG you need: [GitHub - JuliaIO/ImageMagick.jl: Thin Wrapper for the library ImageMagick](https://github.com/JuliaIO/ImageMagick.jl)

I doubt you find _pure_ Julia code for JPEG (unlike e.g. available for the simpler GIF format), but the non-pure dependencies shouldn’t be any problem.

There might be one or more packages for this, be aware of FileIO.jl and Images.jl, but I think this is all you need to know.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [May 27, 2024, 12:54pm UTC](https://discourse.julialang.org/t/reading-jpg-images-for-kaggle-challenge/114772/3 "2024-05-27T12:54:43Z")

</div>

> [@Palli](#):
>
> There might be one or more packages for this, be aware of [FileIO.jl](https://juliahub.com/ui/Packages/General/FileIO) and [Images.jl](https://juliahub.com/ui/Packages/General/Images), but I think this is all you need to know.

GDAL

Packages using GDAL can read georeferenced jpeg2000 and GeoTIFF, which those other packages cannot.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [May 27, 2024, 1:05pm UTC](https://discourse.julialang.org/t/reading-jpg-images-for-kaggle-challenge/114772/4 "2024-05-27T13:05:44Z")

</div>

Thanks, good to know, JPEG2000 is a completely different file format, though people might not realize, so I wasn’t answering for that, since not asked (nor did I now locate it used in the dataset, only JPEG and PNG). TIFF is also completely different (a lossless format, and not the best one, PNG was best, some other even better), and GeoTIFF seems to only apply to that format. Maybe something similar applies for JPEG[2000].

JPEG XL is likely the best JPEG variant (but not for the web, since Google Chrome deprecated it), even best lossless compressed format of all. HEIF is another one, and AVIF maybe best for the web supported by vast majority. The latter (and many more e.g JPEG2000 and PNG and WebP, also video and audio formats) is supported by FileTypes.jl and yet another “jxr”, seemingly the other JPEG variant JPEG XR, but not latest XL.

> **[AVIF image format | Can I use... Support tables for HTML5, CSS3, etc](https://caniuse.com/avif)**
>
> "Can I use" provides up-to-date browser support tables for support of front-end web technologies on desktop and mobile web browsers.

[https://www.fsf.org/blogs/community/googles-decision-to-deprecate-jpeg-xl-emphasizes-the-need-for-browser-choice-and-free-formats](https://www.fsf.org/blogs/community/googles-decision-to-deprecate-jpeg-xl-emphasizes-the-need-for-browser-choice-and-free-formats)

> **[Google's Jpegli Offers ~35% Compression Improvement For High Quality JPEGs](https://www.phoronix.com/news/Google-Jpegli-Library)**
>
> The Google Open-Source Blog today announced Jpegli, a JPEG coding library for encode/decode that maintains backwards compatibility with JPEG while offering around a 35% compression ratio improvement for high quality JPEG compression.

I can’t locate AVIF support for Julia so this may need to be wrapped:

> **[GitHub - AOMediaCodec/libavif: libavif - Library for encoding and decoding...](https://github.com/AOMediaCodec/libavif)**
>
> libavif - Library for encoding and decoding .avif files - AOMediaCodec/libavif

I’m unclear which language or library supports AVIF best, maybe rather Rust (an option to be wrapped or just called):  
[https://crates.io/crates/image/0.24.8](https://crates.io/crates/image/0.24.8)

Python also supports it and has libraries for it available, and more, and it’s easy to use all of them with PythonCall.jl.

Absolute most compression is likely to be with neural-network based methods, e.g. HiFiC (and it needs to be wrapped for Julia):

> **[HiFiC - High-Fidelity Generative Image Compression](https://hific.github.io/)**
>
> We combine Generative Adversarial Networks with learned compression to obtain a state-of-the-art generative lossy compression system.

And even it can be improved upon (e.g. down to “𝐻-𝑅𝑂𝐼: 0.2275𝑏𝑝𝑝” (their example); is possible; even without it down to 0.139bbp possible with good quality, I guess depending on the image):

> **[Super-High-Fidelity Image Compression via Hierarchical-ROI and Adaptive...](https://arxiv.org/abs/2403.13030)**
>
> Learned Image Compression (LIC) has achieved dramatic progress regarding objective and subjective metrics. MSE-based models aim to improve objective metrics while generative models are leveraged to improve visual quality measured by subjective...

> Further, we propose adaptive quantization by non-linear mapping within the channel dimension to constrain the bit rate while maintaining the visual quality. Exhaustive experiments demonstrate that our methods achieve better visual quality on small faces and text with lower bit rates, e.g., 0.7X bits of HiFiC and 0.5X bits of BPG.
