# Import pretrained Keras model in Julia for inference

**URL:** https://discourse.julialang.org/t/import-pretrained-keras-model-in-julia-for-inference/104924
**Category:** Machine Learning
**Created:** [October 13, 2023, 9:28am UTC](https://discourse.julialang.org/t/import-pretrained-keras-model-in-julia-for-inference/104924 "2023-10-13T09:28:04Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![dubosipsl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dubosipsl/32/52998_2.png) [@dubosipsl](https://discourse.julialang.org/u/dubosipsl)
#### Post date: [October 13, 2023, 9:28am UTC](https://discourse.julialang.org/t/import-pretrained-keras-model-in-julia-for-inference/104924/1 "2023-10-13T09:28:04Z")

</div>

Hello all,

I have trained a model in Python/Keras. What is the recommended way with the present Julia ML ecosystem to save it to disk and load it in Julia to do inference (no training) ?

A similar question was asked in 2020. One answer was “export the model to onnx”. Is this still a valid answer ? There are also recent similar questions but with PyTorch.

Thanks for your advice.

---

<div class="post-metadata">

### Author: ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)
#### Post date: [October 13, 2023, 11:44am UTC](https://discourse.julialang.org/t/import-pretrained-keras-model-in-julia-for-inference/104924/2 "2023-10-13T11:44:09Z")

</div>

ONNX is still your best bet

---

<div class="post-metadata">

### Author: ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)
#### Post date: [October 13, 2023, 3:20pm UTC](https://discourse.julialang.org/t/import-pretrained-keras-model-in-julia-for-inference/104924/3 "2023-10-13T15:20:50Z")

</div>

You could look at:

> **[GitHub - chengchingwen/Pickle.jl: An experimental package for loading and saving...](https://github.com/chengchingwen/Pickle.jl)**
>
> An experimental package for loading and saving object in Python Pickle format.

> loading and saving object in Python Pickle and Torch Pickle format.  
> […]  
> We also support loading/saving the tensor data from/for pytorch. […]

From the name of the package, it’s not obvious it has anything to do with neural networks, until you read further. Note, Pickle is Python’s general serialization format, and is not meant to be read by other languages, nor even to be used for all in Python, e.g. a potential security issue (for e.g. neural networks).

> <https://stackoverflow.com/questions/48295661/how-to-pickle-keras-model>

I don’t know in your specific case, in might not use Pickle, could though potentially be converted to. If you read such a file with Pickle.jl then there’s no security issue, since it doesn’t support all of pickle (and likely never will), but I understand enough for (some) neural networks.

You can actually read all Pickle files, and all Keras files by calling to CPython (i.e. using say PythonCall.jl), but then it’s not Julia, or Julia code doing it, only istructing Python to do it, and thus all that Python can do supported. Also the inference, but again you’re just using Python and Keras from Julia.

Some older I found, maybe also helpful:

> <https://stackoverflow.com/questions/65720584/how-to-load-python-pickle-from-julia>

> <https://gist.github.com/RobBlackwell/10a1aeabeb85bbf1a17cc334e5e60acf>

> [@jling](#):
>
> ONNX is still your best bet

That might still be true, or outdated. Do you know for sure contradicting the above, or didn’t know of it? I’m not sure it’s announced yet.

---

<div class="post-metadata">

### Author: ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)
#### Post date: [October 13, 2023, 3:46pm UTC](https://discourse.julialang.org/t/import-pretrained-keras-model-in-julia-for-inference/104924/4 "2023-10-13T15:46:16Z")

</div>

I was not considering PyCall-like solutions because then why doesn’t OP just use Python

---

<div class="post-metadata">

### Author: ![dubosipsl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dubosipsl/32/52998_2.png) [@dubosipsl](https://discourse.julialang.org/u/dubosipsl)
#### Post date: [October 13, 2023, 10:30pm UTC](https://discourse.julialang.org/t/import-pretrained-keras-model-in-julia-for-inference/104924/5 "2023-10-13T22:30:22Z")

</div>

Thanks for the quick replies.

I should have mentioned that the results of the inference would be combined with other computations for which I have Julia code. This is why am not considering a pure Python solution. But I am open to do the inference via PyCall if that is a robust approach.

If I save the model to Pickle I see how to load it but how would I use it ?

I have found FluxML/ONNX.jl . Based on an example, this seems to work despite the statement “ONNX.jl is in the process of a total reconstruction”.

Thanks for your patience.

---

<div class="post-metadata">

### Author: ![dubosipsl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dubosipsl/32/52998_2.png) [@dubosipsl](https://discourse.julialang.org/u/dubosipsl)
#### Post date: [October 21, 2023, 7:34pm UTC](https://discourse.julialang.org/t/import-pretrained-keras-model-in-julia-for-inference/104924/6 "2023-10-21T19:34:32Z")

</div>

To conclude : after trying ONNX and PyCall I am quite happy with PyCall. Here is an example:

```julia
using PyCall: @pyimport
import Images

@pyimport tensorflow as tf
infer(model, x) = model(tf.constant(x)).numpy()

img = Images.imresize(Images.load("data/elephant.jpg"), (224, 224))
rgb = [Float32(getproperty(color, channel)) for z = 1:50, color in img, channel in (:r, :g, :b)]
ResNet50 = tf.keras.applications.resnet50.ResNet50
model = ResNet50(weights = "imagenet")
@time scores = infer(model, rgb);

```

Note that ‘data/elephant.jpg’ is some downloaded picture file.

---

<div class="post-metadata">

### Author: ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)
#### Post date: [October 21, 2023, 9:43pm UTC](https://discourse.julialang.org/t/import-pretrained-keras-model-in-julia-for-inference/104924/7 "2023-10-21T21:43:02Z")

</div>

You might want to rather use PythonCall.jl, and it should be very similar to do.

Regarding ONNX, is it still much used? I’m trying to find out what are the most needed file formats for Julia to support, and I’m not sure we even yet have the have the final best format.

What I think is most important now is the new GGUF file format (and older compatible GGML which seems much used):

> **[What is GGUF and GGML?](https://medium.com/@phillipgimmi/what-is-gguf-and-ggml-e364834d241c)**
>
> GGUF and GGML are file formats used for storing models for inference, particularly in the context of language models like GPT (Generative…

> GGUF and GGML are file formats used for storing models for inference, particularly in the context of language models like GPT (Generative Pre-trained Transformer). Let’s break down the key differences, pros, and cons of each:  
> […]  
> GGUF (GPT-Generated Unified Format)  
> […]
> 
> - Successor to GGML: GGUF aims to address the limitations of GGML and improve the overall user experience. \*
> - No breaking changes: GGUF seeks to eliminate breaking changes, making it easier for users to transition to new versions.
> - Support for various models: GGUF is not limited to llama models, making it more versatile.
> 
> Cons:
> 
> - It may take some time for existing models to be converted to the GGUF format.
> - Users and developers need to adapt to this new format.

> <https://github.com/nomic-ai/gpt4all/issues/1435#issuecomment-1751574160>
>
> \### Feature request
> 
> Is GGUF support coming anytime soon ? Nobody seems to be do…ing GGML anymore as the project was phased out about a month ago.
> 
> \### Motivation
> 
> Yes
> 
> \### Your contribution
> 
> What?

> > WHy cant both be supported ?
> 
> It’s a lot of effort to maintain support for older file formats, since the current version of llama.cpp only supports GGUF. GGUF is intended to be a long-term solution that can easily be extended, so it should be the last new format for a long time.
> 
> There is a script in llama.cpp that you can use to convert LLaMA models to the new format, called ‘convert-llama-ggml-to-gguf.py’.

> **[Python GGUFReader · ggerganov/llama.cpp · Discussion #3718](https://github.com/ggerganov/llama.cpp/discussions/3718)**
>
> The gguf Python package has a GGUFWriter class. Why not also a GGUFReader class? It could be useful for converting models back to torch format or inspecting already converted models.

> **[How GGML format compares to ONNX? · ggerganov/llama.cpp · Discussion #1178](https://github.com/ggerganov/llama.cpp/discussions/1178)**
>
> Hi there, I try to sort machine learning terminology out in my head. Am I right that both ggml and onnx are binary, platform agnostic representation of neural networks? If so, what are use cases fo...

> **[Quantize Llama models with GGML and llama.cpp](https://towardsdatascience.com/quantize-llama-models-with-ggml-and-llama-cpp-3612dfbcc172)**
>
> GGML vs. GPTQ vs. NF4

> <https://github.com/TimDettmers/bitsandbytes/issues/543>
>
> There seems to be no proper explanation for the data type anywhere I've seen (ma…ybe I'm not looking hard enough). What exactly is this data type, and how is FP4 normalized to get NF4?
