# ReactantServer.jl: serve more models per GPU with Reactant.jl and XLA (gauging interest)

**URL:** <https://discourse.julialang.org/t/reactantserver-jl-serve-more-models-per-gpu-with-reactant-jl-and-xla-gauging-interest/137358>\
**Category:** Machine Learning\
**Tags:** inference, server, reactant, xla\
**Created:** [May 31, 2026, 6:00am UTC](https://discourse.julialang.org/t/reactantserver-jl-serve-more-models-per-gpu-with-reactant-jl-and-xla-gauging-interest/137358 "2026-05-31T06:00:58Z")\
**Posts on this page:** 1\
**Showing post:** 14

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [July 23, 2026, 4:11pm UTC](https://discourse.julialang.org/t/reactantserver-jl-serve-more-models-per-gpu-with-reactant-jl-and-xla-gauging-interest/137358/14 "2026-07-23T16:11:27Z")

</div>

It turns out that BERT based models are trivial to export and serve with ReactantServer. I added tutorials here for four common use cases: [Transformer Text Models · ReactantServer.jl](https://enzymead.github.io/ReactantServer.jl/dev/manual/transformers/)

- Dense Embedding: [sentence-transformers/all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2)
- Sparse Embedding: [prithivida/Splade\_PP\_en\_v2](https://huggingface.co/prithivida/Splade_PP_en_v2)
- Cross Encoders: [cross-encoder/ms-marco-MiniLM-L6-v2](https://huggingface.co/cross-encoder/ms-marco-MiniLM-L6-v2)
- Classification: [distilbert/distilbert-base-uncased-finetuned-sst-2-english](https://huggingface.co/distilbert/distilbert-base-uncased-finetuned-sst-2-english)

If your lab/company uses a vector based knowledgebase / retrieval system for RAG, you can already host the models needed for that with ReactantServer.

I also looked into serving LLMs. The main challenges are potentially needing many different compiled program sizes with padding / KV cache using StableHLO semantics / wrapping all of that up together in a way that can be effectively batched. In principle, there shouldn’t be any reason why you couldn’t serve a dense model like [Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B) today. I don’t have time to look into this currently, but serving these sorts of models which can reasonably fit on a single workstation class GPU is something I would like to support in the future. Contributions are welcome!

---

_[View the full topic](https://discourse.julialang.org/t/reactantserver-jl-serve-more-models-per-gpu-with-reactant-jl-and-xla-gauging-interest/137358)._
