# What do you need from a distributed computation framework?

**URL:** <https://discourse.julialang.org/t/what-do-you-need-from-a-distributed-computation-framework/55676>\
**Category:** Julia at Scale\
**Created:** [February 20, 2021, 2:16pm UTC](https://discourse.julialang.org/t/what-do-you-need-from-a-distributed-computation-framework/55676 "2021-02-20T14:16:41Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [February 20, 2021, 2:16pm UTC](https://discourse.julialang.org/t/what-do-you-need-from-a-distributed-computation-framework/55676/1 "2021-02-20T14:16:41Z")

</div>

Recently I was thinking about adding UDF support to [Spark.jl](https://github.com/dfdx/Spark.jl) - a Julia bindings for Apache Spark. UDFs is how you run custom functions on distributed datasets in Python and Scala nowadays. However, adding them to Julia turned to be quite a [huge task](https://github.com/dfdx/Spark.jl/issues/86#issuecomment-782441438), requiring not just one-time investment, but continuous support in all future releases.

This made me reconsider the task we are trying to solve. I personally use distributed computing mostly to prepare datasets from files on AWS S3 + (rarely) to run distributed streaming applications. It also would be great to be able to implement some of the distributed ML algorithms. But these things can be done using many tools, including Spark, Flink, Julia’s ClusterManagers, custom Kubernetes application, etc.

To better understand the need from the community, I wonder what do _you_ need from a distributed computation framework?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [July 20, 2021, 11:30pm UTC](https://discourse.julialang.org/t/what-do-you-need-from-a-distributed-computation-framework/55676/2 "2021-07-20T23:30:26Z")

</div>

Sounds like too hard. I think it’s a numbers’ game. If Julia community got big enough then Spark will make the UDF function available. It doesn’t feel like something that is sustainable to maintain by one or two-person in the Julia community unless they really want it!

Perhaps the right way if for the motivated individual to revive JuliaDB
