# Planing on building an ETL Framework

**URL:** <https://discourse.julialang.org/t/planing-on-building-an-etl-framework/124747>\
**Category:** Data\
**Tags:** data, etl\
**Created:** [January 13, 2025, 11:01pm UTC](https://discourse.julialang.org/t/planing-on-building-an-etl-framework/124747 "2025-01-13T23:01:31Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![th0rben](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/th0rben/32/214682_2.png) [@th0rben](https://discourse.julialang.org/u/th0rben)\
**Post date:** [January 13, 2025, 11:01pm UTC](https://discourse.julialang.org/t/planing-on-building-an-etl-framework/124747/1 "2025-01-13T23:01:31Z")

</div>

I am planing to build a package for ETL processes as my Bachelor thesis.  
My plan is to to have the package be a simple way to build data pipelines following ETL. Having a simple way to load the data from different sources, transforming the data and then loading it back into storage or use it elsewhere.

I am still trying to figure out what i should include into the package. So far i am thinking about the following:

- connectivity to multiple data-sources
- data transformation with custom pipelines
  - might add support for data cleaning
  - might add aggregations, filters, joins, splits

- multi threading/parallel computing
- logging and monitoring
- multi source and target

It would be helpfull to get some insights from you guys.  
Please tell me if i am missing anything important.

---

<div class="post-metadata">

**Author:** ![p-gw](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/p-gw/32/210518_2.png) [@p-gw](https://discourse.julialang.org/u/p-gw)\
**Post date:** [January 14, 2025, 1:01pm UTC](https://discourse.julialang.org/t/planing-on-building-an-etl-framework/124747/2 "2025-01-14T13:01:41Z")

</div>

I tried something similar a while ago when I needed an automation solution for work. I started to build very basic custom pipelines reading data from databases and parquet files (DuckDB.jl), cleaning and aggregating data (DataFrames.jl) and saving the results to parquet files.

I can’t find the threads now, but I remember there were some initial tries to program a framework like this. However, I think they ended up abandoned.

While I found julia quite suitable for the task, in my project I discovered myself basically reinventing [dlt](https://github.com/dlt-hub/dlt) and [dbt](https://www.getdbt.com/), so I just ended up using these existing solutions instead.

---

<div class="post-metadata">

**Author:** ![era127](https://avatars.discourse-cdn.com/v4/letter/e/eb8c5e/32.png) [@era127](https://discourse.julialang.org/u/era127)\
**Post date:** [January 14, 2025, 7:56pm UTC](https://discourse.julialang.org/t/planing-on-building-an-etl-framework/124747/3 "2025-01-14T19:56:27Z")

</div>

The advantage of Julia would be to use the repl to write custom code in Julia and then register it as a UDF to run it inside of Duckdb. This would leverage the c function performance to run models or optimizations within the stream processing.
