# Avoiding global variables while using DataFrames

**URL:** <https://discourse.julialang.org/t/avoiding-global-variables-while-using-dataframes/70314>\
**Category:** General Usage\
**Tags:** question, dataframes\
**Created:** [October 25, 2021, 2:43am UTC](https://discourse.julialang.org/t/avoiding-global-variables-while-using-dataframes/70314 "2021-10-25T02:43:52Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![George9000](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/george9000/32/23619_2.png) [@George9000](https://discourse.julialang.org/u/George9000)\
**Post date:** [October 25, 2021, 2:43am UTC](https://discourse.julialang.org/t/avoiding-global-variables-while-using-dataframes/70314/1 "2021-10-25T02:43:52Z")

</div>

How does one follow the manual’s guidance for performance and avoid global variables while working with DataFrames? For small tables, one could take an approach like this:

```julia
construct1() = DataFrame(A=1:3, B=5:7, fixed=1)

typeof(construct1().A)
names(construct1())
propertynames(construct1())

function foo()
    bar = [1, 2, 3]
    baz = ["a", "b", "c"]
    DataFrame(; bar, baz)
end
foo()

```

When working with large DFs needing to be read from disk each time via `CSV.read` this approach seems impractical — even if a fast reader like Arrow is used. Much of working with a new DF involves exploring, understanding, and cleaning the data. During this process, I see no way to avoid globals. Once that exploration is complete, one could wrap the stable processing code and DF inside a function to be reused with new data.

Do others take a different approach? Or is this a situation to just use global variables?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [October 25, 2021, 1:33pm UTC](https://discourse.julialang.org/t/avoiding-global-variables-while-using-dataframes/70314/2 "2021-10-25T13:33:48Z")

</div>

A DataFrame can be global, as long as you use functions to work with them.

The `transform`, `select`, `combine` functions are written so that the data frame can be global but the operations are fast. DataFramesMeta.jl is the same, and provides more utilities for making working with data frames fast in global scope.

Side note:

```julia
construct1() = DataFrame(A=1:3, B=5:7, fixed=1)

typeof(construct1().A)
names(construct1())
propertynames(construct1())

```

is not doing what you think it’s doing. It’s constructing a new data frame _every time_ you call `construct1`. This will make it impossible to modify a data frame and will be very very slow.

---

<div class="post-metadata">

**Author:** ![George9000](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/george9000/32/23619_2.png) [@George9000](https://discourse.julialang.org/u/George9000)\
**Post date:** [October 25, 2021, 2:59pm UTC](https://discourse.julialang.org/t/avoiding-global-variables-while-using-dataframes/70314/3 "2021-10-25T14:59:14Z")

</div>

Thanks. Yeah, I understood that a new DF is created with each function call. One could still compose it with other functions, `bar(construct1())`, etc. to get transformed results returned from that particular call. However, for large DFs this didn’t seem practical, as you state above. I’d like to understand better how DFs can be global and yet avoid the performance penalty.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [October 25, 2021, 3:15pm UTC](https://discourse.julialang.org/t/avoiding-global-variables-while-using-dataframes/70314/4 "2021-10-25T15:15:49Z")

</div>

[This section](https://docs.julialang.org/en/v1/manual/performance-tips/#kernel-functions) in the performance tips will help.
