# Frustrated using DataFrames

**URL:** <https://discourse.julialang.org/t/frustrated-using-dataframes/67833>\
**Category:** New to Julia\
**Tags:** dataframes, data\_structures\
**Created:** [September 7, 2021, 9:38pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833 "2021-09-07T21:38:10Z")\
**Posts on this page:** 18\
**Page:** 5

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [September 9, 2021, 5:26pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/81 "2021-09-09T17:26:20Z")

</div>

To allow multiple arguments. Plus there are exceptions like `nrow` which can stand alone, might create a lot of method ambiguities when you include these cases.

---

<div class="post-metadata">

**Author:** ![CameronBieganek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cameronbieganek/32/6915_2.png) [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Post date:** [September 9, 2021, 5:39pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/82 "2021-09-09T17:39:49Z")

</div>

@Nathan_Boyer makes a good point. If `source`, `fun`, and `dest` were keyword arguments, I would guess that the internal logic in `select`/`transform` would be approximately the same as it is right now.

---

<div class="post-metadata">

**Author:** ![jzr](https://avatars.discourse-cdn.com/v4/letter/j/eb9ed0/32.png) [@jzr](https://discourse.julialang.org/u/jzr)\
**Post date:** [September 9, 2021, 5:41pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/83 "2021-09-09T17:41:40Z")

</div>

@bkamins @nalimilan Could it be

```julia
transform(df, r"temp" => ByValue(t->((t-32)*5/9)) => (c->c*"celsius"))

```

- `ByValue` is like `ByRow` but the function receives a value instead of a row
- The third component of the pairs is a renamer function.

---

<div class="post-metadata">

**Author:** ![jzr](https://avatars.discourse-cdn.com/v4/letter/j/eb9ed0/32.png) [@jzr](https://discourse.julialang.org/u/jzr)\
**Post date:** [September 9, 2021, 5:46pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/84 "2021-09-09T17:46:54Z")

</div>

With pairs, we can do

```julia
transform(df, 
  :a => :b => :c,
  :d => :e => :f,
  :g => :h => :i,
)

```

which we can’t do with keyword arguments.

---

<div class="post-metadata">

**Author:** ![CameronBieganek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cameronbieganek/32/6915_2.png) [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Post date:** [September 9, 2021, 5:56pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/85 "2021-09-09T17:56:42Z")

</div>

Good point. Unfortunately I’ve been coding in Matlab and Python lately, so I’m a little rusty on DataFrames.jl details.

However, it could probably be handled with the right API. This may not be elegant, but it would work:

```julia
df = DataFrame(a=1:2, b=3:4, c=5:6)

transform(df,
    source = [(:a, ), (:b, :c)],
    fun = [x -> 2x, (x, y) -> x + y],
    dest = [:d, :e]
)

```

This would apply `x -> 2x` to column `:a` and `(x, y) -> x + y` to columns `:b` and `:c`.

One advantage is that it’s a lot more natural to spread keyword arguments over multiple lines than it is to spread a double pair over multiple lines.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [September 9, 2021, 5:58pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/86 "2021-09-09T17:58:35Z")

</div>

I think this would be really cool, but not as an added keyword argument to `transform`, but rather as a function to make pairs.

```julia
make_pair(source = [:a, :b], fun = f, dest = AsTable)

```

etc. Then you can do

```julia
transform(df, make_pair(...))

```

This could probably live in the same package as `Across` and friends.

---

<div class="post-metadata">

**Author:** ![RobertGregg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robertgregg/32/22105_2.png) [@RobertGregg](https://discourse.julialang.org/u/RobertGregg)\
**Post date:** [September 9, 2021, 6:48pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/87 "2021-09-09T18:48:30Z")

</div>

Using a for loop is great if I only have 1 transformation to do, but what if I had 10 other transformations I wanted to do with the data? Doing something like this:

```julia
using DataFrames, Chain

df = DataFrame(Time = [3, 4, 5, 6], TopTemp = [70, 73, 100, missing], BottomTemp = [50, 55, 80, 90])

fahrenheit_to_celsius(t) = Int(round((t - 32) * 5 / 9))

result = @chain df begin
    dropmissing
    filter(row -> row.TopTemp < 90, _)
    transform!(names(df, r"Temp") .=> ByRow(fahrenheit_to_celsius)) #renamecols=false?
    transform!([:TopTemp,:BottomTemp] => (-) => :DiffTemp)
    #etc...
end

```

is a lot easier to read, eliminates the use of temporary variables, and self contains all of data wrangling you performed. I can’t speak on the code performance/comparisons to loops, but I would imagine there would be only be a minor cost, especially if you’re using in-place transformations.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [September 9, 2021, 6:54pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/88 "2021-09-09T18:54:35Z")

</div>

See [my post](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/73) above. You can use a `for` loop inside a `@chain` block easily with the `@aside` macro-flag.

---

<div class="post-metadata">

**Author:** ![RobertGregg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robertgregg/32/22105_2.png) [@RobertGregg](https://discourse.julialang.org/u/RobertGregg)\
**Post date:** [September 9, 2021, 6:58pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/89 "2021-09-09T18:58:47Z")

</div>

That’s actually really cool, seems like you can have the best of both worlds in julia 😁

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [September 9, 2021, 7:14pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/90 "2021-09-09T19:14:17Z")

</div>

> [@Nathan\_Boyer](#):
>
> Similarly, if `fun` contains `ByRow()` or any of the above indexing, then it also cannot be tested alone.

`ByRow` is a fully stand alone thing - unrelated to `DataFrame` object. Simplyfying a bit (to reduce to a single column case) you can think of `ByRow(fun)` as `x -> fun.(x)`.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [September 9, 2021, 7:16pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/91 "2021-09-09T19:16:13Z")

</div>

> [@CameronBieganek](#):
>
> It would be nice if the DataFrames pairs syntax could handle a renaming function in the third position, so that the names of the new columns can be determined programatically

Yes - we are aware of this limitation. It is on a roadmap. If you opened an issue for this probably it will get a higher priority 😃. Thank you!

> [@jzr](#):
>
> `ByValue` is like `ByRow` but the function receives a value instead of a row

It is certainly doable. The only reservation is how much we want to complicate the minilanguage (since it is already complex - and this is part of the reasons this thread was started). Can you please open an issue so we can discuss it there?

> [@CameronBieganek](#):
>
> If `source` , `fun` , and `dest` were keyword arguments, I would guess that the internal logic in `select` / `transform` would be approximately the same as it is right now.

It should be relatively easy to make a companion package investigating this option (wrapper functions would rewrite the specifications passed to the standard `src => fun => dst` style).

---

<div class="post-metadata">

**Author:** ![DoktorMike](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/doktormike/32/2736_2.png) [@DoktorMike](https://discourse.julialang.org/u/DoktorMike)\
**Post date:** [September 10, 2021, 5:46pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/93 "2021-09-10T17:46:47Z")

</div>

Out of curiosity, what is confusing about well established terms like mutate, across etc? Your argument spirals down to that everyone should be coding assembly as far as I can read it. Your : and other special signs are no less magical than the other syntax right? Higher levels of abstractions are not always a bad thing in my mind so I just want to understand where you’re coming from here.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [September 10, 2021, 5:56pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/94 "2021-09-10T17:56:06Z")

</div>

> [@DoktorMike](#):
>
> Out of curiosity, what is confusing about well established terms like mutate, across etc?

the biggest problem with R is that no one ever knows what anything actually means because of “nonstandard evaluation”.

when I see something like `contains("Temp")` I think “that’s a function call on the string “Temp” what does it evaluate to?” But it’s NOT a function call on the string Temp. I actually don’t have the slightest idea what it is. It’s really some macroish magic whose value depends on the context in which it appears.

Now look at across(…) that also looks like "a function called on whatever "contains(“Temp”) returns and whatever `~ (.x - 32)*(5/9)` means. But of course, it’s not that either.

And what does `~(.x - 32)*(5/9)` mean? what is the significance of the symbol .x? Does this evaluate to a thing? And how does that thing work?

So it’s not that I disagree with abstraction, it’s more that I disagree with incredibly obfuscated semantics. The beauty of Julia is that the semantics are usually very clear, and the places where the semantics are different are clearly delineated by `@` macro calls.

I think the `@chain` macro is quite nice, its semantics are clear, and it offers a lot of useful abstraction. I like the functional API with filter, and transform and soforth in Julia, again, because tthe semantics are clear.

---

<div class="post-metadata">

**Author:** ![sijo](https://avatars.discourse-cdn.com/v4/letter/s/da6949/32.png) [@sijo](https://discourse.julialang.org/u/sijo)\
**Post date:** [December 23, 2021, 8:09am UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/95 "2021-12-23T08:09:29Z")

</div>

I just wanted to add that with the awesome 1.3.0 release of DataFrames the temperature example can now be written

```julia
df = DataFrame(Time=[3, 4, 5], TopTemp=[70, 73, 100], BottomTemp=[50, 55, 80])

transform(df, Cols(r"Temp") .=> (t->(t.-32)*5/9), renamecols=false)

# Output
3×3 DataFrame
 Row │ Time TopTemp BottomTemp 
     │ Int64 Float64 Float64    
─────┼────────────────────────────
   1 │ 3 21.1111 10.0
   2 │ 4 22.7778 12.7778
   3 │ 5 37.7778 26.6667

```

and @CameronBieganek’s request for renaming columns has been implemented so we can write

```julia
transform(df, Cols(r"Temp") .=> (t->(t.-32)*5/9) .=> (n->n*"_celsius"))

# Output
3×5 DataFrame
 Row │ Time TopTemp BottomTemp TopTemp_celsius BottomTemp_celsius 
     │ Int64 Int64 Int64 Float64 Float64            
─────┼─────────────────────────────────────────────────────────────────
   1 │ 3 70 50 21.1111 10.0
   2 │ 4 73 55 22.7778 12.7778
   3 │ 5 100 80 37.7778 26.6667

```

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [December 23, 2021, 12:33pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/96 "2021-12-23T12:33:39Z")

</div>

Looks like I need to read the release notes. Cool stuff!

---

<div class="post-metadata">

**Author:** ![lewis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lewis/32/5217_2.png) [@lewis](https://discourse.julialang.org/u/lewis)\
**Post date:** [April 22, 2022, 12:44am UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/97 "2022-04-22T00:44:17Z")

</div>

Everyone was remarkably patient with this except for some criticism of R.

Regardless of language, there are problems with these ad hoc containers layered onto the real datatypes. APIs for doing the same thing (equivalent semantics) have no similarity. Performance is hard to predict.

It’s really time for a heterogeneous matrix: columns can be different types as long as every element is the same type in a column. Memory access from such a matrix must be slower than from a homogeneous matrix, but every memory location can be calculated. Now, just treat this mythical being like a matrix (2D array).

The closest way to get to array manipulation of a collection of dissimilar columns in Julia today is with Typed Tables, which no one mentioned. I use typed tables to hold simulation data with 16 columns and up to 8 million rows. There are some weirdnesses because the structure is immutable, but generally it’s a collection of vectors.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [April 22, 2022, 5:22am UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/98 "2022-04-22T05:22:07Z")

</div>

Indeed, type-stable performant collections are widely useful. Julia is flexible enough to have several reasonably popular implementations of this concept: for example, there is StructArrays in addition to TypedTables you mention. They basically have the same layout, and both implement the Tables interface and can be used in generic tabular functions. For the “inverse” row-based layout, there is the built-in Vector-of-NamedTuples, that is also both a table and a regular array.

---

<div class="post-metadata">

**Author:** ![lewis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lewis/32/5217_2.png) [@lewis](https://discourse.julialang.org/u/lewis)\
**Post date:** [April 22, 2022, 4:09pm UTC](https://discourse.julialang.org/t/frustrated-using-dataframes/67833/99 "2022-04-22T16:09:31Z")

</div>

Thanks.

I’ll take a look at StructArrays. The one problem with the named tuple approach of TypedTables is the horrendous type definition that results.

[Previous page](https://discourse.julialang.org/t/frustrated-using-dataframes/67833.md?page=4)
