# Different ways to calculate rowwise sums?

**URL:** <https://discourse.julialang.org/t/different-ways-to-calculate-rowwise-sums/71026>\
**Category:** New to Julia\
**Tags:** dataframes\
**Created:** [November 5, 2021, 8:36pm UTC](https://discourse.julialang.org/t/different-ways-to-calculate-rowwise-sums/71026 "2021-11-05T20:36:47Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [November 5, 2021, 8:36pm UTC](https://discourse.julialang.org/t/different-ways-to-calculate-rowwise-sums/71026/1 "2021-11-05T20:36:47Z")

</div>

```julia
Random.seed!(1234);
df = DataFrame(randn(10, 4), :auto);

```

```julia
 Row │ x1 x2 x3
 ────┼──────────────────────────────────
   1 │ 0.867347 2.21188 -0.560501
   2 │ -0.901744 0.532813 -0.0192918
   3 │ -0.494479 -0.271735 0.128064
   4 │ -0.902914 0.502334 1.85278
   5 │ 0.864401 -0.516984 -0.827763

```

Any of these three options gives me the rowwise sum:

```julia
sum.(eachrow(df))

combine(df, AsTable(:) .=> sum)

select(df, AsTable(:) => ByRow(sum) => :sum)

```

```julia
 Row │ x1_x2_x3_sum 
─────┼──────────────
   1 │ 2.51872
   2 │ -0.388222
   3 │ -0.63815
   4 │ 1.4522
   5 │ -0.480346

```

What’s the difference, or which one should I use?  
The two latter options work with transform() if I want to add this new column to the original dataframe.  
I don’t know how to include the first one with a transform().

If I want columnwise sums instead:

```julia
sum.(eachcol(df)) 
combine(df,names(df) .=> sum)

```

I don’t know if there is any better alternative.

---

<div class="post-metadata">

**Author:** ![goerch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerch/32/29122_2.png) [@goerch](https://discourse.julialang.org/u/goerch)\
**Post date:** [November 5, 2021, 8:51pm UTC](https://discourse.julialang.org/t/different-ways-to-calculate-rowwise-sums/71026/2 "2021-11-05T20:51:08Z")

</div>

`BenchmarkTools` is your friend:

```julia
using Random, DataFrames, BenchmarkTools

Random.seed!(1234);
df = DataFrame(randn(10000, 40), :auto);

@btime sum.(eachrow(df))
@btime combine(df, AsTable(:) .=> sum)
@btime select(df, AsTable(:) => ByRow(sum) => :sum)

```

shows

```julia
  66.825 ms (2379067 allocations: 42.48 MiB)
  368.600 μs (318 allocations: 2.99 MiB)
  740.400 μs (240 allocations: 93.88 KiB)

```

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [November 5, 2021, 8:54pm UTC](https://discourse.julialang.org/t/different-ways-to-calculate-rowwise-sums/71026/3 "2021-11-05T20:54:34Z")

</div>

OK, that is about the time and memory but what about other considerations?

Why does the third option use much less memory than the second one?

---

<div class="post-metadata">

**Author:** ![goerch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerch/32/29122_2.png) [@goerch](https://discourse.julialang.org/u/goerch)\
**Post date:** [November 5, 2021, 9:00pm UTC](https://discourse.julialang.org/t/different-ways-to-calculate-rowwise-sums/71026/4 "2021-11-05T21:00:49Z")

</div>

> [@Juan](#):
>
> OK, that is about the time and memory but what about other considerations?

> [@Juan](#):
>
> The two latter options work with transform() if I want to add this new column to the original dataframe.

The latter options are the faster ones, also?

> [@Juan](#):
>
> Why does the third option use much less memory than the second one?

Good question, anyone?

> [@Juan](#):
>
> If I want columnwise sums instead:
> 
> ```julia
> sum.(eachcol(df)) 
> combine(df,names(df) .=> sum)
> 
> ```
> 
> I don’t know if there is any better alternative.

Do you want me to investigate these too?

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [November 5, 2021, 11:19pm UTC](https://discourse.julialang.org/t/different-ways-to-calculate-rowwise-sums/71026/5 "2021-11-05T23:19:12Z")

</div>

Yes, but I’m more interested in knowing the best way to append the column of rowwise sums.

```julia
transform(df, AsTable(:) .=> sum)
transform(df, AsTable(:) .=> ByRow(sum))  
hcat(df,combine(df, AsTable(:) .=> sum))

```

Anything better?  
Here ByRow is much slower, though it needs less memory.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [November 8, 2021, 4:54pm UTC](https://discourse.julialang.org/t/different-ways-to-calculate-rowwise-sums/71026/6 "2021-11-08T16:54:48Z")

</div>

These are very small tables so the performance is affected by factors not related to summation.

In DataFrames.jl 1.3 that will be released soon (it is held back the release of Julia 1.7) the fastest option, especially for wide and large tables will be `transform(df, AsTable(:) => ByRow(sum))`.

For the time being an easy (i.e. IMO natural for someone knowing how things in Julia Base work), and reasonably fast option is `df.sum = sum(eachcol(df))`.

* * *

Also note that `.=>` is in this case the same as `=>` the `.` does not do anything in this situation.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [November 8, 2021, 8:17pm UTC](https://discourse.julialang.org/t/different-ways-to-calculate-rowwise-sums/71026/7 "2021-11-08T20:17:19Z")

</div>

In fact my initial question wasn’t about speed but to know if there are other differences or disadvantages. For example if the returned object is more or less useful (dataframes vs other things) for additional operations.

And I have just discovered that…

```julia
df = DataFrame(randn(5, 3), :auto);
allowmissing!(df)
df[1,1] = missing

select(df, AsTable(:) .=> sum∘skipmissing => :sum)

select(df, AsTable(:) => ByRow(sum∘skipmissing) => :sum)  
sum.(skipmissing.(eachrow(df)))

```

The first option, not using ByRow, doesn’t produce the expected output if we have missings. I guess we will have similar problems with other functions.

And another question,  
How can I run more complex functions inside the ByRow()?

I’ve tried  
`select(df, AsTable(:) => ByRow(x -> x.^2) => :sum)`

but it doesn’t work, it says:

> ERROR: ArgumentError: broadcasting over dictionaries and `NamedTuple`s is reserved

@bkamins How can I calculate the sum of the squares of the elements for each row?

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [November 8, 2021, 8:52pm UTC](https://discourse.julialang.org/t/different-ways-to-calculate-rowwise-sums/71026/8 "2021-11-08T20:52:55Z")

</div>

> [@Juan](#):
>
> The first option, not using ByRow, doesn’t produce the expected output if we have missings

In this option you are passing whole columns to `skipmissing` and whole columns are never missing.

`select(df, AsTable(:) => ByRow(sum∘skipmissing) => :sum) ` is correct and will be very fast in DataFrames.jl 1.3.

> How can I run more complex functions inside the ByRow()?

```julia
select(df, AsTable(:) => ByRow(x -> x.^2) => :sum)

```

this fails but not because of DataFrames.jl but because of Julia Base and in general it is incorrect as there is no sum in your expression. You have to write:

```julia
select(df, AsTable(:) => ByRow(x -> sum(v -> v^2, x)) => :sum)

```

(you need to `sum` the squares)

In general I would recommend to handle functions like `x -> sum(v -> v^2, x)` not as anonymous but rather predefine them - as using a lot of anonymous functions can lead to not very readable code (it is like a decision whether one should write one long one line expression or rather define variables to store intermediate values even if they are discarded later).
