# Byrow function to get the mean of row

**URL:** <https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315>\
**Category:** Data\
**Tags:** dataframes, inmemorydatasets\
**Created:** [July 31, 2023, 3:50pm UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315 "2023-07-31T15:50:40Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 31, 2023, 3:50pm UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/1 "2023-07-31T15:50:40Z")

</div>

I want to get the row wise mean of two columns of data using **[byrow](https://docs.juliahub.com/InMemoryDatasets/cS87e/0.5.5/man/byrow/)** function. These two columns are `T_start(s)` and `T_stop(s)`.  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/c/e/ce3c6aba3e6a3c218334e2f2a121f6fedad56d95.png)  
How should i change code given below ? 👇

```julia
using HTTP,CSV,DataFrames,InMemoryDatasets
begin
	for x in 33683
            url="https://gcn.nasa.gov/circulars/$x"
            txt=String((HTTP.get(url)))
        if occursin("report on behalf of the Swift/UVOT team",txt)
             hb,he=findfirst(r"^Filter"im,txt)
			 lr,_=findnext("\n\nThe",txt,he)
			 cltxt=replace(txt[hb:lr], " +/- "=>s"\t", r" +(\w)"=>s"\t\1" ,r" +(&gt;)"=>s"\t",r"\+/?- ?"=>s"\t")
			df=CSV.read(IOBuffer(cltxt), DataFrame, delim='\t')
			if "Column6" in names(df); rename!(df, :Column6 => :Mag_err); end
			byrow(df ,sum ,2:3)
			@show df
        end
    end
end

```

I am getting error shown below :

 ![image](https://global.discourse-cdn.com/julialang/original/3X/6/9/693f856517a8054da4920efdcda2611a339d4a73.png)  
What are the other ways to get the row-wise mean of two columns of data ?

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [July 31, 2023, 4:08pm UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/2 "2023-07-31T16:08:20Z")

</div>

```julia
df.T_mean = (df."T_start(s)" .+ df."T_stop(s)") ./ 2

```

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 31, 2023, 4:10pm UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/3 "2023-07-31T16:10:19Z")

</div>

I think this would be too slow in comparison to `byrow` function. I have run this code many times.

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [July 31, 2023, 4:36pm UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/4 "2023-07-31T16:36:49Z")

</div>

It strikes me as exceedingly unlikely that an element wise addition of two vectors will be the bottleneck in your problem - have you benchmarked/profiled this?

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 31, 2023, 4:39pm UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/5 "2023-07-31T16:39:00Z")

</div>

No, I couldn’t use `byrow` as it is giving error. By using your code it takes 1.2 second.

---

<div class="post-metadata">

**Author:** ![bertschi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bertschi/32/33462_2.png) [@bertschi](https://discourse.julialang.org/u/bertschi)\
**Post date:** [July 31, 2023, 4:47pm UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/6 "2023-07-31T16:47:39Z")

</div>

Have not worked with `InMemoryDatasets`, but skimming the docs it seems to use its own data type `Dataset` instead of `DataFrames.DataFrame`. Thus, you will probably need to convert your data frame in order to use `byrow`, i.e., the error is that no method is defined for a `DataFrame`.

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [July 31, 2023, 6:36pm UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/7 "2023-07-31T18:36:32Z")

</div>

> [@raman\_kumar](#):
>
> What are the other ways to get the row-wise mean of two columns of data ?

The equivalent in DataFrames would be something like

```julia
transform!(df, AsTable(2:3) => ByRow(mean))
# or
map(mean, eachrow(select(df, 2:3)))

```

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [August 1, 2023, 2:37am UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/8 "2023-08-01T02:37:20Z")

</div>

> [@kevbonham](#):
>
> ```julia
> map(mean, eachrow(select(df, 2:3)))
> 
> ```

This is not working as both packages contain `select`.

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [August 1, 2023, 3:20am UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/9 "2023-08-01T03:20:59Z")

</div>

Oh, It requires to convert `DataFrame` to `Dataset`.

```julia
df=CSV.read(IOBuffer(cltxt), DataFrame, delim='\t')
g=Dataset(df)

```

but it does not show mean column 🫤

```julia
@time m=byrow(g ,mean ,2:3)
df.m
@show df

```

and `byrow` is slower than `df.T_mean = (df."T_start(s)" .+ df."T_stop(s)") ./ 2` code.  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/7/d/7d4eb5251547ec488fc5582d222da794c3202c3a.png)

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [August 1, 2023, 12:16pm UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/10 "2023-08-01T12:16:33Z")

</div>

> [@raman\_kumar](#):
>
> This is not working as both packages contain `select`.

When this happens, you can do `DataFrames.select`

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [August 1, 2023, 12:20pm UTC](https://discourse.julialang.org/t/byrow-function-to-get-the-mean-of-row/102315/11 "2023-08-01T12:20:20Z")

</div>

> [@raman\_kumar](#):
>
> but it does not show mean column

You have to set it, they’re different objects, and it doesn’t look like Datasets `byrow` is mutating in any case. Try Eg, `df.m = byrow(...`
