# Loopless calculations

**URL:** <https://discourse.julialang.org/t/loopless-calculations/108628>\
**Category:** Data\
**Created:** [January 10, 2024, 3:49pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628 "2024-01-10T15:49:04Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![SergeantMike](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sergeantmike/32/202613_2.png) [@SergeantMike](https://discourse.julialang.org/u/SergeantMike)\
**Post date:** [January 10, 2024, 3:49pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/1 "2024-01-10T15:49:05Z")

</div>

As an Ecologist, I have a hard time wrapping my head around some of the more esoteric mathematical voo-doo that Julia does, and by that I mean pretty much anything more complicated than x+y=z. With that said, here’s what I want to do.

I set up a DataFrame:  
` df=DataFrame(x1 = [0.0028,0.0136,0.0310,0.0342,0.0466], x2 =[0.0009,0.0092,0.0255,0.0525,0.0813], x3 =[0.0089,0.0299,0.0413,0.0773,0.1147])`

Now what I would like to do is to have Julia do this calculation:  
`t=sum(x.*y)/sqrt(sum(x.^2)*sum(y.^2))`

on the columns in the DataFrame in the following way

first calculation is  
x=df.x1  
y=df.x2

Next calculation is  
x=df.x1  
y=df.x3

Next calculation is  
x=df.x2  
y=df.x3

Note: this is a minimal set I have actually 6 columns (x1:x6 in this example) and want to continue the pattern of for all 6 columns.

I can do this using a series of for loops but it keeps bugging me somewhere in the dark moldy recesses of my brain (and there are a lot of those) that there was a way to do this without the loops. Any ideas?

Thank you  
Mike

---

<div class="post-metadata">

**Author:** ![Jeff\_Emanuel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff_emanuel/32/15440_2.png) [@Jeff\_Emanuel](https://discourse.julialang.org/u/Jeff_Emanuel)\
**Post date:** [January 10, 2024, 3:58pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/2 "2024-01-10T15:58:22Z")

</div>

Change your `t` to a function:

```julia
t(x,y)=sum(x.*y)/sqrt(sum(x.^2)*sum(y.^2))

```

and then you could use Combinatorics.jl to iterate over the combinations of your column names, and apply the `t` function. Something like:

```julia
results = Dict( (x,y) => t(df[!,x],df[!,y)) for (x,y) in combinations(names(df),2))

```

This is untested and likely suboptimal but should get you started. Caveat emptor.

And for the record, loops can be clearer than esoteric voodoo. If you write clear loops, you may thank yourself later.

---

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [January 10, 2024, 4:42pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/3 "2024-01-10T16:42:23Z")

</div>

Here’s an option:

```julia
using DataFrames
using Combinatorics

df=DataFrame(
    x1 = rand(5),
    x2 = rand(5),
    x3 = rand(5),
    x4 = rand(5),
    x5 = rand(5),
    x6 = rand(5)
)

t(x,y) = sum(x .* y) / sqrt(sum(x.^2) * sum(y.^2))

combos = combinations([:x1, :x2, :x3, :x4, :x5, :x6], 2)
	
cos_similarities = [t(df[!, c[1]], df[!, c[2]]) for c in combos]

# Result:

15-element Vector{Float64}:
 0.5039552010150775
 0.9043989038369669
 0.8212958683924352
 0.7143122907695525
 0.7599172627439368
 0.5850956757745355
 0.6471346657792967
 0.5647398171102852
 0.7604339270540581
 0.8210921981937371
 0.9221248721177893
 0.8657435871629356
 0.6768400199441867
 0.6873007136174918
 0.9183546953980369

```

Or, if you want a matrix:

```julia
cos_sim_matrix = [t(df[!, i], df[!, j]) for i in 1:ncol(df), j in 1:ncol(df)]

6×6 Matrix{Float64}:
 1.0 0.503955 0.904399 0.821296 0.714312 0.759917
 0.503955 1.0 0.585096 0.647135 0.56474 0.760434
 0.904399 0.585096 1.0 0.821092 0.922125 0.865744
 0.821296 0.647135 0.821092 1.0 0.67684 0.687301
 0.714312 0.56474 0.922125 0.67684 1.0 0.918355
 0.759917 0.760434 0.865744 0.687301 0.918355 1.0

```

---

<div class="post-metadata">

**Author:** ![simsurace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simsurace/32/30216_2.png) [@simsurace](https://discourse.julialang.org/u/simsurace)\
**Post date:** [January 10, 2024, 4:45pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/4 "2024-01-10T16:45:33Z")

</div>

Isn’t what you want just the correlation matrix (without subtracting the mean)?

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [January 10, 2024, 4:58pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/5 "2024-01-10T16:58:13Z")

</div>

> [@Jeff\_Emanuel](#):
>
> And for the record, loops can be clearer than esoteric voodoo. If you write clear loops, you may thank yourself later.

A simple version with loops that provides a nice vector of tuples output:

```julia
[(i, j, t(df[!,i], df[!,j])) for i in 1:ncol(df) for j in (i+1):ncol(df)]

```

---

<div class="post-metadata">

**Author:** ![SergeantMike67](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sergeantmike67/32/25103_2.png) [@SergeantMike67](https://discourse.julialang.org/u/SergeantMike67)\
**Post date:** [January 10, 2024, 5:44pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/6 "2024-01-10T17:44:57Z")

</div>

I don’t really know. I find a lot of my bottlenecks stem from not having the language to communicate effectively enough to even do a search online.

---

<div class="post-metadata">

**Author:** ![SergeantMike67](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sergeantmike67/32/25103_2.png) [@SergeantMike67](https://discourse.julialang.org/u/SergeantMike67)\
**Post date:** [January 10, 2024, 5:54pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/7 "2024-01-10T17:54:22Z")

</div>

> [@rafael.guerra](#):
>
> > [@Jeff\_Emanuel](#):
> >
> > And for the record, loops can be clearer than esoteric voodoo. If you write clear loops, you may thank yourself later.
> 
> A simple version with loops that provides a nice vector of tuples output:
> 
> ```julia
> [(i, j, t(df[!,i], df[!,j])) for i in 1:ncol(df) for j in (i+1):ncol(df)]
> 
> ```

I understand that loops can be clearer but in all honesty there is more than one learning outcome I’m trying to accomplish besides getting the answer. I am trying to learn some of that Julian Voodoo and therefore make future code better and faster. Loops would have got it done but as I like to call it it is the Excel (as in spreadsheet) way of approaching the problem. Do each step of solving the equation in one cell then arrive at an answer after several columns and sheets making it very loopy and at times very confusing.

---

<div class="post-metadata">

**Author:** ![Jeff\_Emanuel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff_emanuel/32/15440_2.png) [@Jeff\_Emanuel](https://discourse.julialang.org/u/Jeff_Emanuel)\
**Post date:** [January 10, 2024, 6:22pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/8 "2024-01-10T18:22:53Z")

</div>

> [@SergeantMike67](#):
>
> there is more than one learning outcome

Of course, which is why I showed some voodoo. On the flip side, some concise construct may be quicker to comprehend than parsing and digesting several lines of loop code. There’s always tradeoffs.

---

<div class="post-metadata">

**Author:** ![yha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yha/32/3502_2.png) [@yha](https://discourse.julialang.org/u/yha)\
**Post date:** [January 10, 2024, 11:19pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/9 "2024-01-10T23:19:17Z")

</div>

For a matrix result, you can also use `StatsBase.pairwise`:

```julia
using StatsBase
pairwise(t, eachcol(df); symmetric = true)

```

passing `symmetric = true` avoid recalculating each pair in the opposite order, since your function `t` is symmetric.  
Or to get tuples of column names and `t` values:

```julia
pairwise(propertynames(df); symmetric = true) do n1, n2
    (n1, n2, t(df[!,n1], df[!,n2]))
end

```

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [January 11, 2024, 1:46am UTC](https://discourse.julialang.org/t/loopless-calculations/108628/10 "2024-01-11T01:46:51Z")

</div>

I know you said you don’t like math magic voodoo, but

```julia
t(x,y)=sum(x.*y)/sqrt(sum(x.^2)*sum(y.^2))

```

can be simplified (and substantially sped up) with

```julia
using LinearAlgebra
t(x,y)=dot(x, y)/sqrt(dot(x, x)*dot(y, y))

```

by using the [`dot`](https://docs.julialang.org/en/v1/stdlib/LinearAlgebra/#LinearAlgebra.dot) function from the `LinearAlgebra` standard library, which computes the [scalar product](https://en.wikipedia.org/wiki/Dot_product) of two vectors, operation which is highly optimised by BLAS libraries.

---

<div class="post-metadata">

**Author:** ![SergeantMike](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sergeantmike/32/202613_2.png) [@SergeantMike](https://discourse.julialang.org/u/SergeantMike)\
**Post date:** [January 11, 2024, 10:50pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/11 "2024-01-11T22:50:33Z")

</div>

To all of you who responded, I thank you very much for the knowledge. I lave learned much and plan to implement in my next fun Julian adventure. The Julia community is very gracious and generous with their knowledge. I knew Julia was the right community for me.

---

<div class="post-metadata">

**Author:** ![SergeantMike](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sergeantmike/32/202613_2.png) [@SergeantMike](https://discourse.julialang.org/u/SergeantMike)\
**Post date:** [January 11, 2024, 10:57pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/12 "2024-01-11T22:57:54Z")

</div>

> [@giordano](#):
>
> I know you said you don’t like math magic voodoo, but
> 
> ```julia
> t(x,y)=sum(x.*y)/sqrt(sum(x.^2)*sum(y.^2))
> 
> ```
> 
> can be simplified (and substantially sped up) with
> 
> ```julia
> using LinearAlgebra
> t(x,y)=dot(x, y)/sqrt(dot(x, x)*dot(y, y))
> 
> ```
> 
> by using the [`dot`](https://docs.julialang.org/en/v1/stdlib/LinearAlgebra/#LinearAlgebra.dot) function from the `LinearAlgebra` standard library, which computes the [scalar product](https://en.wikipedia.org/wiki/Dot_product) of two vectors, operation which is highly optimised by BLAS libraries.

It’s not that I don’t like math magic voodoo, my brain isn’t wired like that. I am an evolutionary ecologist and my math interest consists of

`Male + Female = Offspring.`

Its how the offspring is different from mom and dad and what did the environment have to do with it that my mind ponders. This is a bit different that math.

All that said, I will try your solution and hopefully learn something from it I can use later on.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [January 11, 2024, 11:27pm UTC](https://discourse.julialang.org/t/loopless-calculations/108628/13 "2024-01-11T23:27:28Z")

</div>

> [@giordano](#):
>
> ```julia
> using LinearAlgebra
> t(x,y)=dot(x, y)/sqrt(dot(x, x)*dot(y, y))
> 
> ```

Better to use `dot(x, y) / (norm(x) * norm(y))`, which avoids spurious overflow/underflow if `x` or `y` have norms that are bigger/smaller than about 10^{\pm 150} (though it is slower – trading off speed for safety). (If you don’t care about this, `sum(abs2, x)` is likely to be faster than `dot(x,x)`.)
