# Looping over previous row efficiency

**URL:** <https://discourse.julialang.org/t/looping-over-previous-row-efficiency/75749>\
**Category:** New to Julia\
**Tags:** loops, dataframes\
**Created:** [February 3, 2022, 7:22pm UTC](https://discourse.julialang.org/t/looping-over-previous-row-efficiency/75749 "2022-02-03T19:22:42Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![korilium](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/korilium/32/48714_2.png) [@korilium](https://discourse.julialang.org/u/korilium)\
**Post date:** [February 3, 2022, 7:22pm UTC](https://discourse.julialang.org/t/looping-over-previous-row-efficiency/75749/1 "2022-02-03T19:22:42Z")

</div>

So I have a dataset with all stock prices and i want to compute the returns.

I have done this the following way:

```julia
test = copy(dollar_portfolio)
returns = copy(dollar_portfolio[2:end,:])
@time for column in 2:length(dollar_portfolio[1,:])
    for row in 2:length(dollar_portfolio[:,1])
        test[row, column] = log(dollar_portfolio[row, column]) - log(dollar_portfolio[row-1,column]) 
    end 
    returns[:, column] = test[2:end, column]
end 

```

I feel like this is not really good coding as I fill in the dataframe in a loop.  
Does someone has a better idea current benchmark is the following:

```julia
  0.002446 seconds (37.77 k allocations: 692.391 KiB

```

thanks in advance

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [February 3, 2022, 7:56pm UTC](https://discourse.julialang.org/t/looping-over-previous-row-efficiency/75749/2 "2022-02-03T19:56:14Z")

</div>

yes, this will be slow. The issue is that accessing a data frame in this way is type-unstable. Julia doesn’t know the types of data frame columns inside hot loops and so can’t generate fast code.

The workaround is to use a function barrier. In general, for fast code with data frame, write a function which acts on _vectors_ and then call that function on the columns you want.

Wait, also are you getting columns and rows confused? It looks like you are generating many columns, each a lag of the previous column. Usually this is done by rows…

EDIT: Sorry I did not read your code carefully enough. Try something like this

```julia
julia> df = DataFrame(rand(1000, 100), :auto);

julia> function get_lag(x)
           out = similar(x)
           out[1] = 0
           for row in 2:length(x)
               out[row] = log(x[row]) - log(x[row-1])
           end
           return out
       end;

julia> test = copy(df)
       for column in 1:ncol(df)
           test[!, column] = get_lag(df[!, column])
       end

```
