# How to reduce redundancy in a list inside a DataFrame

**URL:** <https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585>\
**Category:** New to Julia\
**Tags:** dataframes\
**Created:** [October 11, 2022, 5:03pm UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585 "2022-10-11T17:03:09Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Juan\_Mac\_Donagh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan_mac_donagh/32/31798_2.png) [@Juan\_Mac\_Donagh](https://discourse.julialang.org/u/Juan_Mac_Donagh)\
**Post date:** [October 11, 2022, 5:03pm UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585/1 "2022-10-11T17:03:09Z")

</div>

Hi all. I have two questions, one very basic and one not so basic.

I have a DataFrame that looks like this:

```julia
df = DataFrame([]) 
df.list = [[1,2,3,6,12,13,14],[20,21,23,24,50],[1,3,5,7,10,11,12,13]]

```

And I need to generate a new column, that only stores the “points” of each array. This means, only the first and last number in a sequence of consecutive numbers.  
It would look like this:

```julia
df.new = [[1,3,12,14],[20,21,23,24],[10,13]] #it is important that they are made up of pairs!

```

My solution, so far, looks something like this:

```julia
	df.list2 = [[],[],[]]

	for r in eachrow(df) #loop the df
		j=0 #this is for "walking" down the array
		for i in r.list
			j+=1
			if j == 1 #append the fist number. This needs to go, maybe the fist numbers doesn't have a consecutive pair. 
				append!(r.list2, r[1][j])

			elseif j != 1 
				if (r[1][j]-1) != r[1][j-1] #check if the number is not consecutive to the previous one 
							append!(r.list2, r[1][j])

			elseif j == length(r[1]) #This also should go. It works to append the last value, but it might be an isolated one!
					append!(r.list2,r[1][j])

				
				end
			end
		end
	end
	

	df

```

That kinda works. With the first and last part of the loop, I can append the “points” of each string, and with the middle I can identify some of the numbers that I need. Here is the output:

```julia
df.list2 = [[1,6,12,14],[20,23,50], [1,3,5,7,10,13]]

```

The problem is, **first** , I’m not sure how to create an empty column automatically, so I can append all my values (if I do something like ` df.new = missings(Array, nrow(df))`, that is what I usually do, I get an error, because I am trying to append values, as it were a list).

**Secondly** , I could use some help with the loop. I thought about using ` diff`, so I could save me some trouble, but I am not sure if it will improve the loop.  
Right now I am still getting the “lone” numbers (i e `1,3,5,7` in the last list), and I don’t really know how to get rid of them, and I should find a way to append the last number in the list (in the second list, I can append the ` 20, 23`, but I can’t append the ` 21, 24`, or the `3, 12` on the first one).

I think I can do it with a little more wresting with the for loop, and the different conditions, but any help is welcome.

Thanks a lot! cheers

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [October 11, 2022, 5:23pm UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585/2 "2022-10-11T17:23:10Z")

</div>

Writing a loop is fine. Your logic is sufficiently complicated to merit a loop.

But you should put things in functions to organize your code and make it readable

```julia
julia> function get_consec(x)
           t = [x[1]]
           for i in 2:length(x)
               # Consecutive with last one
               if x[i] == x[i-1] + 1
                   # If its the last one, add it
                   if i == length(x)
                       push!(t, x[i])
                   # If its the end of the chain, push it
                   elseif x[i] != x[i+1] - 1
                       push!(t, x[i])
                   end
               # Not consecutive, start a new chain
               else
                   if i != length(x)
                       if x[i] == x[i+1] -1
                           push!(t, x[i])
                       end
                   end
               end
           end
           return t
       end;

julia> get_consec(df[1,1])
4-element Vector{Int64}:
  1
  3
 12
 14

```

using DataFramesMeta.jl you can do this as

```julia
julia> using DataFramesMeta

julia> @rtransform df :list2 = get_consec(:list)
3×2 DataFrame
 Row │ list list2            
     │ Array… Array…           
─────┼────────────────────────────────────────────────
   1 │ [1, 2, 3, 6, 12, 13, 14] [1, 3, 12, 14]
   2 │ [20, 21, 23, 24, 50] [20, 21, 23, 24]
   3 │ [1, 3, 5, 7, 10, 11, 12, 13] [1, 10, 13]

```

I guess I didn’t get the logic fully right in the function. But the point stands, write a function and use `@rtransform`.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [October 11, 2022, 5:25pm UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585/3 "2022-10-11T17:25:09Z")

</div>

The logic is:

```julia
julia> function identify_pairs(v::AbstractVector)
           local start, stop
           initialized = false
           res = eltype(v)[]
           for x in v
               if initialized
                   if x == stop + 1
                       stop = x
                   else
                       if stop > start
                           push!(res, start, stop)
                           initialized = false
                       end
                       start = stop = x
                       initialized = true
                   end
               else
                   start = stop = x
                   initialized = true
               end
           end
           if stop > start
               push!(res, start, stop)
           end
           return res
       end
identify_pairs (generic function with 1 method)

julia> transform(df, :list => ByRow(identify_pairs) => :new)
3×2 DataFrame
 Row │ list new
     │ Array… Array…
─────┼────────────────────────────────────────────────
   1 │ [1, 2, 3, 6, 12, 13, 14] [1, 3, 12, 14]
   2 │ [20, 21, 23, 24, 50] [20, 21, 23, 24]
   3 │ [1, 3, 5, 7, 10, 11, 12, 13] [10, 13]

```

---

<div class="post-metadata">

**Author:** ![Juan\_Mac\_Donagh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan_mac_donagh/32/31798_2.png) [@Juan\_Mac\_Donagh](https://discourse.julialang.org/u/Juan_Mac_Donagh)\
**Post date:** [October 11, 2022, 5:41pm UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585/4 "2022-10-11T17:41:58Z")

</div>

Thanks a lot! I should try to use functions more often, especially in this cases. I understand the logic, cheers!

---

<div class="post-metadata">

**Author:** ![Juan\_Mac\_Donagh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan_mac_donagh/32/31798_2.png) [@Juan\_Mac\_Donagh](https://discourse.julialang.org/u/Juan_Mac_Donagh)\
**Post date:** [October 11, 2022, 5:42pm UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585/5 "2022-10-11T17:42:54Z")

</div>

Thanks!! As I said in the previous comment, I should tap more often on functions. I am marking this one as the solution. Thanks again!

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [October 11, 2022, 8:09pm UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585/6 "2022-10-11T20:09:06Z")

</div>

Fyi, [this Matlab post](https://www.mathworks.com/matlabcentral/answers/390962-find-the-first-and-last-elements-of-consecutive-values-in-a-vector) provides a compact alternative, probably less efficient than Bogumil’s solution.

A possible Julia translation here into (~10% slower for your input data) :

```julia
function sortedendpoints(v)
    d = diff(eltype(v)[0; diff(v).==1; 0])
    return transpose([v[d.>0] v[d.<0]])[:]
end

transform(df, :list => ByRow(sortedendpoints) => :new) 

```

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [October 11, 2022, 11:48pm UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585/7 "2022-10-11T23:48:32Z")

</div>

This also works:

```julia
transform(df, :list => ByRow(f->filter(x->(x-1 ∈ f) ⊻ (x+1 ∈ f),f)) => :new)

```

(slower in run-time, faster in type-time)

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [October 12, 2022, 7:14am UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585/8 "2022-10-12T07:14:50Z")

</div>

@Dan, what a beautiful one-liner, which also shows the intention most clearly. You might consider adding it to [this list](https://discourse.julialang.org/t/fun-one-liners/28352).

---

<div class="post-metadata">

**Author:** ![Juan\_Mac\_Donagh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan_mac_donagh/32/31798_2.png) [@Juan\_Mac\_Donagh](https://discourse.julialang.org/u/Juan_Mac_Donagh)\
**Post date:** [October 12, 2022, 12:48pm UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585/9 "2022-10-12T12:48:03Z")

</div>

That’s really neat! Thanks! It still amazeme that something that in my (newbie) head was a really complex problem can be solved in many ways, and sometimes like this, in one line. Got lots to learn

---

<div class="post-metadata">

**Author:** ![jules](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@jules](https://discourse.julialang.org/u/jules)\
**Post date:** [October 12, 2022, 1:20pm UTC](https://discourse.julialang.org/t/how-to-reduce-redundancy-in-a-list-inside-a-dataframe/88585/10 "2022-10-12T13:20:29Z")

</div>

Just that this one has quadratic time complexity, so if your arrays get longer you might run into problems quickly.
