# Unsure how to solve error message when applying unstack to DataFrame

**URL:** <https://discourse.julialang.org/t/unsure-how-to-solve-error-message-when-applying-unstack-to-dataframe/114214>\
**Category:** General Usage\
**Tags:** question, package, dataframes, unstack\
**Created:** [May 14, 2024, 3:19am UTC](https://discourse.julialang.org/t/unsure-how-to-solve-error-message-when-applying-unstack-to-dataframe/114214 "2024-05-14T03:19:37Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lee\_Phillips](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lee_phillips/32/203984_2.png) [@Lee\_Phillips](https://discourse.julialang.org/u/Lee_Phillips)\
**Post date:** [May 14, 2024, 3:19am UTC](https://discourse.julialang.org/t/unsure-how-to-solve-error-message-when-applying-unstack-to-dataframe/114214/1 "2024-05-14T03:19:37Z")

</div>

Hello All

I have a dataframe, the first 5 rows are shown below

```julia
5×4 DataFrame
 Row │ PointAID_FK PointBID_FK Distance Time
     │ Any Any Float64 Float64
─────┼─────────────────────────────────────────────
   1 │ 32210470 26402465271 206.12 346.51
   2 │ 26402159021 30461660 190.67 282.95
   3 │ 32761450 25501286700 77.28 117.18
   4 │ 25501310391 25501283840 34.41 20.7
   5 │ 26401205831 26402111690 23.67 17.77
 .
 .
 .

```

And Id like to unstack it into a massive array

When I run the unstack command as follows

```julia
unstack(dataframe, :PointAID_FK, :PointBID_FK, :Distance)

```

I get the following error

```julia
ERROR: LoadError: ArgumentError: Non-unique column names produced. Non equal values in `colkey` were mapped to the same column name.

```

After trying to find the potential source of this error message I have no results.

I am completely at a loss for what the potential problem might be

I’m open to any amount of help you can give

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [May 14, 2024, 6:17am UTC](https://discourse.julialang.org/t/unsure-how-to-solve-error-message-when-applying-unstack-to-dataframe/114214/2 "2024-05-14T06:17:34Z")

</div>

The error message tells you that two different values in the source data have the same string representation, so they produce duplicates in column names.

Here is a minimal example:

```julia
julia> df = DataFrame(x=[0, "0"])
2×1 DataFrame
 Row │ x
     │ Any
─────┼─────
   1 │ 0
   2 │ 0

julia> unstack(df, :x, :x, :x)
ERROR: ArgumentError: Non-unique column names produced. Non equal values in `colkey` were mapped to the same column name.

```

And the source of the problem is that `0` (number) and string `"0"` have the same string representation (`"0"`).

---

<div class="post-metadata">

**Author:** ![Lee\_Phillips](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lee_phillips/32/203984_2.png) [@Lee\_Phillips](https://discourse.julialang.org/u/Lee_Phillips)\
**Post date:** [May 14, 2024, 7:08am UTC](https://discourse.julialang.org/t/unsure-how-to-solve-error-message-when-applying-unstack-to-dataframe/114214/3 "2024-05-14T07:08:44Z")

</div>

Thanks for the reply

I’ll be honest, I’m still a bit confused by your reply

Does this mean that somewhere in my data set there is a PointAID\_FK and a PointBID\_FK combination that have the same string representation

For example

```julia
×4 DataFrame
 Row │ PointAID_FK PointBID_FK Distance Time
     │ Any Any Float64 Float64
─────┼─────────────────────────────────────────────
   1 │ 32210470 26402465271 206.12 346.51
   2 │ '32210470' '26402465271' 206.12 346.51
   3 │ 26402159021 30461660 190.67 282.95
   4 │ 32761450 25501286700 77.28 117.18
   5 │ 25501310391 25501283840 34.41 20.7
   6 │ 26401205831 26402111690 23.67 17.77
 .
 .
 .

```

Would this table produce the same error. if I were to try use PointAID\_FK and PointBID\_FK as column names

Like so

```julia
unstack(dataframe, :PointAID_FK, :PointBID_FK, :Distance)

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [May 14, 2024, 7:30am UTC](https://discourse.julialang.org/t/unsure-how-to-solve-error-message-when-applying-unstack-to-dataframe/114214/4 "2024-05-14T07:30:50Z")

</div>

Yes. I.e. the error says only about the dimension used for specification of target column names, but generally yes, you have in your data something like:

```julia
26
"26"

```

And looking at the element type of your column, which is `Any`, it might be the case (as usually `Any` signals that you have mixed values of different types in your column).

Do the following to see if you have this problem:

```julia
df2 = select(df, :some_problematic_column) # just keep the column of interest
unique(df2) # de-duplicate it based on the original values
df2.test = string.(df2.some_problematic_column) # now add string representation
filter(:nrow => >(1), combine(groupby(df2, :test), nrow))) # find values that have duplicate string representation

```

and you will see the list of values for which you have this situation.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [May 14, 2024, 8:53am UTC](https://discourse.julialang.org/t/unsure-how-to-solve-error-message-when-applying-unstack-to-dataframe/114214/5 "2024-05-14T08:53:39Z")

</div>

> [@Lee\_Phillips](#):
>
> And Id like to unstack it into a massive array

Apart from the clarification of the problem that arose, it’s not clear to me what you would do with an unstacked Dataframe like this.  
What I have seen done most typically in these cases is the opposite operation: stack.

---

<div class="post-metadata">

**Author:** ![Lee\_Phillips](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lee_phillips/32/203984_2.png) [@Lee\_Phillips](https://discourse.julialang.org/u/Lee_Phillips)\
**Post date:** [May 14, 2024, 9:27am UTC](https://discourse.julialang.org/t/unsure-how-to-solve-error-message-when-applying-unstack-to-dataframe/114214/6 "2024-05-14T09:27:41Z")

</div>

Its a big list of distance and times between points on a 2d plane

The plan is to use this as a pairwise distance matrix to run a traveling salesman optimization problem

I’m about 75% certain the extra processing time spent unstacking will result in a greater performance later when lots of distance look ups are needed between random pairs of points

---

<div class="post-metadata">

**Author:** ![Lee\_Phillips](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lee_phillips/32/203984_2.png) [@Lee\_Phillips](https://discourse.julialang.org/u/Lee_Phillips)\
**Post date:** [May 14, 2024, 10:33pm UTC](https://discourse.julialang.org/t/unsure-how-to-solve-error-message-when-applying-unstack-to-dataframe/114214/7 "2024-05-14T22:33:37Z")

</div>

Final Update from me

@bkamins was correct, it was a duplicated combination of ID’s

This was because my dataset was built from a concatenation of several other datasets, and there happened to be duplicated data between them

I solved this using two steps

Make sure all my ID’s are encoded as string data type, this can be done using the below command. The below command is running the string parsing operation on each value selected from the PointAID\_FK column, then overwriting the same column with the result.

```julia
df1[!,:PointAID_FK] = string.(df1[!, :PointAID_FK])

```

The next step was to set the set the combine input argument to be ‘last’. This would mean that my unstack result would prefer whatever value would be the latest value seen from the input table

```julia
unstack(dataframe, :PointAID_FK, :PointBID_FK, :Distance, combine = last)

```
