# Connecting DataFrames with diffrent columns

**URL:** https://discourse.julialang.org/t/connecting-dataframes-with-diffrent-columns/99844
**Category:** New to Julia
**Tags:** question
**Created:** [June 4, 2023, 11:39am UTC](https://discourse.julialang.org/t/connecting-dataframes-with-diffrent-columns/99844 "2023-06-04T11:39:30Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![ardex2000](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@ardex2000](https://discourse.julialang.org/u/ardex2000)
#### Post date: [June 4, 2023, 11:39am UTC](https://discourse.julialang.org/t/connecting-dataframes-with-diffrent-columns/99844/1 "2023-06-04T11:39:31Z")

</div>

I am a moron, who does not know how to connect DataFrames with diffrent columns.

So essentialy the issue is:  
offspring1 = hcat(parent1[1:crossover\_point], parent2[(crossover\_point + 1):end])

The result should be a DataFrame containing selected columns from dataframe parent1 and dataframe parent2. Instead I get a matrix. The solution must be obvious, pointing clearly at my lack of intelligence.

---

<div class="post-metadata">

### Author: ![ardex2000](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@ardex2000](https://discourse.julialang.org/u/ardex2000)
#### Post date: [June 4, 2023, 11:51am UTC](https://discourse.julialang.org/t/connecting-dataframes-with-diffrent-columns/99844/2 "2023-06-04T11:51:08Z")

</div>

Well, it turns out I really am an idiot. It seems that parent1[1:crossover\_point] and parent2[(crossover\_point + 1):end] create an instance of DataFrameRow instead of a DataFrame.

Therefore, the solution is:  
offspring1 = hcat(DataFrame(parent1[1:crossover\_point]), DataFrame(parent2[(crossover\_point + 1):end])).

Let my stupidity be a warning to future generations as I am removed by the real Genetic Algorithm that I am trying to recreate.

---

<div class="post-metadata">

### Author: ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)
#### Post date: [June 4, 2023, 12:15pm UTC](https://discourse.julialang.org/t/connecting-dataframes-with-diffrent-columns/99844/3 "2023-06-04T12:15:01Z")

</div>

Your sense of humour will carry you a long way in the real Genetic Algorithm, but I still don’t understand what type is `parent1`, and how this subsetting `parent1[1:crossover_point]` works, as it errored on Julia 1.9 when I tried it.

---

<div class="post-metadata">

### Author: ![ardex2000](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@ardex2000](https://discourse.julialang.org/u/ardex2000)
#### Post date: [June 4, 2023, 12:32pm UTC](https://discourse.julialang.org/t/connecting-dataframes-with-diffrent-columns/99844/4 "2023-06-04T12:32:59Z")

</div>

The DataFrames parent1 and parent2 essentially contain one row and 15 columns each, with 15 different indexes.

For example, parent1 looks like this:  
8762 11968 2472 10832 4342 14746 5286 18075 5779 1915 9584 9051 16086 1530 15590

And parent2 looks like this:  
4423 16075 13175 16537 18392 14211 12010 47 12065 8230 13449 16864 8410 6076 15923

With column names being: Position\_1 to Position\_15.  
The code above is intended to create one offspring by mixing two candidate solutions in a crossover. The second offspring is created in the opposite way. At this point, before mutation, the future generation is created.

---

<div class="post-metadata">

### Author: ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)
#### Post date: [June 4, 2023, 12:53pm UTC](https://discourse.julialang.org/t/connecting-dataframes-with-diffrent-columns/99844/5 "2023-06-04T12:53:20Z")

</div>

Still Dan makes a good point which is unrelated to the Julia version but relevant to the DataFrames version - indexing as `df[1:something]`, ie one dimension only, was deprecated from DataFrames a long time ago so if that works for you you might be on a very outdated (pre 1.0) Version of DataFrames.

---

<div class="post-metadata">

### Author: ![ardex2000](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@ardex2000](https://discourse.julialang.org/u/ardex2000)
#### Post date: [June 4, 2023, 1:34pm UTC](https://discourse.julialang.org/t/connecting-dataframes-with-diffrent-columns/99844/6 "2023-06-04T13:34:33Z")

</div>

I have a version of DataFrames 1.5.0 and Julia 1.9.0, so it is unlikley to be that outdated. Perahps it works because somehow parent1 and parent 2 are created as an instances of DataFrameRow instead of DataFrame.  
Parents are created like that:  
parent\_pop[c, :], parent\_pop[c + 1, :],  
where for example c=1, and parent\_pop is just a regular DataFrame.

---

<div class="post-metadata">

### Author: ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)
#### Post date: [June 4, 2023, 3:48pm UTC](https://discourse.julialang.org/t/connecting-dataframes-with-diffrent-columns/99844/7 "2023-06-04T15:48:28Z")

</div>

You may actually want to avoid DataFrames and use matrices, and when using matrices, let each agent/creature be a column and not a row. Why? DataFrames and Matrices store each column in a block of memory. And when cutting and stitching you would like to use bunched memory and not many far-flung locations in memory. In addition, DataFrames are good when there are disparate types of data (different column types).  
Having said that, an even better way might be to have a DataFrame with all the agent metadata and a column of ‘genes’ which would hold vectors of all the traits.  
Doing this refactoring would also be a good way to gain experience points with Julia.

---

<div class="post-metadata">

### Author: ![ardex2000](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@ardex2000](https://discourse.julialang.org/u/ardex2000)
#### Post date: [June 4, 2023, 4:02pm UTC](https://discourse.julialang.org/t/connecting-dataframes-with-diffrent-columns/99844/8 "2023-06-04T16:02:48Z")

</div>

That is a very useful idea. Perhaps it will faster the process. I base my current knowledge about DataFrames with this course: [1. Environment Setup | JuliaAcademy](https://juliaacademy.com/courses/1237045/lectures/27575941)
