# Pyjulia - accessing data in a Julia Dataframe that was called from Python

**URL:** <https://discourse.julialang.org/t/pyjulia-accessing-data-in-a-julia-dataframe-that-was-called-from-python/32082>\
**Category:** General Usage\
**Created:** [December 10, 2019, 5:24am UTC](https://discourse.julialang.org/t/pyjulia-accessing-data-in-a-julia-dataframe-that-was-called-from-python/32082 "2019-12-10T05:24:53Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![00krishna](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/00krishna/32/8843_2.png) [@00krishna](https://discourse.julialang.org/u/00krishna)\
**Post date:** [December 10, 2019, 5:24am UTC](https://discourse.julialang.org/t/pyjulia-accessing-data-in-a-julia-dataframe-that-was-called-from-python/32082/1 "2019-12-10T05:24:53Z")

</div>

I had a bunch of help from @stevengj to solve this problem of calling Julia from python. So that was great. I can now write a function in Julia that generates a Dataframe and call that back into python. But I was not sure how to actually access the elements of the Dataframe, since the returned object is a `PyCall.jlwrap`. Let me give an example.

If I have a julia file like so:

```julia
#jinclude.jl
using DataFrames

function myDf(row, columns)
  DataFrame(rand(row, columns))
end

```

Then I have a Python script like this.

```julia
#Pythonscript.py

import julia
jl = julia.Julia(compiled_modules=False)
jl.include('jinclude.jl') # my file with functions
from julia import Main
a = Main.myDf(3,4) # function from `jinclude.jl` file.

```

So `a` will return a DataFrame, but if I go to access its elements, such as `a[1,1]` I get an error `TypeError: 'PyCall.jlwrap' object is not subscriptable`. This is understandable as I did not imagine the data structures would pass smoothly between languages. But does anyone know a good way to destructure or export the PyCall.jlwrap so that I can pull the elements of the dataframe array into Python?

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [December 10, 2019, 5:49am UTC](https://discourse.julialang.org/t/pyjulia-accessing-data-in-a-julia-dataframe-that-was-called-from-python/32082/2 "2019-12-10T05:49:12Z")

</div>

I usually let Pandas.jl handle the conversion and receive the dataframe as `pandas.DataFrame` on Python side.

```julia
#jinclude.jl
using DataFrames
import Pandas

function myDf(row, columns)
  Pandas.DataFrame(DataFrame(rand(row, columns)))
end

```

---

<div class="post-metadata">

**Author:** ![00krishna](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/00krishna/32/8843_2.png) [@00krishna](https://discourse.julialang.org/u/00krishna)\
**Post date:** [December 10, 2019, 6:29am UTC](https://discourse.julialang.org/t/pyjulia-accessing-data-in-a-julia-dataframe-that-was-called-from-python/32082/3 "2019-12-10T06:29:26Z")

</div>

Oh yeah, I just tried this and it worked great. Wow, this is so much easier than trying to do stuff with passing pointers to C arrays and Numpy, etc. Thanks so much for your tip here.

---

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [July 14, 2020, 1:47am UTC](https://discourse.julialang.org/t/pyjulia-accessing-data-in-a-julia-dataframe-that-was-called-from-python/32082/4 "2020-07-14T01:47:16Z")

</div>

I’m dealing with the same issue but I have missing data… How do I get around this?

```julia
JuliaError: Exception 'ArgumentError: Can't create a Pandas.DataFrame from a source that has missing data.' occurred while calling julia code:
df = Pandas.DataFrame(DataFrames.DataFrame(x=1:2, y=[3,missing], dt=[now(),missing]))

```

---

<div class="post-metadata">

**Author:** ![lungben](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lungben/32/12314_2.png) [@lungben](https://discourse.julialang.org/u/lungben)\
**Post date:** [July 14, 2020, 4:44am UTC](https://discourse.julialang.org/t/pyjulia-accessing-data-in-a-julia-dataframe-that-was-called-from-python/32082/5 "2020-07-14T04:44:43Z")

</div>

You could try to coalesce missing data by a dummy value or drop missing rows altogether.

---

<div class="post-metadata">

**Author:** ![00krishna](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/00krishna/32/8843_2.png) [@00krishna](https://discourse.julialang.org/u/00krishna)\
**Post date:** [July 16, 2020, 2:20pm UTC](https://discourse.julialang.org/t/pyjulia-accessing-data-in-a-julia-dataframe-that-was-called-from-python/32082/6 "2020-07-16T14:20:26Z")

</div>

@tk3369 what I did was convert all missing values to NAs or NaNs in Julia, before pushing to python. The error message is really confusing because it does not explain the source of the error. The error is caused because the julia `Pandas` package or something does not have the DataFrame function implemented for the `missing` datatype, so it defaults to an implementation up the dispatch chain and fails.

I opened an issue with the Pandas.jl github on this issue.

[https://github.com/JuliaPy/Pandas.jl/issues/71](https://github.com/JuliaPy/Pandas.jl/issues/71)

In my case I just had a column of missing values, and I converted that column to NAs or NaNs, and then ran `Pandas.Dataframe()` and it worked.
