# How would I write this Pandas column filter code in Julia?

**URL:** <https://discourse.julialang.org/t/how-would-i-write-this-pandas-column-filter-code-in-julia/34310>\
**Category:** General Usage\
**Created:** [February 7, 2020, 4:54pm UTC](https://discourse.julialang.org/t/how-would-i-write-this-pandas-column-filter-code-in-julia/34310 "2020-02-07T16:54:58Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Julia1](https://avatars.discourse-cdn.com/v4/letter/j/db5fbb/32.png) [@Julia1](https://discourse.julialang.org/u/Julia1)\
**Post date:** [February 7, 2020, 4:54pm UTC](https://discourse.julialang.org/t/how-would-i-write-this-pandas-column-filter-code-in-julia/34310/1 "2020-02-07T16:54:58Z")

</div>

Hi how’s t going?

I’ve searched all over the place for regular expression column filtering in Julia, but I can’t seem to figure it out.

Say you have a variable regex that contains a regular expression string.

This works for me in Pandas:  
cols = [val for val in df.columns if df[val].str.contains(regex).any()]

How would I do this in Julia

---

<div class="post-metadata">

**Author:** ![aaowens](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aaowens/32/12101_2.png) [@aaowens](https://discourse.julialang.org/u/aaowens)\
**Post date:** [February 7, 2020, 5:05pm UTC](https://discourse.julialang.org/t/how-would-i-write-this-pandas-column-filter-code-in-julia/34310/2 "2020-02-07T17:05:56Z")

</div>

Are you looking for columns whose name matches a regular expression?

That’s just `df[!, r"x"]`. Go here [Getting Started · DataFrames.jl](http://juliadata.github.io/DataFrames.jl/stable/man/getting_started/#Indexing-syntax-1) and look for the regular expression example.

---

<div class="post-metadata">

**Author:** ![Julia1](https://avatars.discourse-cdn.com/v4/letter/j/db5fbb/32.png) [@Julia1](https://discourse.julialang.org/u/Julia1)\
**Post date:** [February 7, 2020, 5:36pm UTC](https://discourse.julialang.org/t/how-would-i-write-this-pandas-column-filter-code-in-julia/34310/3 "2020-02-07T17:36:34Z")

</div>

No I’m looking columns that have a row value that matches the regular expression. Not the column name itself.

---

<div class="post-metadata">

**Author:** ![dmolina](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmolina/32/5246_2.png) [@dmolina](https://discourse.julialang.org/u/dmolina)\
**Post date:** [February 7, 2020, 6:45pm UTC](https://discourse.julialang.org/t/how-would-i-write-this-pandas-column-filter-code-in-julia/34310/4 "2020-02-07T18:45:34Z")

</div>

It is simple:

```julia
cols = [k for k in names(df) if any(occursin.(r"...", df[:,k]))]

```

names(df) gives the columns.  
occursin(r"…", string) indicates if the regexp is inside the string. If you have a vector you must use it with the “.”.

I recommend the official documentation, and the tutorial of the same author.

---

<div class="post-metadata">

**Author:** ![aaowens](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aaowens/32/12101_2.png) [@aaowens](https://discourse.julialang.org/u/aaowens)\
**Post date:** [February 7, 2020, 7:31pm UTC](https://discourse.julialang.org/t/how-would-i-write-this-pandas-column-filter-code-in-julia/34310/5 "2020-02-07T19:31:47Z")

</div>

I think this won’t work if some of the columns aren’t string type. How about

```julia
function hasmatch(col, regex)
    eltype(col) <: AbstractString || return false
    return any(x -> occursin(regex, x), col)
end
df = DataFrame(A = rand(10), B = "x", C = "y")
regex = r"x"
cols = [col for col in eachcol(df) if hasmatch(col, regex)]

```

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [February 8, 2020, 1:54am UTC](https://discourse.julialang.org/t/how-would-i-write-this-pandas-column-filter-code-in-julia/34310/6 "2020-02-08T01:54:03Z")

</div>

Could also add a conditional to @dmolina’s code:

```julia
cols = [k for k in names(df) if eltype(df[!,k]) <: AbstractString && any(occursin.(r"...", df[!,k]))]

```

Note: I also changed the column selection to `df[!,k]` rather than `df[:,k]`, since the latter makes a (unneeded) copy of the column.
