# Pandas.series.str.extract equivalent?

**URL:** <https://discourse.julialang.org/t/pandas-series-str-extract-equivalent/8016>\
**Category:** Data\
**Tags:** dataframes\
**Created:** [December 27, 2017, 4:35am UTC](https://discourse.julialang.org/t/pandas-series-str-extract-equivalent/8016 "2017-12-27T04:35:14Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![magrathean](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/magrathean/32/20389_2.png) [@magrathean](https://discourse.julialang.org/u/magrathean)\
**Post date:** [December 27, 2017, 4:35am UTC](https://discourse.julialang.org/t/pandas-series-str-extract-equivalent/8016/1 "2017-12-27T04:35:14Z")

</div>

How do I do the following pandas manipulation on a Julia Dataframe column:

```python
>>> s = Series(['a1', 'b2', 'c3'])
>>> s.str.extract('(?P<letter>[ab])(?P<digit>\d)')
  letter digit
0 a 1
1 b 2
2 NaN NaN

```

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [December 27, 2017, 2:19pm UTC](https://discourse.julialang.org/t/pandas-series-str-extract-equivalent/8016/2 "2017-12-27T14:19:30Z")

</div>

I am not aware of an inbuilt method, but to get you started.  
Julia support regexes and so the first parts would be:

```julia
data = ["a1", "b2", "c3"]
regex = r"(?P<letter>[ab])(?P<digit>\d)"
result = match.(regex, data) # notice the "dot" after match
3-element Array{Any,1}:
 RegexMatch("a1", letter="a", digit="1")
 RegexMatch("b2", letter="b", digit="2")
 nothing                                

```

So now you have to convert the `result` array into a `DataFrame`…  
Which is a bit annoying and the below is not the most efficient way of doing it.

```julia
getnames(m::RegexMatch) = collect(values(Base,PCRE.capture_names(m.regex.regex)))
getnames(m::Void) = String[]

columns = Symbol.(unique(reduce(append!, getnames.(result))))
df = DataFrame(fill(String, length(columns)), columns, 0)
for rm in result
    if rm === nothing
        push!(df, fill(NA, length(colums))
        continue
    end
    row = Any[]
    for column in columns
        if column ∉ getnames(rm)
            push!(row, rm[column])
        else
            push!(row, NA)
        end
    end
    push!(df, row)
end

```

---

<div class="post-metadata">

**Author:** ![magrathean](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/magrathean/32/20389_2.png) [@magrathean](https://discourse.julialang.org/u/magrathean)\
**Post date:** [December 27, 2017, 5:18pm UTC](https://discourse.julialang.org/t/pandas-series-str-extract-equivalent/8016/3 "2017-12-27T17:18:01Z")

</div>

@vchuravy looks rather tedious; thanks for the details nevertheless.
