# Arrow's DictEncode to CategoricalArray?

**URL:** <https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451>\
**Category:** Data\
**Tags:** dataframes, arrow\
**Created:** [April 3, 2024, 5:37am UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451 "2024-04-03T05:37:47Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![elenev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elenev/32/18440_2.png) [@elenev](https://discourse.julialang.org/u/elenev)\
**Post date:** [April 3, 2024, 5:37am UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/1 "2024-04-03T05:37:47Z")

</div>

I am using Arrow.jl to store a dataframe to disk. The dataframe has some columns of `CategoricalArray{T}` type, where `T` is either `String` or `Int64`.

When I read the dataframe back in, the categorical-ness of these columns is not preserved, even though the Arrow.Table object does have those columns correctly marked as DictEncode’d?

Mechanically, I think this is because the DataFrames constructor works on any object inheriting the Tables.jl interface, whereas `Arrow.DictEncode` is an Arrow-specific implementation detail that the constructor is not aware of.

Would it be possible to generically convert such columns back to categorical arrays or would this require assumptions about the data (i.e., because multiple original Julia types can get Arrow-serialized with the DictEncode property)?

If not, I can still do the conversions back to CategoricalArray manually, but I notice that `Base.summarysize()` does not decrease and remains considerably larger than the `Base.summarysize()` of the original dataframe, before serialization.

EDIT: Actually, I just realized that the larger memory of the imported dataframe is due to all columns being much bigger in size, not the categoricals. So the previous paragraph doesn’t apply. This can be fixed by just converting them to Vectors from `Arrow.Primitive{T, Vector{T}}`, which is what they get read as. Hmm – why?

---

<div class="post-metadata">

**Author:** ![elenev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elenev/32/18440_2.png) [@elenev](https://discourse.julialang.org/u/elenev)\
**Post date:** [April 3, 2024, 3:51pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/2 "2024-04-03T15:51:25Z")

</div>

OP here.

I figured out the size issue, I think. `copy()`ing the imported dataframe using `df = Arrow.Table(path) |> DataFrame |> copy` shrinks `Base.summarysize()` back to the size of the original dataframe. Now that I think about it, I’m not even sure what `Base.summarysize(df)`'s output means for `df = Arrow.Table(path) |> DataFrame` since `df`’s columns are just views of data still on disk, if I understand correctly. So if `df` won’t be modified, it’s probably best not to worry about `Base.summarysize()`'s output. If `df` will be modified, then `copy()`ing is useful.

My original question – about preserving `CategoricalArray` types – remains. If it turns out to be impossible with Arrow, is there another lossless (i.e., not CSV) and stable (i.e., not generic serialization) method for storing dataframes on disk? Something akin to pickle in pandas?

---

<div class="post-metadata">

**Author:** ![dmbates](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmbates/32/44_2.png) [@dmbates](https://discourse.julialang.org/u/dmbates)\
**Post date:** [April 3, 2024, 5:32pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/3 "2024-04-03T17:32:07Z")

</div>

As I understand it, you are doing something like

```julia
julia> using Arrow, CategoricalArrays, DataFrames

julia> df = DataFrame(a = 1:4, b = string.('a':'d'), c = categorical(["x", "x", "y", "y"]))
4×3 DataFrame
 Row │ a b c    
     │ Int64 String Cat… 
─────┼─────────────────────
   1 │ 1 a x
   2 │ 2 b x
   3 │ 3 c y
   4 │ 4 d y

julia> afn = Arrow.write("./df.arrow", df)
"./df.arrow"

julia> df1 = DataFrame(Arrow.Table(afn))
4×3 DataFrame
 Row │ a b c      
     │ Int64 String String 
─────┼───────────────────────
   1 │ 1 a x
   2 │ 2 b x
   3 │ 3 c y
   4 │ 4 d y

julia> typeof(df1.c)
Arrow.DictEncoded{String, Int8, Arrow.List{String, Int32, Vector{UInt8}}}

```

It won’t be the case that you can “round trip” DataFrame → Arrow → DataFrame and get the same types. Is there a reason that you need a `CategoricalArray` instead of the `Arrow.DictEncoded` result. The `Arrow.DictEncoded` result can in some circumstances take up less storage than the `CategoricalArray`, because it uses the smallest signed integer type available for the `refarray` (`Int8` in this case).

`Arrow.DictEncoded` is more like a `PooledArray` than a `CategoricalArray` but often the distinctions are not important. They can be important for ordered categorical arrays. I think it is still the case that the `Arrow.Table` function does ignores whether `DictEncoded` arrays in the Arrow file have ordered categories.

---

<div class="post-metadata">

**Author:** ![elenev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elenev/32/18440_2.png) [@elenev](https://discourse.julialang.org/u/elenev)\
**Post date:** [April 3, 2024, 5:38pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/4 "2024-04-03T17:38:07Z")

</div>

I think it IS still the case because `DataFrame(afn::Arrow.Table)` dispatches `DataFrame(::Table)`. This is elegant composition because it means that as long as your package can get data in a tabular form, you don’t need to provide a DataFrame constructor for it. But it does mean that DataFrame won’t be aware of any format-specific metadata, such as Arrow’s `DictEncoded`. I think?

As to your first question, I’m using `CategoricalArray` because order matters for some applications, and because some of the statistical analysis I plan on doing with the data needs them to be categorical. But I suppose your implication is correct – the choice of formats for storing data should be mainly driven by performance and size considerations. Then once I start doing analysis, I can make whatever in-memory conversions are appropriate.

---

<div class="post-metadata">

**Author:** ![dmbates](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmbates/32/44_2.png) [@dmbates](https://discourse.julialang.org/u/dmbates)\
**Post date:** [April 3, 2024, 5:43pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/5 "2024-04-03T17:43:35Z")

</div>

If you use the formula/data syntax provided by `StatisticalModels.jl` and specify the contrasts to be used for your categorical columns, you can impose the order there. Often you just need to specify the base level in the contrasts, e.g.

```julia
contrasts = Dict(:c => EffectsCoding(base="y"))

```

That formula/data syntax is used by GLM.jl and MixedModels.jl

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [April 3, 2024, 8:26pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/6 "2024-04-03T20:26:50Z")

</div>

If you want to save and reload data frames without losing any type information, you can use [JLD2.jl](https://github.com/JuliaIO/JLD2.jl). Otherwise you’ll have to recreate `CategoricalArray`s manually like this:

```julia
using DataAPI
mapcols!(df1) do col
     col isa Arrow.DictEncoded ? levels!(categorical(col), DataAPI.refpool(col)) : col
end

```

---

<div class="post-metadata">

**Author:** ![elenev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elenev/32/18440_2.png) [@elenev](https://discourse.julialang.org/u/elenev)\
**Post date:** [April 3, 2024, 8:27pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/7 "2024-04-03T20:27:44Z")

</div>

Is JLD2 serialization compatible across (minor) version changes of Julia or the JLD2 package?

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [April 3, 2024, 9:23pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/8 "2024-04-03T21:23:08Z")

</div>

Yes, it should. Though it could break if the layout of the `DataFrame`, `CategoricalArray` or another type used in the data changes.

It’s kind of unfortunate that reloading data from Arrow is tricky for `CategoricalArray`, as Arrow is a stable format with all the necessary functionality under the hood. Maybe we could have an argument allowing to do this a bit more easily when reloading an Arrow file.

---

<div class="post-metadata">

**Author:** ![palday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palday/32/12640_2.png) [@palday](https://discourse.julialang.org/u/palday)\
**Post date:** [April 4, 2024, 7:47pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/9 "2024-04-04T19:47:08Z")

</div>

This feels like a great case for using `ArrowTypes` and a package extension to `CategoricalArrays`.

---

<div class="post-metadata">

**Author:** ![elenev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elenev/32/18440_2.png) [@elenev](https://discourse.julialang.org/u/elenev)\
**Post date:** [April 4, 2024, 10:27pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/10 "2024-04-04T22:27:43Z")

</div>

@palday Can you explain this in a bit more detail?

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [April 5, 2024, 7:55am UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/11 "2024-04-05T07:55:11Z")

</div>

Indeed, thanks for the pointer. Though I’ve experimented a bit and I couldn’t get Arrow to generate a `CategoricalArray` when loading, only a `Vector{<:CategoricalValue}`. The problem is that ArrowTypes methods are defined for scalars (`CategoricalValue` here), but I couldn’t find a way to have it operate at the array level. Any ideas?

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [April 7, 2024, 2:09pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/12 "2024-04-07T14:09:02Z")

</div>

After some additional investigation, I’ve found code that gives the intended result. But it’s a bit hacky as I need to override the `DictEncoding` constructor so that it stores `CategoricalValue` objects. Otherwise they would have to be created on the fly for each element and they couldn’t share the same pool. Maybe a more generic API could be added to Arrow to make this cleaner @quinnj?

```julia
using Arrow, ArrowTypes, CategoricalArrays, DataFrames

ArrowTypes.ArrowType(::Type{<:CategoricalValue}) = Arrow.DictEncoded

ArrowTypes.arrowname(::Type{<:CategoricalValue}) = Symbol("JuliaLang.CategoricalArray")
ArrowTypes.arrowmetadata(::Type{CategoricalValue{T, R}}) where {T, R} = string(R)

const REFTYPES = Dict(string(T) => T for T in (Int128, Int16, Int32, Int64, Int8, UInt128, UInt16, UInt32, UInt64, UInt8))
function ArrowTypes.JuliaType(::Val{Symbol("JuliaLang.CategoricalArray")}, ::Type{S}, meta::String) where S
    R = REFTYPES[meta]
    return CategoricalValue{S, R}
end

function Arrow.DictEncoding{V,S,A}(id, data::Arrow.List{U, O, B}, isOrdered, metadata) where {T, V<:CategoricalValue{T}, S, O, A, B, U}
    newdata = Arrow.List{T, O, B}(data.arrow, data.validity, data.offsets, data.data, data.ℓ, data.metadata)
    catdata = CategoricalVector{T}(newdata, levels=newdata)
    return Arrow.DictEncoding{V,S,typeof(catdata)}(id, catdata, isOrdered, metadata)
end

function Arrow.DictEncoding{V,S,A}(id, data::Arrow.Primitive{U, B}, isOrdered, metadata) where {T, V<:CategoricalValue{T}, S, A, B, U}
    newdata = Arrow.Primitive{T, B}(data.arrow, data.validity, data.data, data.ℓ, data.metadata)
    catdata = CategoricalVector{T}(newdata, levels=newdata)
    return Arrow.DictEncoding{V,S,typeof(catdata)}(id, catdata, isOrdered, metadata)
end

function Base.copy(x::Arrow.DictEncoded{V}) where {T, R, V<:CategoricalValue{T, R}}
    pool = CategoricalArrays.CategoricalPool{T, R}(x.encoding.data)
    inds = x.indices
    refs = similar(inds, R)
    refs .= inds .+ one(R)
    return CategoricalVector{T}(refs, pool)
end

```

---

<div class="post-metadata">

**Author:** ![palday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palday/32/12640_2.png) [@palday](https://discourse.julialang.org/u/palday)\
**Post date:** [April 10, 2024, 5:14pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/13 "2024-04-10T17:14:37Z")

</div>

This is pretty much exactly what I was expecting.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [April 17, 2024, 12:36pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/14 "2024-04-17T12:36:02Z")

</div>

And do you think it’s clear enough to add to CategoricalArrays?

---

<div class="post-metadata">

**Author:** ![nikolays](https://avatars.discourse-cdn.com/v4/letter/n/45deac/32.png) [@nikolays](https://discourse.julialang.org/u/nikolays)\
**Post date:** [February 7, 2025, 2:15pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/15 "2025-02-07T14:15:51Z")

</div>

Awesome! Thank you very much. It worked as needed, and it should be pushed upstream to `Arrow.jl` or elsewhere appropriate.

The file written by Arrow.jl and @nalimilan code was successfully loaded in R with `arrow::read_ipc_file` function. In R, the categorical column became the factor column out of the box with the same level order.

Even if categories are not true ordinal, the order can be important for printing purposes.

---

<div class="post-metadata">

**Author:** ![nikolays](https://avatars.discourse-cdn.com/v4/letter/n/45deac/32.png) [@nikolays](https://discourse.julialang.org/u/nikolays)\
**Post date:** [February 7, 2025, 3:36pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/16 "2025-02-07T15:36:51Z")

</div>

Need to add handling of missing data to make it perfect. Here is the test code, which is not working:

```jl
# nalimilan addition
using Arrow, ArrowTypes, CategoricalArrays, DataFrames

ArrowTypes.ArrowType(::Type{<:CategoricalValue}) = Arrow.DictEncoded

ArrowTypes.arrowname(::Type{<:CategoricalValue}) = Symbol("JuliaLang.CategoricalArray")
ArrowTypes.arrowmetadata(::Type{CategoricalValue{T, R}}) where {T, R} = string(R)

const REFTYPES = Dict(string(T) => T for T in (Int128, Int16, Int32, Int64, Int8, UInt128, UInt16, UInt32, UInt64, UInt8))
function ArrowTypes.JuliaType(::Val{Symbol("JuliaLang.CategoricalArray")}, ::Type{S}, meta::String) where S
    R = REFTYPES[meta]
    return CategoricalValue{S, R}
end

function Arrow.DictEncoding{V,S,A}(id, data::Arrow.List{U, O, B}, isOrdered, metadata) where {T, V<:CategoricalValue{T}, S, O, A, B, U}
    newdata = Arrow.List{T, O, B}(data.arrow, data.validity, data.offsets, data.data, data.ℓ, data.metadata)
    catdata = CategoricalVector{T}(newdata, levels=newdata)
    return Arrow.DictEncoding{V,S,typeof(catdata)}(id, catdata, isOrdered, metadata)
end

function Arrow.DictEncoding{V,S,A}(id, data::Arrow.Primitive{U, B}, isOrdered, metadata) where {T, V<:CategoricalValue{T}, S, A, B, U}
    newdata = Arrow.Primitive{T, B}(data.arrow, data.validity, data.data, data.ℓ, data.metadata)
    catdata = CategoricalVector{T}(newdata, levels=newdata)
    return Arrow.DictEncoding{V,S,typeof(catdata)}(id, catdata, isOrdered, metadata)
end

function Base.copy(x::Arrow.DictEncoded{V}) where {T, R, V<:CategoricalValue{T, R}}
    pool = CategoricalArrays.CategoricalPool{T, R}(x.encoding.data)
    inds = x.indices
    refs = similar(inds, R)
    refs .= inds .+ one(R)
    return CategoricalVector{T}(refs, pool)
end

function Base.copy(x::Arrow.DictEncoded{Union{Missing, V}}) where {T, R, V<:CategoricalValue{T, R}}
    @info "lets try"
    pool = CategoricalArrays.CategoricalPool{T, R}(x.encoding.data)
    inds = x.indices
    refs = similar(inds, R)
    refs .= inds .+ one(R)
    return CategoricalVector{T}(refs, pool)
end

# test code
df = DataFrame(
    col1 = categorical(["A","B","C","A","A","C", "C"], ordered=false,compress=true),
    col2 = categorical(["A","B","C","A","A",missing, "C"], ordered=false,compress=true)
)

Arrow.write("df.arrow", df, compress=:zstd)
atab = Arrow.Table("df.arrow")
df2 = DataFrame(atab; copycols=true) # failed here on copy of col2 

```

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [February 7, 2025, 10:43pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/17 "2025-02-07T22:43:48Z")

</div>

This code seems to work, but I haven’t checked it carefully yet and it’s not complete. I should really get this into CategoricalArrays or Arrow.jl.

```julia
const CATARRAY_ARROWNAME = Symbol("JuliaLang.CategoricalArray")
ArrowTypes.arrowname(::Type{<:CategoricalValue}) = CATARRAY_ARROWNAME
ArrowTypes.arrowmetadata(::Type{CategoricalValue{T, R}}) where {T, R} = string(R)

ArrowTypes.arrowname(::Type{Union{<:CategoricalValue, Missing}}) = CATARRAY_ARROWNAME
ArrowTypes.arrowmetadata(::Type{Union{CategoricalValue{T, R}, Missing}}) where {T, R} = string(R)

const REFTYPES = Dict(string(T) => T for T in (Int128, Int16, Int32, Int64, Int8, UInt128, UInt16, UInt32, UInt64, UInt8))
function ArrowTypes.JuliaType(::Val{Symbol("JuliaLang.CategoricalArray")}, ::Type{S}, meta::String) where S
    R = REFTYPES[meta]
    return CategoricalValue{S, R}
end

function Arrow.DictEncoding{V,S,A}(id, data::Arrow.List{U, O, B}, isOrdered, metadata) where {T, R, V<:CategoricalValue{T,R}, S, O, A, B, U}
    newdata = Arrow.List{T, O, B}(data.arrow, data.validity, data.offsets, data.data, data.ℓ, data.metadata)
    catdata = CategoricalVector{T,R}(newdata, levels=newdata)
    return Arrow.DictEncoding{V,S,typeof(catdata)}(id, catdata, isOrdered, metadata)
end

function Arrow.DictEncoding{V,S,A}(id, data::Arrow.Primitive{U, B}, isOrdered, metadata) where {T, R, V<:CategoricalValue{T,R}, S, A, B, U}
    newdata = Arrow.Primitive{T, B}(data.arrow, data.validity, data.data, data.ℓ, data.metadata)
    catdata = CategoricalVector{T,R}(newdata, levels=newdata)
    return Arrow.DictEncoding{V,S,typeof(catdata)}(id, catdata, isOrdered, metadata)
end

function Arrow.DictEncoding{Union{Missing,V},S,A}(id, data::Arrow.List{U, O, B}, isOrdered, metadata) where {T, R, V<:CategoricalValue{T,R}, S, O, A, B, U}
    newdata = Arrow.List{Union{Missing,T}, O, B}(data.arrow, data.validity, data.offsets, data.data, data.ℓ, data.metadata)
    levels = collect(skipmissing(newdata))
    catdata = CategoricalVector{Union{Missing,T},R}(newdata, levels=levels)
    return Arrow.DictEncoding{Union{Missing,V},S,typeof(catdata)}(id, catdata, isOrdered, metadata)
end

function Arrow.DictEncoding{Union{Missing,V},S,A}(id, data::Arrow.Primitive{U, B}, isOrdered, metadata) where {T, R, V<:CategoricalValue{T,R}, S, A, B, U}
    newdata = Arrow.Primitive{Union{Missing,T}, B}(data.arrow, data.validity, data.data, data.ℓ, data.metadata)
    levels = collect(skipmissing(newdata))
    catdata = CategoricalVector{Union{Missing,T},R}(newdata, levels=levels)
    return Arrow.DictEncoding{Union{Missing,V},R,typeof(catdata)}(id, catdata, isOrdered, metadata)
end

function Base.copy(x::Arrow.DictEncoded{V}) where {T, R, V<:CategoricalValue{T, R}}
    pool = CategoricalArrays.CategoricalPool{T, R}(x.encoding.data)
    inds = x.indices
    refs = similar(inds, R)
    refs .= inds .+ one(R)
    return CategoricalVector{T}(refs, pool)
end

function Base.copy(x::Arrow.DictEncoded{Union{Missing, V}}) where {T, R, V<:CategoricalValue{T, R}}
    levels = collect(skipmissing(x.encoding.data))
    pool = CategoricalArrays.CategoricalPool{T, R}(levels)
    inds = x.indices
    refs = similar(inds, R)
    if ismissing(x.encoding.data[1])
        refs .= inds
    elseif ismissing(x.encoding.data[end])
        n = length(x.encoding.data) - 1
        refs .= ifelse.(inds .== n, zero(R), inds .+ one(R))
    else
        throw(ErrorException("not implemented"))
    end
    return CategoricalVector{Union{Missing,T}}(refs, pool)
end

```

EDIT: I’ve improved the implementation a bit

---

<div class="post-metadata">

**Author:** ![nikolays](https://avatars.discourse-cdn.com/v4/letter/n/45deac/32.png) [@nikolays](https://discourse.julialang.org/u/nikolays)\
**Post date:** [February 10, 2025, 5:33pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/18 "2025-02-10T17:33:11Z")

</div>

That is excellent and is working within Julia as expected. And R read it as well.

However, during the attempt to read it in Python:

```python
import pyarrow
import pyarrow.ipc
import pandas as pd

# following is an error ("Categorical categories cannot be null")
with pyarrow.ipc.open_file("df.arrow") as reader:
    df2 = reader.read_pandas()

# can read plain arrow format
with pyarrow.ipc.open_file("df.arrow") as reader:
    batches = reader.get_batch(0)

```

I got “Categorical categories cannot be null”, which is particularly funny as pandas docs said “In contrast to R’s `factor` function, categorical data is not converting input values to strings”.

Small experimentation with pyarrow shows that for seemless pandas integration the arrow dict should be:

```julia
pyarrow.RecordBatch
A: dictionary<values=string, indices=int8, ordered=0>
----
A: -- dictionary:
["AA","Bb"]-- indices:
[null,null,1,null,0]

```

and not

```julia
pyarrow.RecordBatch
col1: dictionary<values=string, indices=int8, ordered=0> not null
col2: dictionary<values=string, indices=int8, ordered=0>
----
col1: -- dictionary:
["A","B","C"]-- indices:
[0,1,2,0,0,2,2]
col2: -- dictionary:
[null,"A","B","C"]-- indices:
[1,2,3,1,1,null,3]

```

That is `null` should be outsize of indices

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [February 11, 2025, 2:30pm UTC](https://discourse.julialang.org/t/arrows-dictencode-to-categoricalarray/112451/19 "2025-02-11T14:30:28Z")

</div>

Interesting. AFAICT both representations are allowed (and pyarrow [allows](https://arrow.apache.org/docs/python/generated/pyarrow.compute.dictionary_encode.html) choosing which one you want). I don’t think Arrow.jl allows this currently though, it even has a hack to add `missing` to the pool for CategoricalArrays:

> <https://github.com/apache/arrow-julia/blob/c12899b979f4195788e4d69047fc2c2cd1f81e71/src/arraytypes/dictencoding.jl#L231>

AFAIK having `missing` in the dictionary is more efficient, as least in Julia it avoids allocating additional space to track missing values.

Could be worth filing an issue against pandas.
