# Pandas value\_counts() equivalent in base Julia or in Data frame package?

**URL:** <https://discourse.julialang.org/t/pandas-value-counts-equivalent-in-base-julia-or-in-data-frame-package/93620>\
**Category:** General Usage\
**Created:** [January 27, 2023, 10:16am UTC](https://discourse.julialang.org/t/pandas-value-counts-equivalent-in-base-julia-or-in-data-frame-package/93620 "2023-01-27T10:16:19Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![dove88](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dove88/32/46281_2.png) [@dove88](https://discourse.julialang.org/u/dove88)\
**Post date:** [January 27, 2023, 10:16am UTC](https://discourse.julialang.org/t/pandas-value-counts-equivalent-in-base-julia-or-in-data-frame-package/93620/1 "2023-01-27T10:16:19Z")

</div>

Hello  
I am interested in Datascience and now i am using Julia for Data analysis with python.  
I have a question,  
How can we get the count of each categorical values in an array(dataframe subset) in Julia?

For. Eg.

_we should get answer like this from a dataframe ‘df’,_

_cat 4_  
_dog 5_  
_cow 4_

Thank you

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [January 27, 2023, 10:31am UTC](https://discourse.julialang.org/t/pandas-value-counts-equivalent-in-base-julia-or-in-data-frame-package/93620/2 "2023-01-27T10:31:43Z")

</div>

Here are two ways:

```julia
julia> using DataFrames, StatsBase

julia> df = DataFrame(x = rand('a':'d', 100));

julia> countmap(df.x)
Dict{Char, Int64} with 4 entries:
  'c' => 23
  'a' => 25
  'd' => 25
  'b' => 27

julia> combine(groupby(df, :x), nrow => :count)
4×2 DataFrame
 Row │ x count
     │ Char Int64
─────┼─────────────
   1 │ c 23
   2 │ a 25
   3 │ b 27
   4 │ d 25

```

The first is a generic way which works everywhere and for all vectors but requires a different package, while the second is DataFrames specific and only uses the functionality already available in DataFrames.
