# Reading different types of data in Julia

**URL:** <https://discourse.julialang.org/t/reading-different-types-of-data-in-julia/97906>\
**Category:** General Usage\
**Created:** [April 25, 2023, 3:47pm UTC](https://discourse.julialang.org/t/reading-different-types-of-data-in-julia/97906 "2023-04-25T15:47:01Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![yvikhlya](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yvikhlya/32/3753_2.png) [@yvikhlya](https://discourse.julialang.org/u/yvikhlya)\
**Post date:** [April 25, 2023, 3:47pm UTC](https://discourse.julialang.org/t/reading-different-types-of-data-in-julia/97906/1 "2023-04-25T15:47:01Z")

</div>

Hello All, I am new to Julia, coming from python. I have collections of observational data and output from GCM model in netcdf format, but of different “types” (different grids, frequency, conventions etc, not julia types), which require slightly different treatment after reading from disk, before feeding these data to analysis utils. This may include renaming coordinates, adding missing grid metrics which is required by analysis tools, etc. A reading function should be able to take any “type” of collection as input and return NCDataset (or similar) object with all required metadata/metrics included and coordinates/variable names following the same convention. No regridding/resampling is needed at this step, just putting data to same format. I can think of several ways of how to do this in julia:

1. Using ‘if’ blocks,

```julia
function read_collection(colname, ...; coltype="LATLON")
  if coltype=="LATLON"
    ds=read_latlon(colname)
  elseif coltype="Tripolar" 
    ds=read_tripolar(colname)
  ......................
  end
  return ds
end

```

1. Using dictionary

```julia
function read_collection(colname, ...; coltype="LATLON")
  readers=Dict("LATLON" => read_latlon, "Tripolar" => read_tripolar, ...)
  return readers[coltype](colname)
end

```

I do something like this in python.  
3. Dispatch on a dummy type

```julia
struct LatLon
end

struct Tripolar
end

function read_collection(colname, ::LatLon)
  return read_latlon(colname)
end

function read_collection(colname, ::Tripolar)
  return read_tripolar(colname)
end

```

1. Dispatch on a value type

```julia
function read_collection(colname, ::Val{"LATLON"})
  return read_latlon(colname)
end

function read_collection(colname, ::Val{"Tripolar"})
  return read_tripolar(colname)
end

```

Which way would be preferable in julia in terms of performance and ease of expanding to more collection types in the future?

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 25, 2023, 3:58pm UTC](https://discourse.julialang.org/t/reading-different-types-of-data-in-julia/97906/2 "2023-04-25T15:58:07Z")

</div>

If you really want to handle all kinds of geospatial data arrangements, take a look at the `georef` function from the GeoStats.jl stack:

[https://juliaearth.github.io/GeoStats.jl/stable/data.html](https://juliaearth.github.io/GeoStats.jl/stable/data.html)

We still need to work on some important details like CRS, but already have most flexible domain types you can possibly need in practice, including, grids of cells, unstructured meshes, point sets, geometry sets, etc.

If you follow this approach with `georef` you will gain tons of functionalities and transforms for free, including

[https://juliaearth.github.io/GeoStats.jl/stable/transforms.html](https://juliaearth.github.io/GeoStats.jl/stable/transforms.html)

[https://juliaearth.github.io/GeoStats.jl/stable/splitapplycombine.html](https://juliaearth.github.io/GeoStats.jl/stable/splitapplycombine.html)

and visualization recipes with both the Plots.jl and the Makie.jl stacks, provided by GeoStatsPlots.jl and GeoStatsViz.jl, respectively.

---

<div class="post-metadata">

**Author:** ![Alexander-Barth](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alexander-barth/32/3692_2.png) [@Alexander-Barth](https://discourse.julialang.org/u/Alexander-Barth)\
**Post date:** [April 28, 2023, 11:59am UTC](https://discourse.julialang.org/t/reading-different-types-of-data-in-julia/97906/3 "2023-04-28T11:59:08Z")

</div>

To me the 3rd option seems to be the most idiomatic and is similar to for example to the `parse` and `read` functions:

[https://docs.julialang.org/en/v1/base/io-network/#Base.read](https://docs.julialang.org/en/v1/base/io-network/#Base.read)

When the compiler know the types, the function calls can be inlined.  
Option 1 and 2 are probably not type-stable.

---

<div class="post-metadata">

**Author:** ![yvikhlya](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yvikhlya/32/3753_2.png) [@yvikhlya](https://discourse.julialang.org/u/yvikhlya)\
**Post date:** [April 28, 2023, 4:06pm UTC](https://discourse.julialang.org/t/reading-different-types-of-data-in-julia/97906/4 "2023-04-28T16:06:59Z")

</div>

Thanks All, I’ll try to go to option 3 and look into GeoStats later.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [April 28, 2023, 4:53pm UTC](https://discourse.julialang.org/t/reading-different-types-of-data-in-julia/97906/5 "2023-04-28T16:53:03Z")

</div>

I smell, from those mentions to latlon, that GMT.jl would make your life simpler.
