# Random access to rows of a table

**URL:** <https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386>\
**Category:** Machine Learning\
**Created:** [March 4, 2022, 1:18am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386 "2022-03-04T01:18:49Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [March 4, 2022, 1:18am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/1 "2022-03-04T01:18:49Z")

</div>

I’d like to see certain enhancements to the tabular data ecosystem, and this post is an appeal to maintainers of table-providing packages (eg, DataFrames) to articulate the form they should like these to take, or give other feedback.

A fundamental operation in machine learning is a the extraction of some arbitrary subset of observations in the training set (resampling). Providing **efficient random access** to observations, wherever possible, is therefore crucial.

The Tables.jl interface is [very widely used](https://github.com/JuliaData/Tables.jl/blob/main/INTEGRATIONS.md) in Julia data science. Unfortunately, it only provides _iteration_ over rows (=observations) of a table. So all ML tooling that is designed to work with arbitrary tables that implement the interface are currently stuck with generally inefficient resampling implementations, even for in-memory data. I am aware of several unfortunate workarounds to this issue, not limited to my own contributions to them!

There have been several requests at Tables.jl to add random access methods to the API (with the obvious slow fallbacks) but these have not met with success, as the maintainers understandably wish to limit the scope of the project.

On the other hand, the older LearnBase.jl project provides a well-thought out interface for data containers supporting random access to obervations (more general than tables, eg, a collection of image files). MLDataPattern.jl built on top of that to provide a lot of functionality for resampling in ML (eg, stratified CV). The very nice package DataLoaders.jl also builds on the LearnBase.jl interface to manage data that does not fit into memory. DataLoaders.jl is widely used by the deep learning community (Flux users, FastAI, etc).

[Efforts](https://github.com/JuliaML/MLUtils.jl/issues/2) by some in the ML community are underway to re-organize and re-vitalize the LearnBase API. If this interface could [play well with tables](https://github.com/JuliaML/MLUtils.jl/issues/61), this could help to unify disparate efforts in the julia stats/ml community.

While including tables in the above efforts might be possible, I expect a better option is to extend the Tables.jl interface in a new standalone, lightweight package providing extra methods for row (and other) random access methods. The idea is that existing tables with better-than-iteration random-access **implement the new methods natively**. Would table-providers be prepared to get behind such an effort? How should the API look to get maximum buy-in? Do people have other ideas for achieving the same goals?

@darsnack @samuel_okon

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [March 4, 2022, 2:16am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/2 "2022-03-04T02:16:45Z")

</div>

Sorry, I am not sure if I understand correctly, each column of table is a `Vector`, they support random access. Do you want a convenience function that accesses the same index in all columns and return the result?

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [March 4, 2022, 5:12am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/3 "2022-03-04T05:12:37Z")

</div>

I want efficient row-slicing: Given table `X` and abstract vector of indices `I`, chosen from `1:nrows(X)`, I want a new table `X2` such that the `i`th row of `X2` is the `I[i]`th of `X`. I don’t require that `X2` have the same type as `X`, but iterating over the rows of `X2` should generate objects of the same type as iterating of those of `X` (iteration in the sense of Tables.jl). Or something close to that.

Basically, I want an interface for tables which puts them into the “data container” framework described in great detail [here](https://mldatapatternjl.readthedocs.io/en/latest/documentation/container.html), where for tables “observation” = “row”.

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [March 4, 2022, 5:14am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/4 "2022-03-04T05:14:29Z")

</div>

And I don’t want to have to materialise entire columns for tables for which that is costly, eg so-called row tables, which are vectors of named tuples.

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [March 4, 2022, 8:21am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/5 "2022-03-04T08:21:26Z")

</div>

I’m probably misunderstanding one of the constraints

```julia
julia> df = DataFrame(a=repeat(1:300, 1000), b=repeat(1:6, 50000))
julia> selection = rand(size(df,1)) .> 0.2
300000-element BitVector:
...
julia> @time subset(df, :a=>(a)->selection)
  0.047807 seconds (33.14 k allocations: 7.502 MiB, 90.23% compilation time)
240106×2 DataFrame

julia> @benchmark subset(df, :a=>(a)->selection; view=true)
BenchmarkTools.Trial: 3950 samples with 1 evaluation.
 Range (min … max): 941.200 μs … 7.278 ms ┊ GC (min … max): 0.00% … 80.31%
 Time (median): 1.029 ms ┊ GC (median): 0.00%
 Time (mean ± σ): 1.259 ms ± 978.274 μs ┊ GC (mean ± σ): 14.66% ± 15.18%

  █▇▅▃▁ ▁ ▁
  ███████▇▆▇▃▄▅▅▁▁▁▄▃▁▃▄▁▁▁▁▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▃▄▁▅▇███ █
  941 μs Histogram: log(frequency) by time 6.13 ms <

 Memory estimate: 1.88 MiB, allocs estimate: 158.
```

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [March 4, 2022, 5:13pm UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/6 "2022-03-04T17:13:46Z")

</div>

`subset` is a function from DataFrames for `DataFrame`s. Not all Tables-compatible types that support random access are DataFrames, and I think most consumers of tables (such as MLUtils) would prefer not to take a dep on DataFrames.jl.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [March 4, 2022, 5:42pm UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/7 "2022-03-04T17:42:23Z")

</div>

I’m not sure it has what you want, but maybe this could serve as inspiration for a row-based alternative?

[https://github.com/JuliaData/TableOperations.jl](https://github.com/JuliaData/TableOperations.jl)

---

<div class="post-metadata">

**Author:** ![lewis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lewis/32/5217_2.png) [@lewis](https://discourse.julialang.org/u/lewis)\
**Post date:** [March 4, 2022, 5:57pm UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/8 "2022-03-04T17:57:34Z")

</div>

I needed to do just this. Simply create a vector of the row indices of the table as `rowidx`. Before each pass of training just shuffle!() it. Very fast. Then write your training loop `for i in rowidx`.

```julia
X = rand(9, 10_000) # training data--fake because it wouldn't be rand
rowidx = collect(1:size(X, 1))
shuffle!(rowidx)
for i in rowidx
     # do the training
end

```

This pseudo code is not a complete representation of training obviously, but it provides a fast way to reorder or sample the training matrix. Because the entire vector rowidx is shuffled, you can do a sample with a subset of rowidx: `for i in rowidx[1:1_000]`. This does not repeat the first 1000 samples of X because of shuffle. Some items might reappear because of randomness.

Would this work for you?

---

<div class="post-metadata">

**Author:** ![junder873](https://avatars.discourse-cdn.com/v4/letter/j/e95f7d/32.png) [@junder873](https://discourse.julialang.org/u/junder873)\
**Post date:** [March 4, 2022, 6:17pm UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/9 "2022-03-04T18:17:22Z")

</div>

I think the package that is closest to what you are looking for is [AxisArrays.jl](https://github.com/JuliaArrays/AxisArrays.jl). It is very fast at getting random slices of data. The biggest thing it doesn’t fit is it is not a Tables compatible.

---

<div class="post-metadata">

**Author:** ![Akatz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/akatz/32/15164_2.png) [@Akatz](https://discourse.julialang.org/u/Akatz)\
**Post date:** [March 4, 2022, 7:13pm UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/10 "2022-03-04T19:13:16Z")

</div>

I think @ablaom is looking for a more general interface rather than a specific type

Related work: [https://github.com/FluxML/FastAI.jl/pull/26](https://github.com/FluxML/FastAI.jl/pull/26)

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [March 6, 2022, 10:35pm UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/11 "2022-03-06T22:35:20Z")

</div>

Yes, I’m asking about a general interface, not a specific table type.

The Tables.jl API allows one to iterate over the rows of any table type on [this list](https://github.com/JuliaData/Tables.jl/blob/main/INTEGRATIONS.md) using the same syntax. I’m looking for the same functionality for row slicing. For any table, I want to row slice using a common syntax. For specific table types that buy into the interface, slicing will be as efficient as possible for that type.

Thanks to all for the comments so far.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [March 9, 2022, 10:02pm UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/12 "2022-03-09T22:02:07Z")

</div>

Reference to previous discussions in Tables.jl: [https://github.com/JuliaData/Tables.jl/issues/48](https://github.com/JuliaData/Tables.jl/issues/48), [https://github.com/JuliaData/Tables.jl/issues/123](https://github.com/JuliaData/Tables.jl/issues/123).

What kind of performance would you expect from such a function? For example, random access of rows in a `DataFrame` is relatively fast, but if that’s a time-critical part of your algorithm, you’d probably better convert the table to a named tuple of columns or to a matrix. So what the API should look like depends on the use case.

IMO this kind of thing would deserve living in Tables.jl even if only a subset of table types can support it. I don’t think @quinnj was totally opposed to that when it was discussed, but nobody has proposed a concrete API for inclusion in Tables.jl so far.

---

<div class="post-metadata">

**Author:** ![ablaom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ablaom/32/4889_2.png) [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Post date:** [March 10, 2022, 12:08am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/13 "2022-03-10T00:08:10Z")

</div>

Thanks @nalimilan for your valued input, and for linking in the Tables discussion. My apologies for misrepresenting the position of Tables.jl. I am remembering (or misremembering 😳 ) a different disucussion. I agree an extension within Tables.jl would be ideal.

> [@nalimilan](#):
>
> What kind of performance would you expect from such a function? For example, random access of rows in a `DataFrame` is relatively fast, but if that’s a time-critical part of your algorithm, you’d probably better convert the table to a named tuple of columns or to a matrix. So what the API should look like depends on the use case.

I want the best performance the particular table I am using has to offer, assuming the provider of the table type has implemented the new interface, and otherwise, it should just work.

I don’t understand “So what the API should look like depends on the use case.” Could you elaborate?

I’m not sure if this is what you are getting at, but the reasonable observation that models should just convert tables to the optimal representation for their purposes does not preclude the usefulness of the proposal. For one thing, an operation like resampling might occur external to the model, where changing the representation of the data may be premature. The operation of splitting a table into test observations and training observations is a simple example. I don’t want to be forced at this point of the work-flow to make a decision about the best representation of the data. Indeed, I may be comparing multiple models which have contradictory optimal requirements (KNN wants a matrix where observations are columns, while RandomForest is better if observations are rows). The ultimate choice of representation, and when to change it, could be a complex one, but one useful option is “give me a slice of observations in my table, but keep the representation unchanged if possible” and do it as fast as you can, please.

It’s true that models can be made to take more responsibility for resampling. There is [an attempt](https://alan-turing-institute.github.io/MLJ.jl/dev/adding_models_for_general_use/#Implementing-a-data-front-end) to do this in MLJ. But I haven’t seen this more generally.

One final related thought: I don’t have benchmarks to prove it, but I’m guessing that there are cases (eg, ensemble bagging) where the overheads for poor resampling are getting close to the cost training the model, because that model is very simple (eg, a shallow decision tree).

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [March 10, 2022, 12:36am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/14 "2022-03-10T00:36:23Z")

</div>

maybe MLJ can use the [Home · Tables.jl](https://tables.juliadata.org/stable/#Tables.partitioner) ?

and hope the data source (I have a lazy data source now actually…) implement efficiency partition?

notice some source simply prohibit efficient random access, for example when columns are stored in small chunks on disk, individually compressed. (like Apache Parquet)

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [March 10, 2022, 9:20am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/15 "2022-03-10T09:20:37Z")

</div>

> [@ablaom](#):
>
> I want the best performance the particular table I am using has to offer, assuming the provider of the table type has implemented the new interface, and otherwise, it should just work.
> 
> I don’t understand “So what the API should look like depends on the use case.” Could you elaborate?

For example, for a `DataFrame`, if you just want to split the data in a few subsets, this can be done efficiently by indexing the `DataFrame` object with a vector of indices; so DataFrames could implement that generic API by simply calling its `getindex` method. Even for other table types which are not as suited to random access as `DataFrame` (like SQL tables), indexing with a vector could probably be implemented with an acceptable performance if you only take a few subsets.

On the contrary, if what you want is index single rows repeatedly in a performance-critical part of the code, then maybe you’d better work on a named tuple of columns (or a matrix) than on a `DataFrame`; the API would have to include e.g. a `fastindexingtable` function that you would call to get that named tuple of columns before performing repeated indexing. Likewise, for types such as SQL tables, the only solution would be to copy everything in memory first.

From your description, it seems you care about the first case? Indeed it seems the simplest to support.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [March 10, 2022, 10:20am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/16 "2022-03-10T10:20:36Z")

</div>

I think that the request makes sense, but I would like to better understand the problem.

My current understanding is:

- Tables.jl provides `Tables.rows` method that guarantees row iterator;
- Base Julia defines `AbstractVector` which provides random access API.

Therefore my current workflow is the following if I need random access to rows:

1. get `Tables.rows` object from a table;
2. check if it is an `AbstractVector`; if it is then we are done;
3. if it is not it means that the source table explicitly opts out from providing random access (for whatever reason); then I need to `collect` what `Tables.rows` returns to have an object supporting random access.

For example if your table is a data frame then by using `Tables.rows` you get an `AbstractVector`.

I assume you would want to improve over this workflow, but could you please specify concretely how?  
(in particular note that we have to respect the fact that some table type might explicitly opt-out from providing random access and then `collect` or similar option would be required anyway)

Having said that we might review current implementations of `Tables.rows` method to make sure it returns `AbstractVector` in cases it makes sense to do so.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [March 10, 2022, 11:23am UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/17 "2022-03-10T11:23:15Z")

</div>

That is a super helpful comment @bkamins, thanks!

I sympathize with the issues that @ablaom raised here, specially the one where consumers of the Tables.jl API need to be very explicit about their intentions and convert between different table types too early in the middle of pipelines. It would be nice if we had a set of traits to determine when and how to do things, perhaps even higher-level functions that rely on the current low-level API.

On the other hand, I also sympathize with the idea that some table types aren’t suited for random access. Perhaps we could make this explicit with a trait function? Something that could help consumers do the best they can with a given table not implemented by them. Documentation is one part of the story, but I guess we need something more programmatic to more easily get top performance.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [March 10, 2022, 12:28pm UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/18 "2022-03-10T12:28:56Z")

</div>

@juliohm - the point is that I believe we already have this trait and it is `AbstractVector`. If you want fast random access you essentially need to support this interface (with getting index, making views etc.).

The only challenge might be that some providers of Tables.jl tables might subtype from something else than `AbstractVector` and Julia does not support multiple inheritance currently.

Therefore I have written my comment to kind of “challenge” maintainers of the packages and learn if anyone has a problem with subtyping from `AbstractVector`. If there is no such problem then we are done and the only thing package maintainers need to do is to make sure that `Tables.rows` returns a subtype of `AbstractVector` if it can be supported (I believe it should not be problematic - this is what we did in DataFrames.jl).

However, if people comment that `Tables.rows` must return values that cannot be subtypes of `AbstractVector` (I doubt it but I might not know something) then we indeed need some trait like `Tables.isabstractvector` which would return `true` even for non-`AbstractVector` types that support `AbstractVector` interface.

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [March 10, 2022, 1:35pm UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/19 "2022-03-10T13:35:48Z")

</div>

> [@juliohm](#):
>
> On the other hand, I also sympathize with the idea that some table types aren’t suited for random access.

Tables are basically a collection of `Vector`s, I have yet not understood why people think they are not suited for random access.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [March 10, 2022, 1:46pm UTC](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386/20 "2022-03-10T13:46:53Z")

</div>

> [@bkamins](#):
>
> Therefore I have written my comment to kind of “challenge” maintainers of the packages and learn if anyone has a problem with subtyping from `AbstractVector`

I certainly do with the table types I maintain in Meshes.jl/GeoStats.jl. There we already have a different parent type that is not an AbstractVector. I think AbstractVector is extremely overused in Julia.

So, if we could have a trait function instead of a parent type, that could help maybe.

> [@Henrique\_Becker](#):
>
> Tables are basically a collection of `Vector` s, I have yet not understood why people think they are not suited for random access.

That is the column view of tables @Henrique_Becker . Some tables have “infinite” rows and so we can’t randomly access index 100004320 efficiently. We need to iterate until we get there, the data produces the rows on the fly.

[Next page](https://discourse.julialang.org/t/random-access-to-rows-of-a-table/77386.md?page=2)
