# SASLib v1.1.0 release

**URL:** <https://discourse.julialang.org/t/saslib-v1-1-0-release/32959>\
**Category:** Package Announcements\
**Created:** [January 4, 2020, 7:19am UTC](https://discourse.julialang.org/t/saslib-v1-1-0-release/32959 "2020-01-04T07:19:11Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [January 4, 2020, 7:19am UTC](https://discourse.julialang.org/t/saslib-v1-1-0-release/32959/1 "2020-01-04T07:19:11Z")

</div>

Hi all,

I have just tagged a new major version of [SASLib](https://github.com/tk3369/SASLib.jl). I took the chance to migrate version number to 1.x, as the library/API has been stable and unchanged for quite a while.

The main updates are:

- Fully adopt Julia 1.x by removing old v0.6-specific code and adding Project.toml.
- Implement [Tables.jl](https://github.com/JuliaData/Tables.jl) interface for better integration with the rest of data ecosystem
- An updated performance test results against Python/Pandas and ReadStat C-library

A sneak peak about the latest [performance comparison results](https://github.com/tk3369/SASLib.jl/tree/master/test/perf_results_1.0.0):

| Test | Result |
| --- | --- |
| py\_jl\_homimp\_50.md | 30x faster than Python/Pandas |
| py\_jl\_numeric\_1000000\_2\_100.md | 10x faster than Python/Pandas |
| py\_jl\_productsales\_100.md | 50x faster than Python/Pandas |
| py\_jl\_test1\_100.md | 120x faster than Python/Pandas |
| py\_jl\_topical\_30.md | 30x faster than Python/Pandas |

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [January 4, 2020, 10:33am UTC](https://discourse.julialang.org/t/saslib-v1-1-0-release/32959/2 "2020-01-04T10:33:30Z")

</div>

Being able to read files in chunks would be awesome!

---

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [January 4, 2020, 10:45pm UTC](https://discourse.julialang.org/t/saslib-v1-1-0-release/32959/3 "2020-01-04T22:45:21Z")

</div>

It already supports reading file in chunks (see [Incremental Reading](https://github.com/tk3369/SASLib.jl#incremental-reading) section of README).

Are you referring to the multi-threaded aspect per [issue #5](https://github.com/tk3369/SASLib.jl/issues/5) or the random access idea per [issue #35](https://github.com/tk3369/SASLib.jl/issues/35)?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [January 4, 2020, 10:56pm UTC](https://discourse.julialang.org/t/saslib-v1-1-0-release/32959/4 "2020-01-04T22:56:50Z")

</div>

If I call correctly, the issue is that when reading large files in chunks, it reads the whole file somehow and then chunks it. It’s really slow.

See [https://github.com/tk3369/SASLib.jl/issues/50](https://github.com/tk3369/SASLib.jl/issues/50)

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [January 6, 2020, 6:35pm UTC](https://discourse.julialang.org/t/saslib-v1-1-0-release/32959/5 "2020-01-06T18:35:49Z")

</div>

Awesome!

The performance results made me take another look at [ReadStat.jl](https://github.com/queryverse/ReadStat.jl). Turns out there was something like a type instability festival happening in the most inner loop 🙂 I just released a new version that fixes that and should give much better performance across the board.

I reran the benchmark in SASLib with that new version, and I think the only test where SASLib was still faster than ReadStat was the `data_misc/numeric_1000000_2.sas7bdat` test (but I did this hastily, would be great if you could rerun the results in your repo!). But maybe that is a bit of an apples and oranges comparison, because SASLib doesn’t handle missing values, it just hands them down as `NaN`, right?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [January 6, 2020, 11:54pm UTC](https://discourse.julialang.org/t/saslib-v1-1-0-release/32959/6 "2020-01-06T23:54:07Z")

</div>

I believe ReadStat doesn’t support binary compressed SAS? (RDC compression)

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [January 6, 2020, 11:56pm UTC](https://discourse.julialang.org/t/saslib-v1-1-0-release/32959/7 "2020-01-06T23:56:20Z")

</div>

I think that is so. ReadStat.jl just exposes whatever ReadStat the C library does.

---

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [January 7, 2020, 7:22am UTC](https://discourse.julialang.org/t/saslib-v1-1-0-release/32959/8 "2020-01-07T07:22:53Z")

</div>

I can definitely re-run the performance test again!

Yes, SASLib doesn’t deal with missing data yet. An enhancement request was logged here [https://github.com/tk3369/SASLib.jl/issues/43](https://github.com/tk3369/SASLib.jl/issues/43)

The file `data_misc/numeric_1000000_2.sas7bdat` file does not contain any missing values so it isn’t necessary a bad test.

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [January 7, 2020, 5:10pm UTC](https://discourse.julialang.org/t/saslib-v1-1-0-release/32959/9 "2020-01-07T17:10:58Z")

</div>

I think the problem is that ReadStat’s logic for missing values detection runs always, even if the file doesn’t have missing values in it. As far as I could tell one can’t tell beforehand in the file format whether a column has missing values in it, right? Essentially one has to add a check for every single value that is read that a) checks for NaN, and b) checks for these declared missing ranges (or was that for a different file format that is also supported by ReadStat?). But I think those checks then also run for files that actually don’t have any missing values in them, at least that is what I believe is happening in ReadStat.
