# Reading LIBSVM format

**URL:** <https://discourse.julialang.org/t/reading-libsvm-format/99466>\
**Category:** Data\
**Tags:** libsvm\
**Created:** [May 26, 2023, 9:30pm UTC](https://discourse.julialang.org/t/reading-libsvm-format/99466 "2023-05-26T21:30:25Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ymalitsky](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ymalitsky/32/26262_2.png) [@ymalitsky](https://discourse.julialang.org/u/ymalitsky)\
**Post date:** [May 26, 2023, 9:30pm UTC](https://discourse.julialang.org/t/reading-libsvm-format/99466/1 "2023-05-26T21:30:25Z")

</div>

Hi! How one can read files in the `libsvm` format in Julia? The `LIBSVM.jl` package seems to implement only algorithms. My current solution would be to read the file in python, save it in other format and then load it into Julia, but that’s a bit of an overkill for such a basic task.

---

<div class="post-metadata">

**Author:** ![jeremiedb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeremiedb/32/29150_2.png) [@jeremiedb](https://discourse.julialang.org/u/jeremiedb)\
**Post date:** [August 10, 2023, 7:43pm UTC](https://discourse.julialang.org/t/reading-libsvm-format/99466/2 "2023-08-10T19:43:42Z")

</div>

Following is a naive implementation I’ve made to read libsvm format.  
It’s not high performance but proved suitable for my needs: parsing the Yahoo Laerning to Rank Challenge Set1 train data takes about 2 mins (~ 475 000 observations, 700 columns).  
Note that It returns a dense matrix, not a sparse one:

```julia
function read_libsvm(raw::Vector{UInt8}; has_query=false)

    io = IOBuffer(raw)
    lines = readlines(io)

    nobs = length(lines)
    nfeats = 0 # number of features

    y = zeros(Float64, nobs)

    if has_query
        offset = 2 # offset for feature idx: y + query entries
        q = zeros(Int, nobs)
    else
        offset = 1 # offset for feature idx: y
    end

    vals = [Float64[] for _ in 1:nobs]
    feats = [Int[] for _ in 1:nobs]

    for i in eachindex(lines)
        line = lines[i]
        line_split = split(line, " ")

        y[i] = parse(Int, line_split[1])
        has_query ? q[i] = parse(Int, split(line_split[2], ":")[2]) : nothing

        n = length(line_split) - offset
        lfeats = zeros(Int, n)
        lvals = zeros(Float64, n)
        @inbounds for jdx in 1:n
            ls = split(line_split[jdx+offset], ":")
            lvals[jdx] = parse(Float64, ls[2])
            lfeats[jdx] = parse(Int, ls[1])
            lfeats[jdx] > nfeats ? nfeats = lfeats[jdx] : nothing
        end
        vals[i] = lvals
        feats[i] = lfeats
    end

    x = zeros(Float64, nobs, nfeats)
    @inbounds for i in 1:nobs
        @inbounds for jdx in 1:length(feats[i])
            j = feats[i][jdx]
            val = vals[i][jdx]
            x[i, j] = val
        end
    end

    if has_query
        return (x=x, y=y, q=q)
    else
        return (x=x, y=y)
    end
end

```

---

<div class="post-metadata">

**Author:** ![ymalitsky](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ymalitsky/32/26262_2.png) [@ymalitsky](https://discourse.julialang.org/u/ymalitsky)\
**Post date:** [August 11, 2023, 5:09pm UTC](https://discourse.julialang.org/t/reading-libsvm-format/99466/3 "2023-08-11T17:09:18Z")

</div>

Thanks! So does it mean that there is indeed no dedicated package for this?

---

<div class="post-metadata">

**Author:** ![jeremiedb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeremiedb/32/29150_2.png) [@jeremiedb](https://discourse.julialang.org/u/jeremiedb)\
**Post date:** [August 12, 2023, 6:59pm UTC](https://discourse.julialang.org/t/reading-libsvm-format/99466/4 "2023-08-12T18:59:30Z")

</div>

Correct, at least I’m not aware of any dedicated package.  
I think that the above function could be polished a little and added into LIBSVM, or even adapted into a dedicated lightweight package within [JuliaIO · GitHub](https://github.com/JuliaIO).
