# How do I make the julia code efficient?

**URL:** <https://discourse.julialang.org/t/how-do-i-make-the-julia-code-efficient/87560>\
**Category:** General Usage\
**Tags:** question\
**Created:** [September 21, 2022, 9:39am UTC](https://discourse.julialang.org/t/how-do-i-make-the-julia-code-efficient/87560 "2022-09-21T09:39:49Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![zhangchunyong](https://avatars.discourse-cdn.com/v4/letter/z/d9b06d/32.png) [@zhangchunyong](https://discourse.julialang.org/u/zhangchunyong)\
**Post date:** [September 21, 2022, 9:39am UTC](https://discourse.julialang.org/t/how-do-i-make-the-julia-code-efficient/87560/1 "2022-09-21T09:39:49Z")

</div>

```julia
function conver(ind, str)
    ls = length(str)
    ls < ind + 2 && return nothing
    s = SubString(str, ind+1, ind+2)
    if 'N' in s
        return missing
    else
        return s=="CG"
    end
end
function classify_reads(index::Int64
                       ,match_read_cpg::GroupedDataFrame
                       ,starts_cpgs::Int64
                       ,starts_reads::Int64
                       ,seqs_reads::LongDNASeq
                       ,overlapcopy::DataFrame)
    covered_cpgs = match_read_cpg[index][:,2]
    if length(covered_cpgs)<4
        return missing
    end
    start_cpgs=starts_cpgs[covered_cpgs]
    start_of_read=starts_reads[index]
    start_cpgs = start_cpgs .- start_of_read
    sequence=seqs_reads[index]
    representation=Vector{Bool}()
    s=String(sequence)
    for i in start_cpgs
        c=conver(i,s)
        if isequal(c,missing)|isequal(c,nothing)
        	deleteat!(overlapcopy,findall(overlapcopy.queryHits.==index .&& overlapcopy.subjectHits.==covered_cpgs[findfirst(x->x==i,start_cpgs)]))
        else
            push!(representation,c)
        end
    end
    if length(representation)<4
        return missing
    end
    concordant = (all(representation) || all(.!representation))
    return !concordant
end        
        
function calculatestate(classified_reads,match_read_cpg,starts_cpgs,starts_reads,seqs_reads,overlapcopy)
    p=Vector{Union{Missing,Bool}}()
    for i in classified_reads
        #println(i)
        a=classify_reads(i,match_read_cpg,starts_cpgs,starts_reads,seqs_reads,overlapcopy)
        push!(p,a)
    end
    p
end

```

match\_read\_cpg (the **second** parameter to the second and third function) is a GroupedDataframe that I need iterate one dataframe at a time to handle. When I execute the calculatestate function, I use the first for and there is a function called _a=classify\_reads (i,match\_read\_cpg,starts\_cpgs,starts\_reads,seqs\_reads, overlapcopy)_ with a for inside. I actually want to broadcast conver(the first function) in the second function classify\_reads, but I have to make a judgment after execution to delete the rows of overlapcopy. So you have two nested fors. This step will slow down. Would you please show me how to solve the problem to make it faster? Is there anything else I can change to improve the performances? Thanks for helping me!

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [September 21, 2022, 2:11pm UTC](https://discourse.julialang.org/t/how-do-i-make-the-julia-code-efficient/87560/2 "2022-09-21T14:11:53Z")

</div>

The low hanging fruits are (I have no idea how they affect performance, because the example is not runnable):

1. Add a view here

> [@zhangchunyong](#):
>
> ` covered_cpgs = @view(match_read_cpg[index][:,2]`)

1. Replace this with a non-allocating generator of the indexes to run over:

> [@zhangchunyong](#):
>
> ` start_cpgs = (start_cpgs[i] - start_of_read[i] for i in eachindex(starg_cpgs))`

This is probably the worst line:

> [@zhangchunyong](#):
>
> ` deleteat!(overlapcopy,findall(overlapcopy.queryHits.==index .&& overlapcopy.subjectHits.==covered_cpgs[findfirst(x->x==i,start_cpgs)]))`

Why not (I even think it is simpler, not sure if correct):

```julia
for i in eachindex(overlapcopy.queryHits)
    if overlapcopy.queryHits[i] == index &&
       overlapcopy.subjectHIts[i] == covered_cpgs[findfirst(x->x==i,start_cpgs)]
       deleteat!(overlapcopy, i)
    end
end

```

Finally, in this one I think you are allocating an array unnecessarily in `all(.!representation)`, because the `.!representation` will create a new array:

> [@zhangchunyong](#):
>
> ` concordant = (all(representation) || all(.!representation))`

use

```julia
concordant = (all(representation) || all(==(false), representation))

```

(or `all(!, representation)`, seems to work)

---

<div class="post-metadata">

**Author:** ![zhangchunyong](https://avatars.discourse-cdn.com/v4/letter/z/d9b06d/32.png) [@zhangchunyong](https://discourse.julialang.org/u/zhangchunyong)\
**Post date:** [September 21, 2022, 2:32pm UTC](https://discourse.julialang.org/t/how-do-i-make-the-julia-code-efficient/87560/4 "2022-09-21T14:32:22Z")

</div>

Thanks very much.I will test it.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [September 21, 2022, 2:45pm UTC](https://discourse.julialang.org/t/how-do-i-make-the-julia-code-efficient/87560/5 "2022-09-21T14:45:06Z")

</div>

A really basic thing is to remember that DataFrames and GroupedDataFrames are not type-stable objects. So you should write a function which acts on the columns from those objects and pass columns to those columns to the function.
