# Some questions about eachoverlap in GenomicFeatures.jl

**URL:** <https://discourse.julialang.org/t/some-questions-about-eachoverlap-in-genomicfeatures-jl/85977>\
**Category:** Biology, Health, and Medicine\
**Created:** [August 19, 2022, 7:05am UTC](https://discourse.julialang.org/t/some-questions-about-eachoverlap-in-genomicfeatures-jl/85977 "2022-08-19T07:05:24Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![zhangchunyong](https://avatars.discourse-cdn.com/v4/letter/z/d9b06d/32.png) [@zhangchunyong](https://discourse.julialang.org/u/zhangchunyong)\
**Post date:** [August 19, 2022, 7:05am UTC](https://discourse.julialang.org/t/some-questions-about-eachoverlap-in-genomicfeatures-jl/85977/1 "2022-08-19T07:05:24Z")

</div>

During using the eachoverlap function in GenomicFeatures,I cannot know how to solve this problem.  
My problem is :if each interval in “_col_” contains any interval in “_hhh_” ,I will output interval "_col_ ".Here is the picture(below). For example _col_’s **10628-10683** contains **10631** in _hhh_,I will output the first line in _col_.Next **10643-10779** in _col_ contains _hhh_’s **10648**.If _col_’s interval does not contain any _hhh_’s interval,we will discard this _col_’s line

 ![image](https://global.discourse-cdn.com/julialang/original/3X/e/c/eca15711883798628f41746f4a0e098629532901.png)

 ![image](https://global.discourse-cdn.com/julialang/original/3X/e/c/ec47e0d47caafa7e70fe71b13ad16ee71d28a4c3.png)  
I typed in this code.and cannot solve it.

```julia
eachoverlap(col,hhh,isless)

```

 ![image](https://global.discourse-cdn.com/julialang/original/3X/e/8/e8ea01b183ebd9211e80f0f3e8e4cb50fa3a8d6f.png)  
And I also tried isoverlappiing,unfortunately it was slow.

```julia
k=IntervalCollection{String}()
for a in col ,b in hhh
    if isoverlapping(a,b)
        if (a in k)==false
            push!(k,a)
        end
    end
end

```

So would you please tell me how I can handle this issue?I will be grateful to you.

---

<div class="post-metadata">

**Author:** ![CiaranOMara](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ciaranomara/32/2589_2.png) [@CiaranOMara](https://discourse.julialang.org/u/CiaranOMara)\
**Post date:** [August 19, 2022, 12:50pm UTC](https://discourse.julialang.org/t/some-questions-about-eachoverlap-in-genomicfeatures-jl/85977/2 "2022-08-19T12:50:53Z")

</div>

I may not understand your question. However, below are some approaches for determining whether intervals in col have at least one intersection with intervals in hhh.

```julia
result = Vector{eltype(col)}()

for (c, h) in eachoverlap(col, hhh)
	push!(result, c)
end

result |> unique

```

```julia

result = Dict{eltype(col), Vector{eltype(hhh)}}()

for (c, h) in eachoverlap(col, hhh)
	intervals = get!(result, c, [])
	push!(intervals, h)
end

keys(result) |> collect |> sort

```

```julia
function hasintersection(interval, col)

	# Return early if chromosome is not in the interval collection.
	if !haskey(col.trees, seqname(interval))
		return false
	end

	# Setup intersection iterator.
	iter = GenomicFeatures.IntervalTrees.intersect(col.trees[seqname(interval)], (leftposition(interval), rightposition(interval)))

	# Attempt first iteration.
	if iterate(iter) === nothing
		return false
	end

	# Intersection exists.
	return true 

end

result = Vector{eltype(col)}()

for interval in col
	if hasintersection(interval, hhh)
		push!(result, interval)
	end
end

result

```

---

<div class="post-metadata">

**Author:** ![zhangchunyong](https://avatars.discourse-cdn.com/v4/letter/z/d9b06d/32.png) [@zhangchunyong](https://discourse.julialang.org/u/zhangchunyong)\
**Post date:** [August 21, 2022, 2:07am UTC](https://discourse.julialang.org/t/some-questions-about-eachoverlap-in-genomicfeatures-jl/85977/4 "2022-08-21T02:07:21Z")

</div>

Many thanks.And I made it in this way. But there is another question:How can I count the intervals in hhh in col’s interval(such as hhh’s first interval is in col’s first interval,hhh’s second interval is in col’s first interval,too)?  
For example ,the hhh’s 10631 10633 10636 … are in col’s first line like the picture below.Probably 10644 in hhh is also in col’s first line and second line, we will return like the second picture.(It means the fifth and sixth line in hhh are in the first line’s interval in col and also in the second line’s interval in col.) But I cannot solve the problem by myself.Would you please to help me?  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/2/f/2fd42e9b0877d481f9465916dc08a751cd2b463a.png)

 ![46246366642bc0beca0c4393c8e5970](https://global.discourse-cdn.com/julialang/original/3X/2/2/2204797e914f7de3652c214659391012d7d3e534.jpeg)

---

<div class="post-metadata">

**Author:** ![CiaranOMara](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ciaranomara/32/2589_2.png) [@CiaranOMara](https://discourse.julialang.org/u/CiaranOMara)\
**Post date:** [August 21, 2022, 12:05pm UTC](https://discourse.julialang.org/t/some-questions-about-eachoverlap-in-genomicfeatures-jl/85977/5 "2022-08-21T12:05:26Z")

</div>

That output does not appear to count the intervals; rather it lists the pairs of enumerated intervals that overlap.

```julia
using GenomicFeatures
using DataFrames

col = [
	Interval("chr1", 10628, 10683, '?', "abc")
	Interval("chr1", 10643, 10779, '?', "abc")
	Interval("chr1", 10645, 10748, '?', "abc")
	Interval("chr1", 10648, 10786, '?', "abc")
	Interval("chr1", 10676, 10767, '?', "abc")
	Interval("chr1", 10690, 10731, '?', "abc")
	Interval("chr1", 10707, 10793, '?', "abc")
	Interval("chr1", 10742, 10814, '?', "abc")
] |> IntervalCollection

hhh = [
	Interval("chr1", 10631, 10631)
	Interval("chr1", 10633, 10633)
	Interval("chr1", 10636, 10636)
	Interval("chr1", 10638, 10638)
	Interval("chr1", 10641, 10641)
	Interval("chr1", 10644, 10644)
	Interval("chr1", 10645, 10645)
	Interval("chr1", 10650, 10650)
] |> IntervalCollection

function number_interval(tp::Tuple)
	# Unpack Tuple.
	(i, interval) = tp

	# Setup numbered metadata.
	new_metadata = (
		i = i,
		original = metadata(interval)
	)

	# Create new interval with numbered metadata.
	return Interval(
		seqname(interval),
		leftposition(interval),
		rightposition(interval),
		strand(interval),
		new_metadata
	)
end

# Convert intervals.
col_numbered = col |> enumerate .|> number_interval |> IntervalCollection
hhh_numbered = hhh |> enumerate .|> number_interval |> IntervalCollection

# Reproduce example.
df = DataFrame()

for (c, h) in eachoverlap(col_numbered, hhh_numbered)

	result = (
		queryHits = GenomicFeatures.metadata(c).i, 
		subjectHits = GenomicFeatures.metadata(h).i
	)

	push!(df, result)
end

df

```

---

<div class="post-metadata">

**Author:** ![zhangchunyong](https://avatars.discourse-cdn.com/v4/letter/z/d9b06d/32.png) [@zhangchunyong](https://discourse.julialang.org/u/zhangchunyong)\
**Post date:** [August 21, 2022, 12:25pm UTC](https://discourse.julialang.org/t/some-questions-about-eachoverlap-in-genomicfeatures-jl/85977/6 "2022-08-21T12:25:06Z")

</div>

> [@CiaranOMara](#):
>
> ```julia
> function number_interval(tp::Tuple)
> # Unpack Tuple.
> (i, interval) = tp
> 
> # Setup numbered metadata.
> new_metadata = (
> i = i,
> original = metadata(interval)
> )
> 
> # Create new interval with numbered metadata.
> return Interval(
> seqname(interval),
> leftposition(interval),
> rightposition(interval),
> strand(interval),
> new_metadata
> )
> end
> 
> # Convert intervals.
> col_numbered = col |> enumerate .|> number_interval |> IntervalCollection
> hhh_numbered = hhh |> enumerate .|> number_interval |> IntervalCollection
> 
> # Reproduce example.
> df = DataFrame()
> 
> for (c, h) in eachoverlap(col_numbered, hhh_numbered)
> 
> result = (
> queryHits = GenomicFeatures.metadata(c).i, 
> subjectHits = GenomicFeatures.metadata(h).i
> )
> 
> push!(df, result)
> end
> 
> df
> 
> ```

Thanks for helping me again.It works.Best wishes.
