# Box coordinates from image segmentation

**URL:** <https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987>\
**Category:** General Usage\
**Tags:** images, image-processing, juliaimages, segmentation\
**Created:** [February 9, 2024, 3:26pm UTC](https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987 "2024-02-09T15:26:00Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![sardinecan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sardinecan/32/29141_2.png) [@sardinecan](https://discourse.julialang.org/u/sardinecan)\
**Post date:** [February 9, 2024, 3:26pm UTC](https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987/1 "2024-02-09T15:26:00Z")

</div>

Hi Everyone,  
I’m new to the field of image segmentation, and I need your help to find a way of determining the coordinates of a box from image segmentation.

I have a batch of images like thie one:

 ![Capture d’écran 2024-02-09 à 16.09.18](https://global.discourse-cdn.com/julialang/original/3X/8/5/85984c1ce67ceba22b5180439b891643fcc446cd.jpeg)

I try to find a way of determining the zones/boxes corresponding to each page to “crop” these images directly with the IIIF API.

The ImageSegmentation package seems to be efficient for this task:

```julia
using Image
using ImageSegmentation
using HTTP
using FileIO
using Random

imgUrl = "https://api.nakala.fr/data/10.34847/nkl.027b840e/5c8e77a046216ab6aed848b2f781deb9495fea76"
file = download(imgUrl) |> load

function get_random_color(seed)
    Random.seed!(seed)
    rand(RGB{N0f8})
end

seg = fast_scanning(file, 0.2)
map(i->get_random_color(i), labels_map(seg))

```

ouput:

 ![Capture d’écran 2024-02-09 à 16.09.31](https://global.discourse-cdn.com/julialang/original/3X/d/3/d35543bd04c801de22e3ea0e1fe9c3ca66a1ebe3.jpeg)

What I need to do now is find a way to determine the coordinates of the top left and bottom right points for both blue and green areas. Do you have any idea to achieve this task, or with another method?

Thanks for your help!  
Best,  
Josselin

---

<div class="post-metadata">

**Author:** ![sardinecan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sardinecan/32/29141_2.png) [@sardinecan](https://discourse.julialang.org/u/sardinecan)\
**Post date:** [February 10, 2024, 8:44pm UTC](https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987/2 "2024-02-10T20:44:32Z")

</div>

I found a solution.  
First, we take the segmentation matrix

```julia
segMap = labels_map(seg)

```

Then, from this matrix, we can retrieve the coordinates of all the pixels in a given segment. Then we can store the X and Y coordinates in two separate variables to finally determine a box around a segment (the black crosses on the image below):

```julia
coordinates = findall(x -> x == 1, segMap)
x = Vector()
y = Vector()

for c in coordinates
    push!(x, c[2])
    push!(y, c[1])
end

xStart = x[1]
xEnd = last(x)

sort!(y)
yStart = y[1]
yEnd = last(y)

```

 ![Capture d’écran 2024-02-10 à 21.11.13](https://global.discourse-cdn.com/julialang/original/3X/5/8/58c78a1cde52f0f7843c1b3aa3a6fbf02471adae.png)

Hope this makes sens!  
Best

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 11, 2024, 4:27am UTC](https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987/3 "2024-02-11T04:27:48Z")

</div>

You don’t need to sort an array to find the minimum and maximum. You should be able to just do:

```julia
segMap = labels_map(seg)
coordinates = findall(==(1), segMap)
xmin, xmax = extrema(c -> c[1], coordinates)
ymin, ymax = extrema(c -> c[2], coordinates)

```

(With a bit more cleverness, you could do it in a single pass over `segMap`, without ever constructing a `coordinates` array explicitly, but it’s late and I’m lazy.)

PS. Note that calling `x = Vector()` and then repeatedly calling `push!` is fairly inefficient. For one thing, `Vector()` returns an untyped array `Any[]`, which is [inefficient to work with](https://docs.julialang.org/en/v1/manual/performance-tips/#man-performance-abstract-container). For another thing, repeatedly calling `push!`, while it is still amortized linear time, is less efficient than allocating an array of the correct length to begin with. A simpler, more efficient construction would be `x = getindex.(coordinates, 1)`, or alternatively `x = map(c -> c[1], coordinates)`. But it is even better to avoid allocating the `x` array entirely, as in my code above.

---

<div class="post-metadata">

**Author:** ![sardinecan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sardinecan/32/29141_2.png) [@sardinecan](https://discourse.julialang.org/u/sardinecan)\
**Post date:** [February 11, 2024, 12:19pm UTC](https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987/4 "2024-02-11T12:19:20Z")

</div>

Thanks for your advice and improvements @stevengj. I still have a lot to learn!

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [February 11, 2024, 8:07pm UTC](https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987/5 "2024-02-11T20:07:17Z")

</div>

If the heuristic is deemed useful for the case at hand, the following form should be faster in carving out the “good” part

```julia
tr=10^3
 
cmin=findfirst(c->sum(c.!=@view segMap[:,1])>tr,eachcol(segMap))
cmax=findfirst(c->sum(c.!=@view segMap[:,end])>tr,reverse(eachcol(segMap)))
rmin=findfirst(r->sum(r.!=@view segMap[1,:])>tr,eachrow(segMap))
rmax=findfirst(r->sum(r.!=@view segMap[end,:])>tr,reverse(eachrow(segMap)))

file[rmin:end-rmax,cmin:end-cmax]

```

---

<div class="post-metadata">

**Author:** ![mpeters2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mpeters2/32/202247_2.png) [@mpeters2](https://discourse.julialang.org/u/mpeters2)\
**Post date:** [February 12, 2024, 3:10am UTC](https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987/6 "2024-02-12T03:10:23Z")

</div>

I might have missed this, so feel free to ignore me, but does that algorithm return two points (e.g. left-top and bottom-right) or 4 points? The reason I ask is that the images might not be exactly 90-degrees aligned, and it might make more sense to return 4 points so that you have parallelogram.

---

<div class="post-metadata">

**Author:** ![sardinecan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sardinecan/32/29141_2.png) [@sardinecan](https://discourse.julialang.org/u/sardinecan)\
**Post date:** [February 12, 2024, 10:57am UTC](https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987/7 "2024-02-12T10:57:48Z")

</div>

Hi @mpeters2,

You’re right. It returns only two points, which are the intersections of the extrema (xmin/ymin and xmax/ymax). So it does not “fit” exactly the segments… but [region with IIIF API](https://iiif.io/api/image/3.0/#41-region) only takes one point, the upper left, with two other parameters: width and height.

> The region of the full image to be returned is specified in terms of absolute pixel values. The value of _`x`_ represents the number of pixels from the 0 position on the horizontal axis. The value of _`y`_ represents the number of pixels from the 0 position on the vertical axis. Thus the _`x,y`_ position 0,0 is the upper left-most pixel of the image. _`w`_ represents the width of the region and _`h`_ represents the height of the region in pixels.

You are also right about the alignment. If we could get the exact coordinates of the 4 points, it would be possible to rotate the image and increase precision (although my 18th c. papers are not square 😄). But in any case, I have no idea how to obtain these 4 exact coordinates, segmentation may not be accurate enough to identify them? Any suggestion is welcome!

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 12, 2024, 1:33pm UTC](https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987/8 "2024-02-12T13:33:48Z")

</div>

> [@sardinecan](#):
>
> But in any case, I have no idea how to obtain these 4 exact coordinates, segmentation may not be accurate enough to identify them?

You could use a [corner identification algorithm](https://en.wikipedia.org/wiki/Corner_detection). (e.g. via ImageCorners.jl)

---

<div class="post-metadata">

**Author:** ![sardinecan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sardinecan/32/29141_2.png) [@sardinecan](https://discourse.julialang.org/u/sardinecan)\
**Post date:** [February 15, 2024, 9:08am UTC](https://discourse.julialang.org/t/box-coordinates-from-image-segmentation/109987/9 "2024-02-15T09:08:04Z")

</div>

Thanks @stevengj, I’ll take a look at ImageCorners. For my use case, it might be more effective than ImageSegmentation!?
