# Problem of unnecessarily iterating many times in for loop

**URL:** <https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319>\
**Category:** Performance\
**Tags:** question, strings, csv, http\
**Created:** [July 7, 2023, 5:05pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319 "2023-07-07T17:05:24Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 7, 2023, 5:05pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/1 "2023-07-07T17:05:25Z")

</div>

For the following code it gives output many times(about 40 times) for only three _true_ matches . Unnecessary iteration will decrease performance as far as i know it should give only three output .

```julia
julia> using CSV, DataFrames,HTTP

julia> for x in 34000:34135
               url="https://gcn.nasa.gov/circulars/$x"
               txt=String((HTTP.get(url)))
           if occursin("report on behalf of the Swift/UVOT team",txt)==true || continue
                       url
                       for e in url
                               @show hb,he=findfirst(r"Filter"i,txt)
                       end
               else
               end
       end

(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491

```

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 7, 2023, 5:13pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/2 "2023-07-07T17:13:03Z")

</div>

What are you trying to do? What’s your question? This code is doing exactly what you told it to do, in other wordsyou’re making unnecessary loops.

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 7, 2023, 5:14pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/3 "2023-07-07T17:14:08Z")

</div>

where i am making unnecessary loops ? 🤔

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 7, 2023, 5:16pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/4 "2023-07-07T17:16:39Z")

</div>

> [@raman\_kumar](#):
>
> `",txt)==true || continue`

This is probably not doing what you think it’s doing.

You’re trying to skip the e in URL loop here right?

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 7, 2023, 5:20pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/5 "2023-07-07T17:20:40Z")

</div>

minor note: the `==true` is not necessary as `occursin` already outputs the boolean value needed for the `if ` statement.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 7, 2023, 5:23pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/6 "2023-07-07T17:23:16Z")

</div>

> [@raman\_kumar](#):
>
> ` for e in url`

Why are you iterating every character or your URL btw?

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 7, 2023, 5:26pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/7 "2023-07-07T17:26:22Z")

</div>

I am iterating in those url because i want to extract data of all webpages of those satisfying _occursin_ condition. and removing ` ==true` does not make any good changes.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 7, 2023, 5:28pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/8 "2023-07-07T17:28:44Z")

</div>

> [@raman\_kumar](#):
>
> `url="https://gcn.nasa.gov/circulars/$x"`

URL is just this fixed string

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 7, 2023, 5:32pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/9 "2023-07-07T17:32:50Z")

</div>

but `x` is changing every time.

```julia
julia> for x in 34000:34135
               url="https://gcn.nasa.gov/circulars/$x"
               txt=String((HTTP.get(url)))
           if occursin("report on behalf of the Swift/UVOT team",txt)
                   @show url
               else
               end
       end
url = "https://gcn.nasa.gov/circulars/34008"
url = "https://gcn.nasa.gov/circulars/34049"
url = "https://gcn.nasa.gov/circulars/34135"

```

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 7, 2023, 5:35pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/10 "2023-07-07T17:35:45Z")

</div>

As far as I can tell, you should only need.

```julia
for x in 34000:34135
    url="https://gcn.nasa.gov/circulars/$x"
    txt=String((HTTP.get(url)))
    if occursin("report on behalf of the Swift/UVOT team",txt)
        @show hb,he=findfirst(r"Filter"i,txt)
    end
end

```

If you do

```julia
for e in url
   println(e)
end

```

then you will see what @jling is saying. That loop is just counting over the individual characters in the URL string, it is _not_ iterating over the individual URLs. That is what your first loop is doing.,

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 7, 2023, 5:55pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/11 "2023-07-07T17:55:39Z")

</div>

Try code given below in Pluto notebook. I see it giving correct answer again and again repeatedly and restarting again .

```julia
using CSV, DataFrames,HTTP

```

```julia
begin
	for x in 34000:34135
        url="https://gcn.nasa.gov/circulars/$x"
        txt=String((HTTP.get(url)))
        if occursin("report on behalf of the Swift/UVOT team",txt)
            @show hb,he=findfirst(r"Filter"i,txt)
        end
    end
end

```

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 7, 2023, 6:26pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/12 "2023-07-07T18:26:11Z")

</div>

Pluto issues?

```julia
julia> using HTTP

julia> for x in 34000:34135
               url="https://gcn.nasa.gov/circulars/$x"
               txt=String((HTTP.get(url)))
               if occursin("report on behalf of the Swift/UVOT team",txt)
                   @show hb,he=findfirst(r"Filter"i,txt)
               end
           end
(hb, he) = findfirst(r"Filter"i, txt) = 10545:10550
(hb, he) = findfirst(r"Filter"i, txt) = 10610:10615
(hb, he) = findfirst(r"Filter"i, txt) = 10486:10491

```

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 26, 2023, 12:13pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/13 "2023-07-26T12:13:13Z")

</div>

```julia
using HTTP,DataFrames,CSV

```

```julia
begin
	for x in 33211:33212
            url="https://gcn.nasa.gov/circulars/$x"
            txt=String((HTTP.get(url)))
        if occursin("report on behalf of the Swift/UVOT team",txt)
             hb,he=findfirst(r"^Filter"im,txt)
			 lr,_=findnext("\n\nThe",txt,he)
			 cltxt=replace(txt[hb:lr], " +/- "=>"+/-", r" +(\w)"=>s"\t\1",r" +(&gt;)"=>s"\t>")
			 @show df=CSV.read(IOBuffer(cltxt), DataFrame, delim='\t')
        end
    end
end

```

![image](https://global.discourse-cdn.com/julialang/original/3X/f/6/f64e7ff252305ff7436bf33f35f85c456e060554.png)

The above code gives correct table but inserting it in below code it produces _missing_ column 🙃

```julia
using HTTP,DataFrames,CSV

```

```julia
function doanalysis()
       dfg=nothing
       for x in 33211:33212
           print("\r peeking at GCN $x ")
           try
               url = "https://gcn.nasa.gov/circulars/$x/raw"
               resp = HTTP.get(url) 
               status=resp.status
               print(" ",status," "); 
               if status == 404 ; println("status=",status); continue; end          
               txt = String(resp.body)

			   if occursin(r"GRB ?\d{6}([A-G]|(\.\d{2}))?",txt)
				    m=match(r"GRB ?\d{6}([A-G]|(\.\d{2}))?",txt)
				    print(m.match)
			   end
			  if occursin("report on behalf of the Swift/UVOT team",txt)
             hb,he=findfirst(r"^Filter"im,txt)
			 lr,_=findnext("\n\nThe",txt,he)
			 cltxt=replace(txt[hb:lr], " +/- "=>"+/-", r" +(\w)"=>s"\t\1",r" +(&gt;)"=>s"\t>")
			 @show df=CSV.read(IOBuffer(cltxt), DataFrame, delim='\t')
			
				#df.GCN=[x for i in 1:nrow(df)]
				#df.GRB=[m.match for i in 1:nrow(df)]
				#rename!(df, :Mag => :Magnitude)
			#@show DataFrame(df)
           end
		   catch e
			   print("Error")
		   end
	   end
end   

```

```julia
doanalysis()

```

![image](https://global.discourse-cdn.com/julialang/original/3X/e/b/ebc7dbb5fba830267ba5a28fc49c8c111a66272f.jpeg)

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 26, 2023, 1:17pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/14 "2023-07-26T13:17:57Z")

</div>

The URL and `txt =` lines are different between the two codes. What happens is you use the `url = ` and `txt = ` lines from `doanalysis` in the original (upper) version of the code?

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 26, 2023, 2:38pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/15 "2023-07-26T14:38:55Z")

</div>

What changes should i do in code if i want to keep _/raw_ in URL ?

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 26, 2023, 3:00pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/16 "2023-07-26T15:00:49Z")

</div>

What happened when you used the `raw` url in the code that worked?

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [July 26, 2023, 3:46pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/17 "2023-07-26T15:46:12Z")

</div>

> [@raman\_kumar](#):
>
> `url = "https://gcn.nasa.gov/circulars/$x/raw"`

It contains missing column. See last output above.

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 27, 2023, 11:31am UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/18 "2023-07-27T11:31:58Z")

</div>

I’m not great with web scraping or regex, so I’m not sure how much I can help there. But maybe the first thing to do is grab the `txt` for both URLs (with and without `/raw` and see if the `cltxt` is the same. It seems like it might not be, then you might need to update `replace`.

---

<div class="post-metadata">

**Author:** ![Jeff\_Emanuel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff_emanuel/32/15440_2.png) [@Jeff\_Emanuel](https://discourse.julialang.org/u/Jeff_Emanuel)\
**Post date:** [August 10, 2023, 7:51pm UTC](https://discourse.julialang.org/t/problem-of-unnecessarily-iterating-many-times-in-for-loop/101319/19 "2023-08-10T19:51:06Z")

</div>

It looks like your missing Mag data are in the Exp(s) column with some whitespace between the desired Exp(s) and the Mag… You probably accidentally removed or escaped the separating ‘\t’.
