# Web scraping of GCN NASA circulars TEXT

**URL:** <https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809>\
**Category:** General Usage\
**Tags:** http\
**Created:** [June 25, 2023, 12:29pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809 "2023-06-25T12:29:15Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [June 25, 2023, 12:29pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/1 "2023-06-25T12:29:15Z")

</div>

I want to scrap [GCN - Circulars](https://gcn.nasa.gov/circulars) for optical events.

For example i want to get following information this web [page](https://gcn.nasa.gov/circulars/34030)😊

![image](https://global.discourse-cdn.com/julialang/original/3X/e/b/eb6fe74f5546244de91bc03d7b34f4d8070d04d1.png)

and similarly for other web pages containing optical events and store in ccv or excel format.

---

<div class="post-metadata">

**Author:** ![cormullion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cormullion/32/49131_2.png) [@cormullion](https://discourse.julialang.org/u/cormullion)\
**Post date:** [June 25, 2023, 12:39pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/2 "2023-06-25T12:39:57Z")

</div>

Fortunately many NASA pages provide JSON versions as well, so you can just use a JSON package such as [GitHub - quinnj/JSON3.jl](https://github.com/quinnj/JSON3.jl).

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [June 25, 2023, 1:05pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/3 "2023-06-25T13:05:45Z")

</div>

How can i extract information and compile in csv table for many web pages automatically ? 🧐

---

<div class="post-metadata">

**Author:** ![cormullion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cormullion/32/49131_2.png) [@cormullion](https://discourse.julialang.org/u/cormullion)\
**Post date:** [June 25, 2023, 1:39pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/4 "2023-06-25T13:39:43Z")

</div>

> [@raman\_kumar](#):
>
> automatically

By “automatically” do you mean “without writing any code”? You would probably still have to write some code to guide the process.

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [June 25, 2023, 2:07pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/5 "2023-06-25T14:07:35Z")

</div>

My code for web page looks like :

```julia
using HTTP , Gumbo , Cascadia, AbstractTrees
url="https://gcn.nasa.gov/circulars/34030"
r=HTTP.get(url)
h=parsehtml(String(r.body))
body=h.root[2]
eachmatch(Selector("p"), body)
Div=eachmatch(Selector(".usa-accordion __button.usa-banner__ button"), body)
Div[1]

```

see below Pluto 😊

 ![image](https://global.discourse-cdn.com/julialang/original/3X/f/5/f5b0bec204ace8b6e58e5e9f0df0c5280748c89c.png)

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [June 26, 2023, 5:18am UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/6 "2023-06-26T05:18:37Z")

</div>

I want to scrap that text information for every web page show below : 🧐

![image](https://global.discourse-cdn.com/julialang/original/3X/e/b/eb6fe74f5546244de91bc03d7b34f4d8070d04d1.png)

---

<div class="post-metadata">

**Author:** ![algunion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/algunion/32/51630_2.png) [@algunion](https://discourse.julialang.org/u/algunion)\
**Post date:** [June 26, 2023, 6:26am UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/7 "2023-06-26T06:26:38Z")

</div>

At this point, it looks like you have already managed to extract the relevant content/text from HTML.

Gumbo/Cascadia will not help to get the text into formatted data (since you have raw text, not some HTML table or other elements).

Gumbo.jl conveniently provides the `text` function that extracts the text from any HTML element. In your scenario, `text(Div[1])`.

However, this will output a string that is still not yet formatted per your needs - and Gumbo.jl has no helper functions for transforming a raw string into structured data.

A very simple parser for the format above can look like this:

```julia
using DataFrames

txt = """
JD (mid) | Telescope | Filter | Exposure (s) | Magnitude (AB) |
----------------------------------------------------------------------
2460115.3875 | OHP-T120 | R | 3900 | 20.70 +/- 0.12 | 
2460115.413706 | OHP-T193/MISTRAL | r' | 4560 | 20.84 +/- 0.04 | 
2460115.440972 | OHP-T120 | V | 4200 | 20.85 +/- 0.07 |"""

lines = split(txt, "\n")
parseline(line) = strip.(split(line, "|"))[1:end-1]
header = parseline(lines[1])
rows = parseline.(lines[3:end])
d = Dict(k => [getindex(row, i) for row in rows] for (i, k) in enumerate(header))
DataFrame(d)

```

And will produce something like this:

 ![disc](https://global.discourse-cdn.com/julialang/original/3X/a/6/a627888170f006be885efc6012fdb19b847b3abd.png)

Now, if the pages contain the same text somewhere in the content, you can create some matching pattern to get the start and the end of the desired text and use something similar to the code above to extract it as a data frame (and finally as CSV).

However, please note that this is beyond Gumbo.jl capabilities.

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [June 26, 2023, 6:51am UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/8 "2023-06-26T06:51:03Z")

</div>

What are the other possible ways to extract data ? I mean are there some other library which can help in getting out desired data? 🤔

---

<div class="post-metadata">

**Author:** ![Jose\_Diaz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jose_diaz/32/43750_2.png) [@Jose\_Diaz](https://discourse.julialang.org/u/Jose_Diaz)\
**Post date:** [June 26, 2023, 10:46am UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/9 "2023-06-26T10:46:20Z")

</div>

I think with the matching pattern thing you’re pretty much thone.

Btw, for the web scraping section, you can also use Harbest

```julia
using Harbest

html = read_html("https://gcn.nasa.gov/circulars/34030")

data = html_elements(html,["main","div"])[12]
html_text3(data)

```

---

<div class="post-metadata">

**Author:** ![algunion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/algunion/32/51630_2.png) [@algunion](https://discourse.julialang.org/u/algunion)\
**Post date:** [June 26, 2023, 11:03am UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/10 "2023-06-26T11:03:56Z")

</div>

For simple examples like the one you provided, I think the best way is to stick with pure Julia (see the example I provided - you can improve on that, I just put something together for you to get you started). For example, you can see that I didn’t convert things from strings to numerical values - the goal was to give you a minimal example regarding data extraction.

I don’t have working experience with specific parser libraries in Julia, but I stumbled upon [andrewcooke/ParserCombinator.jl](https://github.com/andrewcooke/ParserCombinator.jl) a few times.

I think it is easier and faster just to put together your own parser for scenarios like the above. However, if somebody else is aware of a better way to do this, I am curious myself if there are some good parsing libraries in Julia (besides the language-related parsers).

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [June 26, 2023, 6:21pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/11 "2023-06-26T18:21:17Z")

</div>

Ok, so i will learn ParserCombinator.jl . Thank You

---

<div class="post-metadata">

**Author:** ![algunion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/algunion/32/51630_2.png) [@algunion](https://discourse.julialang.org/u/algunion)\
**Post date:** [June 26, 2023, 6:23pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/12 "2023-06-26T18:23:21Z")

</div>

Don’t get me wrong - I mentioned that package because of the familiarity alone - I am not suggesting that is the right solution for your specific problem: as I said, I consider writing some Julia specific for your use-case is a better approach (a parser combinator might be overkill).

However, learning will not hurt - I am just ensuring I am not pointing you in the wrong direction.

Also - please check the code I shared - that is already working for your specific use case (although - it might not be complete - you might want to add some conversations and make sure you isolate the snippet from the larger text content).

Have fun.

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [June 27, 2023, 10:05am UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/13 "2023-06-27T10:05:54Z")

</div>

I have been able to do following things using ParserCombinator.jl 😊

 ![image](https://global.discourse-cdn.com/julialang/original/3X/6/9/694dc2694b20359362f6966db604fec4b8a39b52.png)

It shows Parser Exception for line 1 of text. Please see last line of picture above . 🤔

---

<div class="post-metadata">

**Author:** ![algunion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/algunion/32/51630_2.png) [@algunion](https://discourse.julialang.org/u/algunion)\
**Post date:** [June 28, 2023, 5:22pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/14 "2023-06-28T17:22:10Z")

</div>

I suggest you go deeper into the documentation of the package.

For example, `parse_one` returns _a single result or throws a `ParserException`_. This is exactly what happened when the parser failed to detect the intended pattern (“Date”) - and the behavior is the correct/intended one.

Maybe try to review the tutorials and examples provided by the package documentation.

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [June 28, 2023, 5:38pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/15 "2023-06-28T17:38:28Z")

</div>

Why parse\_one terminate after only first index and don’t move forward in the _for_ loop ? I mean it should give output for line 5 when marching in _for_ loop.

---

<div class="post-metadata">

**Author:** ![algunion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/algunion/32/51630_2.png) [@algunion](https://discourse.julialang.org/u/algunion)\
**Post date:** [June 28, 2023, 5:40pm UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/16 "2023-06-28T17:40:58Z")

</div>

1. It throws an exception: if you want it to continue iterating, you must manage the exception.
2. See 1.

---

<div class="post-metadata">

**Author:** ![raman\_kumar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raman_kumar/32/26782_2.png) [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Post date:** [June 30, 2023, 10:02am UTC](https://discourse.julialang.org/t/web-scraping-of-gcn-nasa-circulars-text/100809/17 "2023-06-30T10:02:33Z")

</div>

I am able to handle exception. 😃

![image](https://global.discourse-cdn.com/julialang/original/3X/3/5/35dc6398f00f29caf2623640ea31091b362c7582.png)
