# Extracting information from https://caps.fool.com/Ticker/MSFT.aspx

**URL:** <https://discourse.julialang.org/t/extracting-information-from-https-caps-fool-com-ticker-msft-aspx/56017>\
**Category:** New to Julia\
**Created:** [February 25, 2021, 2:24pm UTC](https://discourse.julialang.org/t/extracting-information-from-https-caps-fool-com-ticker-msft-aspx/56017 "2021-02-25T14:24:10Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nash/32/14482_2.png) [@Nash](https://discourse.julialang.org/u/Nash)\
**Post date:** [February 25, 2021, 2:24pm UTC](https://discourse.julialang.org/t/extracting-information-from-https-caps-fool-com-ticker-msft-aspx/56017/1 "2021-02-25T14:24:11Z")

</div>

I am trying to extract information from Caps run by Fool. However, using Gumbo and Cascadia, the html page appears to contain no useful information (and certainly not what I am after in a direct way). How can I get the information that is shown in my browser (for example, the Caps rating of 4)?

Here is my my code:

```
using Gumbo
using Cascadia

url = "https://caps.fool.com/Ticker/MSFT.aspx"

page = parsehtml(read(download(url), String))

```

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [February 25, 2021, 2:44pm UTC](https://discourse.julialang.org/t/extracting-information-from-https-caps-fool-com-ticker-msft-aspx/56017/2 "2021-02-25T14:44:21Z")

</div>

Does the example in the Cascadia readme not work for you?

```julia
julia> using Cascadia, Gumbo, HTTP

julia> r = HTTP.get("https://caps.fool.com/Ticker/MSFT.aspx");

julia> h = parsehtml(String(r.body));

julia> qs = eachmatch(Selector("#tickerRating"), h.root)
1-element Vector{HTMLNode}:
 HTMLElement{:div}:<div class="subtle marginT" id="tickerRating">
  CAPS Rating:
  <img alt="4 out of 5" class="capsStarRating" id="ctl00_ctl00_ctl00_ctl00_cphContent_cphContent_cphContent_cphCrossBar_TickerHeader_ctlLargeImageStars" src="https://g.foolcdn.com/art/ratings/stars/trans/4stars-trans-lg.png" title="4 Stars: Favorite"/>
</div>

```

Where it seemed from cursory inspection of the page source that the star rating is in the `tickerRating` div. You can then split the result like

```julia
julia> first(split(split(string(qs[1]), "title=\"")[2], ":"))
"4 Stars"

```

---

<div class="post-metadata">

**Author:** ![Nash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nash/32/14482_2.png) [@Nash](https://discourse.julialang.org/u/Nash)\
**Post date:** [February 25, 2021, 7:20pm UTC](https://discourse.julialang.org/t/extracting-information-from-https-caps-fool-com-ticker-msft-aspx/56017/3 "2021-02-25T19:20:27Z")

</div>

What if I wanted to access information further down on the website. For example, the information that of “All Players” 14,714 think MSFT will “Outperform” while 2,044 think it will “Underperform”. I am stuck “digging in” to the website beyond .root.

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [February 25, 2021, 7:30pm UTC](https://discourse.julialang.org/t/extracting-information-from-https-caps-fool-com-ticker-msft-aspx/56017/4 "2021-02-25T19:30:09Z")

</div>

```julia
julia> eachmatch(Selector(".jointSentimentGroup"), h.root)[1]
HTMLElement{:div}:<div class="jointSentimentGroup">
  <div class="perform out">
    <span>
      14,716
    </span>
    <span class="legend">
      Outperform
    </span>
  </div>
  <div class="perform under">
    <span>
      2,044
    </span>
    <span class="legend">
      Underperform
    </span>
  </div>
  <div class="sentimentBar underperformBar">
    <div class="outperformBar" style="width:87.80429594272076372315035800%">
    </div>
...

```

I don’t think there’s any magic here, just `Ctrl+F` for the information you’re after in the site’s source code and then check whether there’s a CSS element that you can select which contains what you need.

---

<div class="post-metadata">

**Author:** ![Nash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nash/32/14482_2.png) [@Nash](https://discourse.julialang.org/u/Nash)\
**Post date:** [February 25, 2021, 8:01pm UTC](https://discourse.julialang.org/t/extracting-information-from-https-caps-fool-com-ticker-msft-aspx/56017/5 "2021-02-25T20:01:33Z")

</div>

> [@nilshg](#):
>
> `eachmatch(Selector(".jointSentimentGroup"), h.root)[1]`

I See. Thank you!

---

<div class="post-metadata">

**Author:** ![Nash](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nash/32/14482_2.png) [@Nash](https://discourse.julialang.org/u/Nash)\
**Post date:** [February 25, 2021, 10:37pm UTC](https://discourse.julialang.org/t/extracting-information-from-https-caps-fool-com-ticker-msft-aspx/56017/6 "2021-02-25T22:37:16Z")

</div>

On a very similar note, I am trying to extract the “SmartScore” for MSFT (it’s the 8) on the following URL:

> **[Microsoft (MSFT) Stock Price, Quote, News & Analysis - TipRanks.com](https://www.tipranks.com/stocks/msft)**
>
> Microsoft (MSFT) has a Smart Score of 9 based on an analysis of 8 unique data sets, including Analyst Recommendations, Crowd Wisdom, and Hedge Fund Activity.​​

However, this one is tougher, I think. I find the source in the website code (see the diagram) but I can’t seem to extract it. Some help would again be appreciated.

 ![SmartScore](https://global.discourse-cdn.com/julialang/original/3X/0/1/0131c4b8c625494be41dcdaeb4327b69ed35d73e.jpeg)
