# Seeking advice on learning ML

**URL:** <https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016>\
**Category:** Machine Learning\
**Created:** [November 12, 2019, 7:36pm UTC](https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016 "2019-11-12T19:36:45Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [November 12, 2019, 7:36pm UTC](https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016/1 "2019-11-12T19:36:45Z")

</div>

I’ve done plenty of dabbling in ML and written some models that work (albeit, not very well) but I’m finding it difficult to 1) navigate the many ML-related topics and decide which ones to spend more time learning and which ones to gloss over 2) understand how to fine-tune a model once it’s working.

I’m hoping there is enough charity here in the Julia community to help guide me in the right direction 😊. The idea I have is to take some data from the Census Bureau’s American Community Survey (ACS) and explore different ways to predict a person’s income. There are millions of rows of data in the 2013 - 2017 ACS file that include tons of measurements about the individual survey respondents (age, educational attainment, occupation, the industry in which they work, race, gender, etc.).

It _seems_ like there should be enough information in this dataset to make fairly accurate predictions. I’m thinking I should start out simply trying to predict whether or not a person earns above or below some threshold amount by building a logistic regression model, a random forest, a neural network, and maybe some other model that’s good for this kind of problem.

Does this sound like a good start? Is there any reason I should start with one type of algorithm over another? Any tips/advice you can give? Does this sound like a decent ‘beginner’ problem to solve, or is it too complex?

Lastly, if anyone is interested in learning these topics, please reach out as I’d love to collaborate and learn together.

Thanks!!!

---

<div class="post-metadata">

**Author:** ![jbrea](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbrea/32/3879_2.png) [@jbrea](https://discourse.julialang.org/u/jbrea)\
**Post date:** [November 12, 2019, 8:22pm UTC](https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016/2 "2019-11-12T20:22:21Z")

</div>

I can highly recommend [The Elements of Statistical Learning](https://web.stanford.edu/~hastie/ElemStatLearn/) or [An Introduction to Statistical Learning](http://faculty.marshall.usc.edu/gareth-james/ISL/index.html). These books come with many rather small data sets that I find useful as examples to develop an intuition. I would not necessarily start with a huge dataset, because I don’t want to wait for my computer to finish a fit; I prefer when the results arrive within a fraction of a second when I am learning.

But the ACS data looks nice too. Or any competition or dataset on [kaggle](https://kaggle.com).

---

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [November 12, 2019, 8:32pm UTC](https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016/3 "2019-11-12T20:32:51Z")

</div>

Thank you!

---

<div class="post-metadata">

**Author:** ![fipelle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fipelle/32/4772_2.png) [@fipelle](https://discourse.julialang.org/u/fipelle)\
**Post date:** [November 12, 2019, 11:18pm UTC](https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016/4 "2019-11-12T23:18:13Z")

</div>

It would probably be best to start from some more standard datasets. The Ames, Boston and California Housing datasets are easier to approach and with less layers of complexity. For classification, MNIST is quite standard. You can find them on [https://www.kaggle.com](https://www.kaggle.com).

Depending on your interest and background you could focus on different topics. In addition to the books mentioned from @jbrea, you can also read [Goodfellow, Bengio, Courville (2016)](http://www.deeplearningbook.org) on deep learning, any of Vladimir Vapnik’s books on statistical learning and for a real classic in time series (not an easy book) you can take a look at [Priestley (1981)](https://www.elsevier.com/books/spectral-analysis-and-time-series-two-volume-set/priestley/978-0-08-057055-6).

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [November 13, 2019, 12:05am UTC](https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016/5 "2019-11-13T00:05:17Z")

</div>

If you’re inte the process of learning, I would reach for an ML framework that allows you to test many different things. I’ve experimented woth MLJ and believe it to have great potential. It’s similar in spirit to scikit learn, but have been rethought to solve some common pains of scikit. The kind of data you’re talking about lends itself well to prediction using tree and forest models, unless you specifically want to experiment with deep learning.

---

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [November 13, 2019, 1:48am UTC](https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016/6 "2019-11-13T01:48:20Z")

</div>

@baggepinnen @fipelle Thank you!

---

<div class="post-metadata">

**Author:** ![pontus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pontus/32/10438_2.png) [@pontus](https://discourse.julialang.org/u/pontus)\
**Post date:** [November 13, 2019, 2:49am UTC](https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016/7 "2019-11-13T02:49:14Z")

</div>

Thanks for raising the question! There is this recent text co-authored by @yoninazarathy and made available in draft form here: [https://people.smp.uq.edu.au/YoniNazarathy/julia-stats/StatisticsWithJulia.pdf](https://people.smp.uq.edu.au/YoniNazarathy/julia-stats/StatisticsWithJulia.pdf)

 ![2019-%20Klok%26Nazarathy-StatswithJulia-DS_ML_AI](https://global.discourse-cdn.com/julialang/original/3X/a/0/a0159c437517101b4523340a1bbf41ecc29dc960.png)  
(caught that in this post: [Multivariate Normal Distribution - #8 by yoninazarathy](https://discourse.julialang.org/t/multivariate-normal-distribution/29973/8))  
This seems relevant to your question in particular:

```julia
9 Machine Learning Basics - DRAFT . . . . . . . . . . . . . . . 311
9.1 Training, Validation and Testing . . . . . . . . . . . . . 311
9.2 Bias, Variance and Regularization . . . . . . . . . . . . . 312
9.3 Supervised Learning Methods . . . . . . . . . . . . . . . . 315
9.4 Unsupervised Learning Methods . . . . . . . . . . . . . . 324
9.5 Reinforcement Learning and MDP . . . . . . . . . . . . . . 333
9.6 A Taste of Generational Adversarial Networks . . . . . . . 340

```

It’s a beautifully typeset text, and you can copy the Julia examples right out of the pdf and generate the same output and plots to validate.  
There’s another good stats ref to pass along … thinking, or unthinking, let me dig it up…

---

<div class="post-metadata">

**Author:** ![pontus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pontus/32/10438_2.png) [@pontus](https://discourse.julialang.org/u/pontus)\
**Post date:** [November 13, 2019, 3:08am UTC](https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016/8 "2019-11-13T03:08:32Z")

</div>

Ah, found it! More on the statistics fundamentals. **REthinking** was the keyword… This may be more on the Bayesian side than what you’re looking for.  
I knew it was by @Tamas_Papp - was in a post he made as part of his “ANN: DynamicHMC 2.0” here: [ANN: DynamicHMC 2.0 - #15 by Tamas\_Papp](https://discourse.julialang.org/t/ann-dynamichmc-2-0/28314/15)  
It’s “StatisticalRethinkingJulia” associated with the book [Statistical Rethinking](https://xcelab.net/rm/statistical-rethinking/) by Richard McElreath.  
Dr. McElreath is Director of the Department of Human Behavior, Ecology, and Culture at the Max Planck Institute for Evolutionary Anthropology. He’s “an evolutionary ecologist who studies humans”, and as such, his lectures are engaging and fascinating!!  
There’s a plethora of vids on his Tube channel here:  
[https://www.youtube.com/channel/UCNJK6\_DZvcMqNSzQdEkzvzA/videos](https://www.youtube.com/channel/UCNJK6_DZvcMqNSzQdEkzvzA/videos)  
I’ve spent quite a few hours listening to those thanks to Mr. Papp! 🤨 Seriously, watch one, they’re great!

---

<div class="post-metadata">

**Author:** ![pontus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pontus/32/10438_2.png) [@pontus](https://discourse.julialang.org/u/pontus)\
**Post date:** [November 13, 2019, 3:34am UTC](https://discourse.julialang.org/t/seeking-advice-on-learning-ml/31016/9 "2019-11-13T03:34:35Z")

</div>

Sorry for the rapid-fire replies, but this just popped to mind, did you see the announcement the other day that Julia Academy is now free and available online?! There are topics specific to ML.

> [@JuliaAcademy is now free](https://discourse.julialang.org/t/juliaacademy-is-now-free/30457):
>
> We have exciting news: [Julia Academy](https://juliaacademy.com) is now entirely free! Perhaps even better, we’ve simultaneously open-sourced the raw [materials we used to generate the courses](https://github.com/JuliaComputing/JuliaAcademyMaterials/). Julia Academy bundles interactive notebooks with instructor videos so you can follow along at your own pace. We now have five courses spanning from a basic introduction to machine learning and advanced parallel computing. We hope that making these courses and their materials free to use and share will help you all continue learnin…

And you might consider spending your lunch hour next Tue with Dr. @mbauman:  
Webinar : Machine Learning with Julia  
Julia Computing Webinar  
Tuesday November 19 2019, 12:00 pm - 1:00 pm US Eastern Time  
Mode : Online  
Presenter : Dr. Matt Bauman, Senior Research Scientist, Julia Computing  
This may be a personal link as it came via email to my Julia Computing linked-address, but the form looks fine (if this should be pulled down, please alert me). Signed up but haven’t heard back yet:

> **[Webinar : Machine Learning with Julia](https://docs.google.com/forms/d/1ORfyYgMNvOBAq66ethQ75LchcizocnzIm8RTzPYd_yA/viewform?edit_requested=true)**
>
> Julia Computing Webinar
> Tuesday November 19 2019, 12:00 pm - 1:00 pm US Eastern Time
> Mode : Online 
> Presenter : Dr. Matt Bauman, Senior Research Scientist, Julia Computing
