# US elections model implementation in Turing

**URL:** <https://discourse.julialang.org/t/us-elections-model-implementation-in-turing/49864>\
**Category:** Probabilistic Programming\
**Created:** [November 9, 2020, 7:45pm UTC](https://discourse.julialang.org/t/us-elections-model-implementation-in-turing/49864 "2020-11-09T19:45:53Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![entr0pidelic](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/entr0pidelic/32/15386_2.png) [@entr0pidelic](https://discourse.julialang.org/u/entr0pidelic)\
**Post date:** [November 9, 2020, 7:45pm UTC](https://discourse.julialang.org/t/us-elections-model-implementation-in-turing/49864/1 "2020-11-09T19:45:53Z")

</div>

Hello guys, hope you are doing well. I am trying to implement the US forecasting election model from the paper by Linzer [https://www.ime.usp.br/~abe/lista/pdfpWRrt4xFLt.pdf](https://www.ime.usp.br/~abe/lista/pdfpWRrt4xFLt.pdf), and it would be great if someone could give me a little help  
Here is my implementation:

```julia
@model function linzer_model(state_polls_dict, hist_state_forecast, poll_dates, states)
	n_states = length(hist_state_forecast)
	n_dates = length(poll_dates)
	σ_δ ~ Uniform(0, 10)
	σ_β ~ Uniform(0, 10)
	δ = Vector{Real}(undef, n_dates)
	β = Array{Real, 2}(undef, n_dates, n_states)
	# Random walks of the parameters β and δ
	for i in 1:(n_dates)
		if i == 1
			δ[i] = 0
			for j in 1:(n_states)
				β[i,j] = logit(hist_state_forecast[j])
			end
			continue
		end
		δ[i] ~ Normal(δ[i-1], σ_δ)
		for j in 1:(n_states)
			β[i,j] ~ Normal(β[i-1,j], σ_β)
		end
	end
	π = Array{Real,2}(undef, n_dates, n_states)
	for i in 1:n_dates
		π[i,:] = logistic.(β[i,:] .+ δ[i])
	end
	# rows -> days
	# columns -> states
	for i in 1:n_states
		for j in 1:n_dates
			if !(poll_dates[j] in keys(state_polls_dict[states[i]]))
				continue
			end
			n_polls = length(state_polls_dict[states[i]][poll_dates[j]][1,:])
			polls = state_polls_dict[states[i]][poll_dates[j]][1,:] # Hillary polls
			sample_sizes = state_polls_dict[states[i]][poll_dates[j]][3,:]
			for k in 1:n_polls
				polls[k] ~ Binomial(sample_sizes[k], π[j,i])
			end
		end
	end
end

```

The model basically goes like this:

- Pre-election polls are generated by a Binomial distribution. We have polls from every state and days before the election.
- the probabilities π for each poll of each states at each time, is given by π\_ij = logit−1 (β\_ij + δ\_j).
- β and δ are obtained by random walks, i.e., β\_ij∼ N(βi, j+1, (σ^2)\_β), δj ∼ N(δj+1, (σ^2)\_δ)
- σ\_β and σ\_δ have uniform priors.

The arguments of the model are:

- state\_polls\_dict: A dictionary of dictionaries, containing information about the polls at every state and every date.
- hist\_state\_forecast: An array with some historical forecasting information of each state
- poll\_dates: An array whith all the days where a poll was performed
- states: An array with all the states

The model is quite complex and it would be great if someone could tell me if you see any errors in the implementation or if there is a simpler way to write it. Also, performance tips will be appreciated, as my model is taking too much to run (I am using HMC to sample).  
Thanks a lot in advance for your time

---

<div class="post-metadata">

**Author:** ![grero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/grero/32/109_2.png) [@grero](https://discourse.julialang.org/u/grero)\
**Post date:** [November 10, 2020, 5:15am UTC](https://discourse.julialang.org/t/us-elections-model-implementation-in-turing/49864/2 "2020-11-10T05:15:42Z")

</div>

I haven’t had a chance to look through your model, but I think it would be cool if you put your implementation on github. I’d love to look into a bit more when I have more time. Where do you get the polling data that go into the model, by the way?

---

<div class="post-metadata">

**Author:** ![Emmanuel-R8](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/emmanuel-r8/32/11839_2.png) [@Emmanuel-R8](https://discourse.julialang.org/u/Emmanuel-R8)\
**Post date:** [November 10, 2020, 7:40am UTC](https://discourse.julialang.org/t/us-elections-model-implementation-in-turing/49864/3 "2020-11-10T07:40:56Z")

</div>

Data, I would recommend [https://data.fivethirtyeight.com/](https://data.fivethirtyeight.com/).

For a STAN implementation to compare: [https://github.com/TheEconomist/us-potus-model](https://github.com/TheEconomist/us-potus-model)

---

<div class="post-metadata">

**Author:** ![entr0pidelic](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/entr0pidelic/32/15386_2.png) [@entr0pidelic](https://discourse.julialang.org/u/entr0pidelic)\
**Post date:** [November 10, 2020, 3:47pm UTC](https://discourse.julialang.org/t/us-elections-model-implementation-in-turing/49864/4 "2020-11-10T15:47:05Z")

</div>

hello guys, thank you very much for your answers. Here you can see my implementation in a [jupyter notebook](https://github.com/marianLambda/elections_forecast/blob/main/elections_forecast.ipynb)  
I’m using polling data of the 2016 US election from FiveThirtyEight.

I think the biggest problem here is the implementation of the random walks of my model. I can’t find a better way to optimize the code for this (like using filldist() or distarray()). Any suggestion will be truly appreciated

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [January 12, 2025, 4:01pm UTC](https://discourse.julialang.org/t/us-elections-model-implementation-in-turing/49864/5 "2025-01-12T16:01:39Z")

</div>

This post was temporarily hidden by the community for possibly being off-topic, unfocused, inappropriate, or spammy.
