# Weights in gaussian process

**URL:** <https://discourse.julialang.org/t/weights-in-gaussian-process/70646>\
**Category:** Statistics\
**Tags:** gaussian-process, stheno\
**Created:** [October 30, 2021, 1:11am UTC](https://discourse.julialang.org/t/weights-in-gaussian-process/70646 "2021-10-30T01:11:03Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![jzr](https://avatars.discourse-cdn.com/v4/letter/j/eb9ed0/32.png) [@jzr](https://discourse.julialang.org/u/jzr)\
**Post date:** [October 30, 2021, 1:11am UTC](https://discourse.julialang.org/t/weights-in-gaussian-process/70646/1 "2021-10-30T01:11:03Z")

</div>

I have a model like @willtebbutt’s TemporalGPs example [https://github.com/JuliaGaussianProcesses/TemporalGPs.jl#learning-kernel-parameters-with-optimjl-parameterhandlingjl-and-zygotejl](https://github.com/JuliaGaussianProcesses/TemporalGPs.jl#learning-kernel-parameters-with-optimjl-parameterhandlingjl-and-zygotejl)

In my dataset, different observations should have different weights to account for the sampling procedure. In StatsBase, there are functions like `fit(formula, tbl, wts=weightvec)` ([`StatsBase.fit`](https://juliastats.org/GLM.jl/stable/api/#StatsBase.fit)). How can I apply sampling weights to data points in the GP fit?

I realize there are different types of weighting and I don’t want to use the wrong invocation by accident (as [cautioned here](https://github.com/JuliaStats/GLM.jl/issues/327#issuecomment-515429469)).

---

<div class="post-metadata">

**Author:** ![willtebbutt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/willtebbutt/32/6790_2.png) [@willtebbutt](https://discourse.julialang.org/u/willtebbutt)\
**Post date:** [October 31, 2021, 10:14pm UTC](https://discourse.julialang.org/t/weights-in-gaussian-process/70646/2 "2021-10-31T22:14:13Z")

</div>

> In my dataset, different observations should have different weights to account for the sampling procedure.

Could you elaborate a bit on this? My first thought would be to incorporate your weights through the observation variance, but I’d like to figure out whether this is a reasonable thing to do or not.

---

<div class="post-metadata">

**Author:** ![jzr](https://avatars.discourse-cdn.com/v4/letter/j/eb9ed0/32.png) [@jzr](https://discourse.julialang.org/u/jzr)\
**Post date:** [October 31, 2021, 11:08pm UTC](https://discourse.julialang.org/t/weights-in-gaussian-process/70646/3 "2021-10-31T23:08:24Z")

</div>

The dataset is a phone survey produced by a survey company, conducted over a period of time. It includes demographic information about respondents such as sex, age, race. I am producing estimates of a certain variable over time. In the collected survey data, some demographic groups appear more or less often than they do in the real population (as determined by census data). In order to produce estimates representative of the population, the survey company adds to each row a weight, indicating how much that response should be weighted in order to match the population distribution, which is generated by [iterative proportional fitting](https://en.wikipedia.org/wiki/Iterative_proportional_fitting). For example, on a certain day of surveying if 77-year-old white women appeared more often than they do in the population, such a response might be assigned a sampling weight of 0.8. Then my estimates and predictions will produce results that match the population distribution, rather than the distribution of responses that happened to occur on that day.

---

<div class="post-metadata">

**Author:** ![willtebbutt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/willtebbutt/32/6790_2.png) [@willtebbutt](https://discourse.julialang.org/u/willtebbutt)\
**Post date:** [November 1, 2021, 4:47pm UTC](https://discourse.julialang.org/t/weights-in-gaussian-process/70646/4 "2021-11-01T16:47:23Z")

</div>

Hmm interesting. I’m really not sure what the appropriate way to think about these weights is 😬 . It’s not obvious to me that it’s appropriate to downweight observations via changing the likelihood variance based on your description, and there’s not really another way that I’m aware of to re-weight observations in the log marginal likelihood / for doing posterior inference.

edit: this isn’t a problem I’ve got a lot of experience with though, so I might be missing something.

---

<div class="post-metadata">

**Author:** ![Eric](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eric/32/25565_2.png) [@Eric](https://discourse.julialang.org/u/Eric)\
**Post date:** [November 4, 2021, 10:09am UTC](https://discourse.julialang.org/t/weights-in-gaussian-process/70646/5 "2021-11-04T10:09:30Z")

</div>

For the question about weights, I often go back to that sas article [https://blogs.sas.com/content/iml/2017/10/02/weight-variables-in-statistics-sas.html](https://blogs.sas.com/content/iml/2017/10/02/weight-variables-in-statistics-sas.html)

With this terminology, Julia’s GLM provides frequencies weights, and my LinearRegression package provides analytical weights. To my knowledge at this time, there is no implementation of regression with survey weights.  
Although StatsBase defines Probability weights which I think are survey weights.  
[https://juliastats.org/StatsBase.jl/stable/weights/[https://juliastats.org/StatsBase.jl/stable/weights/](https://juliastats.org/StatsBase.jl/stable/weights/)](https://juliastats.org/StatsBase.jl/stable/weights/%5Bhttps://juliastats.org/StatsBase.jl/stable/weights/%5D(https://juliastats.org/StatsBase.jl/stable/weights/)) this could be a starting point.

---

<div class="post-metadata">

**Author:** ![willtebbutt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/willtebbutt/32/6790_2.png) [@willtebbutt](https://discourse.julialang.org/u/willtebbutt)\
**Post date:** [November 4, 2021, 10:21am UTC](https://discourse.julialang.org/t/weights-in-gaussian-process/70646/6 "2021-11-04T10:21:36Z")

</div>

Correct me if I’m wrong, but don’t these various weighting schemes assume models of the form

\log p(y\_{1:N} | x\_{1:N}, \theta) = \prod\_{n=1}^N \log p(y\_n | x\_n, \theta)

? If so, this isn’t really what’s going on with GPs, at least not obviously (maybe there’s a way to cast them in this framework then marginalise over \theta?)

---

<div class="post-metadata">

**Author:** ![Eric](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eric/32/25565_2.png) [@Eric](https://discourse.julialang.org/u/Eric)\
**Post date:** [November 4, 2021, 12:41pm UTC](https://discourse.julialang.org/t/weights-in-gaussian-process/70646/7 "2021-11-04T12:41:53Z")

</div>

I can’t comment on the math formulation.  
I would also expect the tool needed to have features that enable to correctly:

- relate the sampled population to the overall population accounting for the strata of interest in the study.
- and also some methodology/strategy to account for the non-respondent (in this case the ones that do not answer the phone) and how they relate to the population and/or its strata.  
While this is far from my comfort zone, I guess there  
Is voluminous literature on the subject.

I have seen two packages that appear to be related to the topic:  
[https://github.com/jamanrique/SurveyAnalysis.jl](https://github.com/jamanrique/SurveyAnalysis.jl)  
And  
[https://github.com/grahamstark/SurveyDataWeighting.jl](https://github.com/grahamstark/SurveyDataWeighting.jl)

Hope this helps a little.

---

<div class="post-metadata">

**Author:** ![opera\_malenky](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/opera_malenky/32/8213_2.png) [@opera\_malenky](https://discourse.julialang.org/u/opera_malenky)\
**Post date:** [November 4, 2021, 1:51pm UTC](https://discourse.julialang.org/t/weights-in-gaussian-process/70646/8 "2021-11-04T13:51:07Z")

</div>

This might be terribly hacky, but perhaps just do a separate GP for each demographic cell of interest? Then maybe look at the weighted mean and variance of those individual processes?
