# Normalizing Values and Floating Point Error

**URL:** <https://discourse.julialang.org/t/normalizing-values-and-floating-point-error/99213>\
**Category:** Statistics\
**Created:** [May 22, 2023, 11:26am UTC](https://discourse.julialang.org/t/normalizing-values-and-floating-point-error/99213 "2023-05-22T11:26:41Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![ianroddis](https://avatars.discourse-cdn.com/v4/letter/i/22d042/32.png) [@ianroddis](https://discourse.julialang.org/u/ianroddis)\
**Post date:** [May 22, 2023, 11:26am UTC](https://discourse.julialang.org/t/normalizing-values-and-floating-point-error/99213/1 "2023-05-22T11:26:41Z")

</div>

Sorry if this isn’t the right place to post this question, but it was as close as I could think for the intersection of concerns.

Most sources I’ve read on data modelling suggest that, when required, values be normalized/standardized ~ N(0,1). If those values are going to be used in matrices for modelling, though, what are the considerations for floating point error? On my system, the machine epsilon is ~ 1e-7 for 64-bit floats (doubles). For very large data sets, there could be many values affected by roundoff error.

My question is: does it make sense to use a different parameterization of N, say N(0, 100), to avoid roundoff error in this case?

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [May 22, 2023, 11:32am UTC](https://discourse.julialang.org/t/normalizing-values-and-floating-point-error/99213/2 "2023-05-22T11:32:09Z")

</div>

> [@ianroddis](#):
>
> On my system, the machine epsilon is ~ 1e-7 for 64-bit floats (doubles).

Do you mean for 32-bit floats? Machine epsilon for 64-bit numbers (at 1.0) is 2e-16, and that doesn’t depend on the system, it’s a property of floating point numbers themselves.

---

<div class="post-metadata">

**Author:** ![ianroddis](https://avatars.discourse-cdn.com/v4/letter/i/22d042/32.png) [@ianroddis](https://discourse.julialang.org/u/ianroddis)\
**Post date:** [May 22, 2023, 11:41am UTC](https://discourse.julialang.org/t/normalizing-values-and-floating-point-error/99213/3 "2023-05-22T11:41:55Z")

</div>

[You are right](https://en.wikipedia.org/wiki/Machine_epsilon), I misinterpreted another value and thought epsilon was 1e-7, so my concerns are likely invalid.

I’d still be curious to know if the conventional wisdom of standardizing to N(0,1) stands for very large data sets (200m+ rows)?

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [May 22, 2023, 11:48am UTC](https://discourse.julialang.org/t/normalizing-values-and-floating-point-error/99213/4 "2023-05-22T11:48:27Z")

</div>

The machine epsilon is a relative measure. Scaling all values (within reasonable limits) won’t make any difference to floating point roundoff errors.

---

<div class="post-metadata">

**Author:** ![ianroddis](https://avatars.discourse-cdn.com/v4/letter/i/22d042/32.png) [@ianroddis](https://discourse.julialang.org/u/ianroddis)\
**Post date:** [May 22, 2023, 11:55am UTC](https://discourse.julialang.org/t/normalizing-values-and-floating-point-error/99213/5 "2023-05-22T11:55:01Z")

</div>

@GunnarFarneback Could you expand on that a bit, please? From my understanding, machine epsilon is the smallest difference that the type supports between two values … i.e. there are ranges of the real numberline where a value will get rounded up or down. The impact of roundoff will increase the smaller the value being rounded.

If I scale all my features / targets to N(0,1), and do many multiplication operations on them, the values between (-1,1) will tend to get smaller and the effect of roundoff error will be magnified, no?

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [May 22, 2023, 12:04pm UTC](https://discourse.julialang.org/t/normalizing-values-and-floating-point-error/99213/6 "2023-05-22T12:04:03Z")

</div>

The distance between two “adjacent” floating point values depends on the magnitude of the values.

```julia
julia> nextfloat(1.0)
1.0000000000000002

julia> nextfloat(1.0) - 1.0
2.220446049250313e-16

julia> nextfloat(1024.0)
1024.0000000000002

julia> nextfloat(1024.0) - 1024.0
2.2737367544323206e-13

julia> eps(1.0)
2.220446049250313e-16

julia> eps(1024.0)
2.2737367544323206e-13

```

---

<div class="post-metadata">

**Author:** ![ianroddis](https://avatars.discourse-cdn.com/v4/letter/i/22d042/32.png) [@ianroddis](https://discourse.julialang.org/u/ianroddis)\
**Post date:** [May 22, 2023, 12:04pm UTC](https://discourse.julialang.org/t/normalizing-values-and-floating-point-error/99213/7 "2023-05-22T12:04:46Z")

</div>

I see, that’s really helpful, thank you!
