# A conformal LSTM classifier to detect anomalies

**URL:** <https://discourse.julialang.org/t/a-conformal-lstm-classifier-to-detect-anomalies/127580>\
**Category:** VS Code\
**Tags:** lstm, conformal-prediction\
**Created:** [April 1, 2025, 12:37am UTC](https://discourse.julialang.org/t/a-conformal-lstm-classifier-to-detect-anomalies/127580 "2025-04-01T00:37:56Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Unnathi\_R\_Shaiva](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/unnathi_r_shaiva/32/216097_2.png) [@Unnathi\_R\_Shaiva](https://discourse.julialang.org/u/Unnathi_R_Shaiva)\
**Post date:** [April 1, 2025, 12:37am UTC](https://discourse.julialang.org/t/a-conformal-lstm-classifier-to-detect-anomalies/127580/1 "2025-04-01T00:37:56Z")

</div>

a conformal LSTM classifier to detect anomalies: and here is the code can somebody help me with the code correction I am new to julia:

# Install required packages

using Pkg  
Pkg.add([“Flux”, “ConformalPrediction”, “DataFrames”, “CSV”, “Plots”, “Statistics”, “StatsPlots”, “Random”])

# Import libraries

using Flux, ConformalPrediction, DataFrames, CSV, Plots, Statistics, Random

# Load the dataset

df = CSV.read(“C:/Users/unnat/OneDrive/Desktop/Individual project/Dataset/ec2\_cpu\_utilization\_24ae8d.csv”, DataFrame)

# Convert timestamps

using Dates  
df[!, “timestamp”] = DateTime.(df[!, “timestamp”], dateformat"yyyy-mm-dd HH:MM:SS")

# Define anomalies (based on dataset reference)

anomalies\_timestamp = [  
“2014-02-26 22:05:00”,  
“2014-02-27 17:15:00”  
]

# Assign labels: -1 (anomaly), 1 (normal)

df[!, “is\_anomaly”] .= 0  
for each in anomalies\_timestamp  
df[df[!, “timestamp”] .== each, “is\_anomaly”] .= 1  
end

# Normalize the data

X = df.value  
X\_normalized = (X .- mean(X)) ./ std(X)  
y = df.is\_anomaly

# Split dataset: Train (80%), Calibration (10%), Test (10%)

n = length(X)  
n\_train = floor(Int, 0.8n)  
n\_calib = floor(Int, 0.1n)

X\_train = X\_normalized[1:n\_train]  
y\_train = y[1:n\_train]

X\_calib = X\_normalized[n\_train+1:n\_train+n\_calib]  
y\_calib = y[n\_train+1:n\_train+n\_calib]

X\_test = X\_normalized[n\_train+n\_calib+1:end]  
y\_test = y[n\_train+n\_calib+1:end]

# Prepare sequences for LSTM (window size = 10)

window\_size = 10  
function create\_sequences(data, window\_size)  
sequences =   
for i in 1:length(data)-window\_size  
push!(sequences, data[i:i+window\_size-1])  
end  
return sequences  
end

X\_train\_seq = create\_sequences(X\_train, window\_size)  
X\_calib\_seq = create\_sequences(X\_calib, window\_size)  
X\_test\_seq = create\_sequences(X\_test, window\_size)

# Convert to Flux format

X\_train\_flux = [reshape(x, 1, window\_size, 1) for x in X\_train\_seq]  
X\_calib\_flux = [reshape(x, 1, window\_size, 1) for x in X\_calib\_seq]  
X\_test\_flux = [reshape(x, 1, window\_size, 1) for x in X\_test\_seq]

# Prepare labels (align with sequences)

y\_train\_seq = y\_train[window\_size+1:end]  
y\_calib\_seq = y\_calib[window\_size+1:end]  
y\_test\_seq = y\_test[window\_size+1:end]

# Define LSTM model

# Add proper initialization and dropout

model = Chain(  
LSTM(1 =\> 32, init=Flux.glorot\_uniform),  
Dropout(0.3),  
LSTM(32 =\> 16, init=Flux.glorot\_uniform),  
Dense(16 =\> 1, sigmoid)  
)

# Reduce learning rate

opt = Adam(0.0001)

# Loss function

# Weight anomalies 10x more than normal points

function weighted\_binary\_crossentropy(y\_pred, y\_true, weights)  
return -mean(weights[1] \* y\_true .\* log.(y\_pred .+ eps()) .+ weights[2] \* (1 .- y\_true) .\* log.(1 .- y\_pred .+ eps()))  
end

loss(x, y) = weighted\_binary\_crossentropy(model(x), y, [10.0f0, 1.0f0])

# Optimizer

using Flux: Adam, params  
opt = Adam(0.001)

# Training function

function train!(model, data, labels, opt)  
for (x, y) in zip(data, labels)  
gs = gradient(() → loss(x, y), params(model))  
Flux.update!(opt, params(model), gs)  
end  
end

# Train the model

epochs = 50  
for epoch in 1:epochs  
Flux.train!(loss, params(model), zip(X\_train\_flux, y\_train\_seq), opt)  
train\_loss = loss(X\_train\_flux[1], y\_train\_seq[1])  
println("Epoch $epoch - Loss: ", train\_loss)  
end

# Compute nonconformity scores for calibration (absolute residuals)

# Compute calibration scores

calib\_scores = [y == 1 ? 1 - model(x)[1] : model(x)[1] for (x, y) in zip(X\_calib\_flux, y\_calib\_seq)]

# Compute test scores

test\_scores = [y == 1 ? 1 - model(x)[1] : model(x)[1] for (x, y) in zip(X\_test\_flux, y\_test\_seq)]

# Compute p-values for test set

function compute\_p\_value(score, calib\_scores)  
return sum(calib\_scores .\>= score) / length(calib\_scores)  
end  
p\_values = [compute\_p\_value(s, calib\_scores) for s in test\_scores]

# Set confidence level (alpha = 0.05 for 95% confidence)

alpha = 0.05  
threshold = quantile(calib\_scores, 1 - alpha)

# Predict anomalies (p-value \< alpha)

y\_pred = [p \< alpha ? -1 : 1 for p in p\_values]

# Confusion Matrix

function confusion\_matrix(y\_true, y\_pred)  
tp = fp = fn = tn = 0  
for (t, p) in zip(y\_true, y\_pred)  
if t == -1 && p == -1  
tp += 1  
elseif t == 1 && p == -1  
fp += 1  
elseif t == -1 && p == 1  
fn += 1  
else  
tn += 1  
end  
end  
return (tp=tp, fp=fp, fn=fn, tn=tn)  
end

cm = confusion\_matrix(y\_test\_seq, y\_pred)

# Print Confusion Matrix

println(“”"  
Confusion Matrix:  
True Positives (Anomalies): (cm.tp) False Positives: (cm.fp)  
False Negatives: (cm.fn) True Negatives: (cm.tn)  
“”")

# Calculate Metrics

tp, fp, fn, tn = cm.tp, cm.fp, cm.fn, cm.tn  
prec = tp / (tp + fp + eps())  
rec = tp / (tp + fn + eps())  
f1 = 2 \* (prec \* rec) / (prec + rec + eps())  
fpr = fp / (fp + tn + eps())

println(“\nAnomaly Detection Metrics:”)  
println("Precision: ", round(prec, digits=4))  
println("Recall: ", round(rec, digits=4))  
println("F1-Score: ", round(f1, digits=4))  
println("FPR: ", round(fpr, digits=4))

# Plot test scores distribution

using StatsPlots  
histogram(test\_scores, bins=20, label=“Model Scores”, color=:blue)  
vline!([threshold], label=“Conformal Threshold”, color=:red)

---

<div class="post-metadata">

**Author:** ![eteppo](https://avatars.discourse-cdn.com/v4/letter/e/90db22/32.png) [@eteppo](https://discourse.julialang.org/u/eteppo)\
**Post date:** [April 2, 2025, 11:45am UTC](https://discourse.julialang.org/t/a-conformal-lstm-classifier-to-detect-anomalies/127580/2 "2025-04-02T11:45:26Z")

</div>

Hey and welcome! Some might be happy and more able to help you if you post a formatted (also put your code between two triple-backticks ```), copy-pasteble version of your script, or at least post the errors and stack traces that you got and found difficult to fix. Also you should post it in a more accurate category, like New to Julia or Machine learning.
