# Automatic Speech Recognition in Julia

**URL:** <https://discourse.julialang.org/t/automatic-speech-recognition-in-julia/79003>\
**Category:** Machine Learning\
**Tags:** question\
**Created:** [April 4, 2022, 12:13pm UTC](https://discourse.julialang.org/t/automatic-speech-recognition-in-julia/79003 "2022-04-04T12:13:58Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [April 5, 2022, 5:15pm UTC](https://discourse.julialang.org/t/automatic-speech-recognition-in-julia/79003/2 "2022-04-05T17:15:07Z")

</div>

I don’t find anything either in Julia. Why do you need it implemented (fully) in Julia? Isn’t it good enough to call say PyTorch from Julia and use whatever is available (or Avalon.jl then helpful)? I’m not up-to-speed on moving models from Python to Julia, i.e. just the parameters, weights and biases, shouldn’t that be possible, and wasn’t there even a standard for it ONNX? Might likely just work for certain types of networks, e.g. I believe it’s an older standard than Transformers, so those excluded?

What I did find however brand-new from 31 March 2022:

> **[Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired...](https://arxiv.org/abs/2203.17113)**
>
> This paper studies a novel pre-training technique with unpaired speech data, Speech2C, for encoder-decoder based automatic speech recognition (ASR). Within a multi-task learning framework, we introduce two pre-training tasks for the encoder-decoder...

> Comprehensive experiments on the LibriSpeech corpus show that the proposed Speech2C can relatively reduce the word error rate (WER) by 19.2% over the method without decoder pre-training, and also outperforms significantly the state-of-the-art wav2vec 2.0 and HuBERT on finetuning subsets of 10h and 100h

“Wav2Vec 2.0” was state-of-the-art in 2020, according to its paper 2020 paper, is it still so, even though this other Feb 2022, states so (or is it just an evaluation/survay paper, and they tend to repeat claims?):

If someone DOES want to reimplement something in Julia, I at least would want them to find the state-of-the-art and use that…

> **[Wav2Vec2.0 on the Edge: Performance Evaluation](https://arxiv.org/abs/2202.05993)**
>
> Wav2Vec2.0 is a state-of-the-art model which learns speech representations through unlabeled speech data, aka, self supervised learning. The pretrained model is then fine tuned on small amounts of labeled data to use it for speech-to-text and machine...

Might be a helpful thread:

> [@Speech-based Emotion Recognition](https://discourse.julialang.org/t/speech-based-emotion-recognition/37846):
>
> Hello, I wasn’t able to find any package for Speech-based Mood/Emotion recognition to use in a project I’m working on, so I need some help setting up something myself. I know that LSTM’s are in use for this, I’ve seen WaveNet adapted for almost every other audio problem at this point and I have even seen some of the slightly outdated Spatio-Temporal Box Filters. I guess what I really want to ask is, what would be a good place to start from? Not necessarily the best, most accurate or even fast…

SincNet was also intriguing when I noticed it (might be outdated, or not, hadn’t heard of SpeechBrain):

> **[GitHub - mravanelli/SincNet: SincNet is a neural architecture for efficiently...](https://github.com/mravanelli/SincNet)**
>
> SincNet is a neural architecture for efficiently processing raw audio samples. - GitHub - mravanelli/SincNet: SincNet is a neural architecture for efficiently processing raw audio samples.

> SincNet is implemented in the SpeechBrain ([https://speechbrain.github.io/](https://speechbrain.github.io/)) project as well.

> **[Sinc as a Neural Networks Activation Function - Sefik Ilkin Serengil](https://sefiks.com/2018/01/07/sinc-as-a-neural-networks-activation-function/)**
>
> Sinc function is a sinusoidal activation function in neural networks. In contrast to other common activation functions, it has rises and falls. However, the function saturated and its output converges to zero for large positive and negative...

sinc (and sin) looked intriguing for periodic functions, but may actually be outdated. SIREN is if I recall newer and better, and even something more recent, even better (applications I saw however for computer vision).

I hadn’t heard of conformers (thanks for the tip), only transformers, which it’s a variant of, but might also be too old:

> **[GitHub - sooftware/conformer: \[Unofficial\] PyTorch implementation of...](https://github.com/sooftware/conformer)**
>
> \[Unofficial\] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020) - GitHub - sooftware/conformer: \[Unofficial\] PyTorch implem...

---

_[View the full topic](https://discourse.julialang.org/t/automatic-speech-recognition-in-julia/79003)._
