- Auralis STT v0.1
- At a glance
- What makes it different
- Where it sits next to Whisper
- Architecture
- Where it runs inside Auralis
- Quickstart
- Training recipe and provenance
- Experiments we ran, including the one that did not help
- Limitations (read this before you rely on it)
- Roadmap
- Files
- Licence and attribution
- Citation
- Links
- At a glance
Auralis STT v0.1
A 33.5M-parameter English speech-to-text model trained from random initialisation, with its own tokenizer, its own audio frontend and its own training loop. It runs on a laptop CPU inside a Rust runtime, with no cloud call anywhere.
It is the first model in the Auralis project: a local-first, push-to-talk dictation tool (GitHub). This is a v0.1 research release. Read Limitations: it is accurate on clean read speech and clearly behind Whisper base.en.
At a glance
| Architecture | Conformer encoder + CTC head (14 blocks, d_model 320, 4 heads, conv kernel 15) |
| Parameters | 33,471,264 (138 MB as fp32 ONNX) |
| Input | 80-band log-mel, 16 kHz, 25 ms window, 10 ms hop, per-utterance normalised |
| Output | 256-token BPE vocabulary, greedy CTC, lower-case, no punctuation |
| Languages | English |
| Longest input | about 80 s per forward pass (2048 encoder steps) |
| Training data | LibriSpeech only (CC-BY-4.0): about 550 h in the final stage |
| LibriSpeech WER | 7.98% test-clean, 18.25% test-other (full sets, no language model) |
| CPU speed | about 0.09x real time through the Rust runtime (about 11x faster than real time) |
| Licence | MIT (weights and code) |
What makes it different
Most small open STT models are fine-tunes of someone else's weights. This one is not, and the project is built around being able to show its work.
| Auralis STT | Typical open small model (e.g. Whisper base.en) | |
|---|---|---|
| Weights | Random init, trained here. No third-party weights anywhere in the release | Pretrained by a third party; you inherit their data and licence terms |
| Training data | Every dataset passes a licence gate with a recorded licence, training-use and redistribution status. UNKNOWN is rejected, never defaulted in |
Web-scale data, provenance not itemised |
| Frontend | Hand-written log-mel (DFT, STFT, mel bank) in Python, ported to Rust and parity-tested against the reference | Library frontend |
| Runtime | Pure-Rust ONNX (tract), same code path in the desktop app, CLI and local HTTP API |
Python or C++ runtime |
| Privacy | On-device by construction: no cloud STT, no cloud text cleanup, no telemetry | Depends on deployment |
| Reproducibility | Every checkpoint stores git commit, seed, dataset manifest with licences and the full train config | Usually not published with the weights |
| Claims | Scoped: "WER X on dataset Y at configuration Z", never "better than" | Varies |
| Accuracy today | Behind (see next section) | Ahead |
The honest summary: the difference right now is ownership, transparency and size, not accuracy. The point of v0.1 is a model whose every input is known and replaceable, which is the precondition for improving it in the open.
Where it sits next to Whisper
For context, here are LibriSpeech test WERs for Whisper models of similar scale. The Whisper rows are the figures published on the OpenAI model cards on Hugging Face, not something we measured; evaluation settings (text normaliser, decoding) differ from ours, so read the gap as indicative.
| Model | Params | Trained on | test-clean WER | test-other WER |
|---|---|---|---|---|
| Auralis STT v0.1 (measured here, full sets) | 33.5M | LibriSpeech only, 550 h | 7.98% | 18.25% |
| Whisper tiny (card) | 39M | 680k h web audio, 99 languages | 7.54% | 17.15% |
| Whisper base.en (card) | 74M | 680k h web audio | about 4.27% | about 12.80% |
So Auralis v0.1 lands close to Whisper tiny and clearly behind base.en, on about 0.08% of the training audio and with a model less than half the size of base.en. That is the honest starting line for a from-scratch model, and the gap is what the roadmap below is for.
We also scored Auralis end to end through the Rust runtime (ONNX via tract, the path the app uses) on every 5th LibriSpeech clip: 8.56% clean (524 clips) and 18.0% other (588 clips), about 0.09x real time on a laptop CPU, which agrees with the PyTorch numbers. We did not complete a same-clip head-to-head against the Whisper build bundled with the app, because that build decoded far too slowly on our test machine to score at useful scale.
Architecture
16 kHz mono PCM
│
▼
log-mel frontend ── 80 bands, n_fft 512, win 400, hop 160, per-utterance mean/var norm
│ (Python reference ⇄ Rust port, parity-tested)
▼
Conv2D subsampling ×4 ───── 40 ms per encoder step
│
▼
┌─ Conformer block ×14 ──────────────────────────────────┐
│ ½ FFN → multi-head self-attention → conv module → ½ FFN → LayerNorm │
│ d_model 320 · 4 heads · depthwise conv kernel 15 · SiLU · dropout 0.1 │
└────────────────────────────────────────────────────────┘
│
▼
Linear (320 → 256) → log-softmax → greedy CTC → BPE detokenise → text
| Hyper-parameter | Value |
|---|---|
| Blocks / d_model / heads | 14 / 320 / 4 |
| Feed-forward multiplier | 4 |
| Conv module kernel | 15 (depthwise) |
| Subsampling | 2 × Conv2D stride 2 (×4 in time) |
| Positional encoding | sinusoidal |
| Vocabulary | 256 BPE pieces (<blank>, <unk>, language tokens, word-mark ▁) |
Where it runs inside Auralis
MIC → VAD → DENOISER → STT → TEXT CLEANUP → KEYBOARD INSERTION
│ │ │ │
webrtc-vad RNNoise ★ this raw / clean / polished / developer
(trim) (Rust) model modes, local rules only
The desktop app (Tauri 2, Windows) holds Ctrl+Space for push-to-talk and Ctrl+Shift+Space for hands-free dictation, and types the result into whatever window has focus. The same pipeline is exposed as a loopback HTTP API (POST /v1/transcriptions).
Quickstart
1. Download
hf download Sagexd/auralis-stt --local-dir auralis-stt # Hugging Face
# or: kaggle models instances versions download sagexd08/auralis-stt/onnx/v0-1-onnx/1
# or: grab auralis-stt-v0.1.zip from the GitHub release
git clone https://github.com/Sagexd08/Auralis
2. Rust CLI (the same engine the app uses)
cargo run --release -p auralis-runtime --bin auralis-bench-cli -- \
--model auralis-stt/auralis.onnx --input clip.wav --preprocess vad-denoise
3. Python with onnxruntime (tested against the published files)
import sys
sys.path.insert(0, "Auralis/training") # the cloned repo
import numpy as np, onnxruntime as ort, soundfile as sf, torch
from auralis_stt.features import LogMel, normalize_features
from auralis_stt.tokenizer import Tokenizer
wave, sr = sf.read("clip.wav", dtype="float32") # mono, 16 kHz
assert sr == 16000 and wave.ndim == 1
frames = 1 + len(wave) // 160
feats = LogMel()(torch.from_numpy(wave)[None])
feats = normalize_features(feats, torch.tensor([frames])).numpy()
sess = ort.InferenceSession("auralis-stt/auralis.onnx")
log_probs = sess.run(None, {"features": feats, "lengths": np.array([frames], dtype=np.int64)})[0]
tok = Tokenizer.load("auralis-stt/auralis.tokenizer.json")
ids, prev = [], -1
for t in log_probs[0].argmax(-1): # greedy CTC
if t != prev and t != tok.blank_id:
ids.append(int(t))
prev = t
print(tok.decode(ids))
On the bundled demo clip this prints and so my fellow american ask not what your country can do for you as what you can do for your country (plus a stray trailing token, see Limitations).
4. Tensor contract
| Tensor | dtype | shape | notes |
|---|---|---|---|
features |
float32 | [1, 80, frames] |
log-mel, normalised per utterance; frames 16 to 6000 |
lengths |
int64 | [1] |
number of valid frames |
log_probs (out) |
float32 | [1, steps, 256] |
steps ≈ frames / 4; blank is <blank> |
The ONNX graph was checked against the PyTorch model after export: max absolute difference 3.6e-5 on log-probabilities.
Training recipe and provenance
| Stage | Data | Notes |
|---|---|---|
| 1. From scratch | LibriSpeech train-clean-100 + train-clean-360 (460 h) | random init, own BPE tokenizer, checkpoint step 25,108 |
| 2. Continue (this release) | LibriSpeech train-clean-100 + train-other-500, 550.6 h, 166,909 utterances | init from stage 1, final step 26,824 |
Stage-2 settings: batch 32 × grad-accum 2, AdamW (betas 0.9/0.98, weight decay 0.01), peak LR 4e-4 with 500 warm-up steps then cosine decay to 5% of peak, mixed precision, gradient clip 5.0, SpecAugment, max utterance 16 s, seed 1004, 1% speaker-disjoint validation split (981 utterances). Peak VRAM was about 6.3 GB, so it trains on a single consumer or Kaggle GPU.
| Dataset | Licence | Training allowed | Commercial use | Redistribution |
|---|---|---|---|---|
| LibriSpeech train-clean-100 | CC-BY-4.0 | yes | yes | yes |
| LibriSpeech train-other-500 | CC-BY-4.0 | yes | yes | yes |
Each checkpoint on GitHub / Hugging Face carries its config.json and tokenizer; the training manifest (corpus hash 91a2fc7e…bb38) records exactly which datasets, licences and revisions went in.
Experiments we ran, including the one that did not help
We trained a second branch fine-tuned on People's Speech and tried weight-averaging it with this model (both start from the same stage-1 checkpoint, so averaging is meaningful). Full LibriSpeech test sets, greedy CTC:
| Weight on this model | test-clean WER | test-other WER |
|---|---|---|
| 1.00 (this release) | 7.98% | 18.37% |
| 0.90 | 7.94% | 18.47% |
| 0.75 | 8.07% | 19.10% |
| 0.50 | 8.45% | 20.75% |
| 0.00 (People's Speech branch alone) | 10.29% | 27.26% |
Averaging did not improve LibriSpeech accuracy, so v0.1 ships the single model. (The 0.9 row is a 0.04-point test-clean gain, which is within noise, paired with a worse test-other.) The local test-other figure is 18.37% against the 18.25% reported on Kaggle for the same checkpoint; the gap most likely comes from differences between the two evaluation scripts (for example length filtering); we did not chase it down.
Limitations (read this before you rely on it)
- Accuracy: about 8% WER on clean read speech and 18% on LibriSpeech's harder "other" set, greedy decoding, no language model. That is behind Whisper base.en (table above). Expect worse on noisy, conversational, accented or far-field speech: it has only seen audiobook-style read speech.
- English only. No other language was trained.
- No punctuation or casing. Output is lower-case words. The Auralis app adds capitalisation and punctuation with local rules, not with the model.
- Not streaming. The encoder sees the whole utterance; the app segments speech with VAD before decoding. Inputs are capped at about 80 s.
- Trailing artefacts: on the bundled demo clip a stray token (
pap) appears after the final word. VAD trimming in the app mostly hides this; raw model output can show it. - No language model or beam search in the published decoder. Rare names and numbers suffer.
- Evaluation scope: numbers are LibriSpeech test-clean and test-other only. There is no result yet on conversational, noisy or non-US-accent test sets.
- Not a safety or identity system. It transcribes; it does not verify speakers.
Roadmap
This is the plan from the project's build plan. It is a direction, not a schedule or a promise; each step is gated on the previous step's measured results.
| Next | Intent |
|---|---|
| Launch languages | Hindi, Bengali, Japanese on top of English: shared multilingual tokenizer with language tokens, FLEURS and Common Voice data, reported per language |
| Broader evaluation | Held-out FLEURS and Common Voice sets, noisy and conversational tests, and a documented comparison against Whisper base on the same sets |
| Streaming | Causal or chunked encoder with stable partial results for live dictation |
| Smaller and faster | INT8 / INT4 quantisation, a "Nano" model for mobile |
| Our own VAD and denoiser | Replace the borrowed webrtc-vad and RNNoise stages with trained Auralis components |
| More, better data | Scale from hundreds to thousands of eligible hours, hard-example mining, owned recordings with consent |
| Text layer | Model-assisted punctuation and casing, dictated symbols in developer mode ("open paren" → () |
| Later | Text-to-speech research, mobile apps, self-hosted compute; deferred until the model is worth serving |
If a step does not beat the current numbers, it will be reported as a miss, not dropped.
Files
| File | What |
|---|---|
auralis.onnx |
The model (fp32, ONNX opset 18, dynamic frames) |
auralis.tokenizer.json |
256-piece BPE vocabulary and merges |
auralis.json |
Metadata: mel bands, parameter count, languages, source checkpoint |
test_results.json |
Full-set evaluation of this checkpoint |
checkpoint/ (Hugging Face) / auralis-stt-v0.1-checkpoint.zip (GitHub) |
PyTorch model.safetensors, config.json, tokenizer.json for fine-tuning |
Licence and attribution
Weights and code are released under the MIT licence. The model was trained only on LibriSpeech (Panayotov, Chen, Povey, Khudanpur, ICASSP 2015), licensed CC-BY-4.0; please credit it when you use results derived from this model. No third-party model weights are used or redistributed.
Citation
@misc{auralis_stt_2026,
title = {Auralis STT v0.1: a small English CTC speech-to-text model trained from scratch},
author = {{Auralis contributors}},
year = {2026},
howpublished = {\url{https://github.com/Sagexd08/Auralis}},
note = {Release stt-v0.1.0}
}
Links
GitHub · Hugging Face · Kaggle · Benchmark harness · Build plan
Dataset used to train Sagexd/auralis-stt
Evaluation results
- Test WER (greedy CTC on LibriSpeech (clean)test set self-reported7.980
- Test CER on LibriSpeech (clean)test set self-reported2.900
- Test WER (greedy CTC on LibriSpeech (other)test set self-reported18.250
- Test CER on LibriSpeech (other)test set self-reported8.270