Draft preview: not yet available for purchase

The address notify@substrike.com.au is a placeholder: the mailbox is not live yet, so a message sent now may bounce.

How we compared models

The method behind the model comparison on the home page: which clips, how they were scored, which versions ran, and why it is a small test.

Measured 27 September 2026, except where a row says otherwise.

The clips

Three clips of our own multiplayer voice chat, recorded in OBS, 383 reference words and about 3.2 minutes of audio in total.

  • Clip one: 120 seconds, 176 words. Answer key made in the Substrike editor: the app’s transcript, corrected by hand and marked as checked.
  • Clip two: 27 seconds, 51 words. Answer key made the same way.
  • Clip three: 45 seconds, 156 words, three players talking over each other with game music. Answer key typed from scratch by hand, with no model output to start from.

All three were checked by one person, the developer. Because two answer keys began as whisper output, a word whisper heard and nobody changed counts as right for whisper based setups and wrong for other models. Clip three has no such head start.

How it was scored

  • Word error rate: every substituted, missing or extra word against the answer key counts as one error. Case and punctuation are ignored, and bracketed sound tags are removed.
  • The same scoring code for every model. It is the word error maths the app itself uses.
  • Local models were given the same audio: the app’s own filtered, mono 16 kHz version of each clip. Models built for short audio were fed the clip in pieces under 30 seconds, cut at quiet points.
  • ElevenLabs Scribe was run on both the filtered audio and the original, unfiltered recording. The home page shows its better result, on the original audio.
  • Each model ran once, with its default settings. No names list or custom vocabulary was given to any model, Substrike included.

Results

A selection from the 26 local model setups scored, plus the cloud service. Fewer is better.

Errors per clip and in total, 383 words, 27 September 2026
ModelClip oneClip twoClip threeTotalPer 100 words
Substrike as shipped whisper large-v3-turbo with repeat cleanup39115110126.4
Mistral Voxtral Mini 3B44124910527.4
whisper large-v3-turbo without the repeat cleanup51115211429.8
Qwen3-ASR 1.7B56156713836.0
ElevenLabs Scribe v1 cloud, original audio67185814337.3
OpenAI whisper medium.en58147314537.9
ElevenLabs Scribe v1 cloud, filtered audio67237016041.8
NVIDIA Parakeet TDT 0.6B v265159417445.4

Differences under about 30 errors in total are within the noise of a test this size: small changes that should not matter moved one clip by up to 36 errors in earlier runs. Read Substrike and Voxtral Mini as level.

Model versions

  • Substrike: whisper large-v3-turbo through whisper.cpp 1.9.3 on the CPU, English, beam search 5, with the app’s repeat cleanup. This is what the app ships today.
  • Mistral Voxtral Mini 3B (2507 release), Hugging Face transformers, bf16, whole clip.
  • Qwen3-ASR 1.7B, Hugging Face transformers, bf16, English forced.
  • OpenAI whisper medium.en, whisper.cpp 1.9.3, same settings as Substrike.
  • NVIDIA Parakeet TDT 0.6B v2, int8, sherpa-onnx 1.12.15, whole clip.
  • ElevenLabs Scribe v1 (eleven_scribe_v1), through the ElevenLabs connector, language detected automatically. That route offers only v1 and no settings.
  • All local models ran on the CPU of one desktop PC. Speed was not compared, because other work shared the machine.

The earlier cloud test, 17 September 2026

A different measurement, so it cannot be set against the table above. Clip three only (152 words in that version of the answer key), the whisper filtered audio for every service, scored as the share of answer key words each transcript matched. One run each.

  • Substrike (whisper large-v3-turbo): 68.4% matched.
  • ElevenLabs Scribe v2: 65.8% matched.
  • AssemblyAI Universal: 56.6% matched.
  • Deepgram Nova-3: 52.0% matched.

The filtered audio was tuned for whisper and may not suit other services. On 27 September the original audio helped Scribe v1 (143 errors against 160), so the cloud results here may understate those services a little.

Cloud transcription prices

Raw pay-as-you-go transcription API rates from the services in the earlier cloud test above, checked from each vendor’s own pricing page on 27 September 2026. These are developer APIs, not finished captioning apps: no editor, no caption styling, no export, you would still have to build all of that. Speechmatics was checked too, but its own pricing page shows two different numbers for the same plan, so it is left out here as unverified.

Price per hour of audio, and the cost at 10 hours a month, checked 27 September 2026.
ServicePrice per hourCost per year at 10 h/month
ElevenLabs Scribe v2US$0.22US$26.40
AssemblyAI Universal-3.5 Pro with diarizationUS$0.23US$27.60
Deepgram Nova-3 mono, pre-recorded, diarization includedUS$0.258US$30.96
OpenAI whisper-1 / gpt-4o-transcribe no diarizationUS$0.36US$43.20
Google Cloud Speech-to-Text V2 Standard, diarization includedUS$0.96US$115.20

Word timing and voice matching

  • Word timing: clip three is our only clip with word times set by hand (149 words). Six larger timing models were run through the app’s own alignment steps. None placed words measurably closer than the model Substrike ships, which put the typical word start 53 ms from the hand timing.
  • Voice matching: 19 speaker models from the sherpa-onnx release were tried on short samples from our clips. The one Substrike ships was in a four way statistical tie at the top. Samples from different nights were not tested.

How thin this is

  • 3 clips, 383 words, one reviewer, our own gaming voice chat from one recording setup and one game genre. It is not a general benchmark, and results on your audio may differ.
  • Public leaderboards rank several of these models above whisper on read speech and meetings. Quiet, overlapping voice chat over game audio is a different test.
  • One setting per model, one run each. A model with a better long audio recipe, a names list or tuned settings could do better.
  • Cloud models change over time. These are the versions we reached on the dates given.

Scores and censored transcripts are available on request through the support page. The clips hold our friends’ voices, so they are shared only with their agreement.

Model and company names are trademarks of their owners. Substrike is not affiliated with or endorsed by them.

Back to the comparison