Blabb Get Blabb free

Cloud dictation has
bad days. Local doesn't.

If your dictation tool suddenly got slower, or its accuracy quietly changed, you're not imagining it — and you're not alone. Cloud dictation runs on shared servers, swaps models server-side, and reaches you over the public internet, so its performance moves for reasons that have nothing to do with you. Blabb runs entirely on your PC: the same hardware gives you the same speed and the same accuracy today, on deadline day, and a year from now. This page lays out the public evidence, and how we hold our own models to a fixed standard.

01The evidence

Three ways cloud dictation degrades — all documented by the providers themselves

This isn't speculation; it's the providers' own status pages and lifecycle policies. If a cloud service transcribes your voice and rewrites your text, every one of these applies to both stages at once.

Shared capacity means someone else's peak is your lag

Cloud speech runs on GPU fleets shared across every customer. OpenAI's public status history logged three dozen incidents in four months (May–August 2026) — including "Selected Model is at Capacity" errors, "increased server-overload errors", elevated latency with interrupted streaming, and voice-mode availability degraded twice. Even 99.9% uptime is roughly nine hours a year where dictation simply doesn't work — and latency dips don't show up in uptime numbers at all.

The network path is part of the product

Deepgram — a serious speech infrastructure company — logged a nearly four-hour degradation of both its batch and streaming speech-to-text on 4 August 2026. The cause wasn't their models: it was an ISP network failure on the path between users and their servers. With cloud dictation, the internet between you and the datacentre is a component you depend on and can't see. On-device dictation has no such component.

Models get swapped — or switched off — server-side

Microsoft's published model lifecycle retires generally-available models on a fixed ~18-month clock, after which inference returns "410 Gone"; preview deployments are "force-upgraded" with "no option to remain". Their own migration guidance tells customers to re-apply prompt engineering "to match prior accuracy" — a provider acknowledging that the swap changes your output. Your dictation tool's accuracy can change overnight, in either direction, because someone else refreshed a deployment.

02Why it's structural

It isn't bad luck. It's the architecture.

A cloud dictation tool puts two AI models on someone else's servers: the speech model that hears you, and the rewrite model that tidies the text. Both inherit the same dependencies — shared GPU capacity, a model lifecycle the provider controls, and a network path you don't. None of them are things you can schedule around, and none of them show up in a feature list when you're choosing the tool.

The rewrite degrades too, not just the transcription

Providers market the AI polish as a feature, but it's the same cloud dependency wearing a nicer name. When the rewrite model is swapped, upgraded, or re-prompted server-side, your cleanup behavior changes with it — a rewrite that was restrained last month might be breezier this month, and no setting on your end explains it.

The economics point one direction

Cloud providers pay for GPU capacity by the hour, so faster and cheaper models get adopted on infrastructure grounds — speed and cost per request — whether or not they transcribe better. Capacity is elastic and rationed at peak; models consolidate as old ones retire. For the provider that's prudence. For you, it means the product you bought is a moving target.

None of it is visible to you

There's no changelog entry for "we rebalanced capacity" or "we moved you to a smaller model" — just a tool that feels a bit slower, or output that reads a bit different, and a support forum where other users describe the same thing. When the variable lives on someone else's server, you can't even confirm what changed.

03What local changes

Your PC's resources are dedicated to exactly one user

Blabb's speech recognition and its rewrite model both run on your machine. That single architectural fact removes every dependency above — not by promising harder, but by not having the dependency in the first place.

No other users, no peak hours

Your CPU and GPU don't get busier because a million other people started dictating. Tuesday at 9am, Sunday at midnight, deadline day: the resources available to your dictation are the same, because they're yours. Performance is a property of your hardware, and your hardware doesn't have a rush hour.

No network path, no outages

Nothing you say leaves the machine, so there's no route to degrade, no ISP between you and your transcription, and no provider status page to check. Dictation works the same with the network down — an already-verified licence keeps working offline for 30 days at a time.

You control the upgrade path

Models are pinned in the app you installed. Nothing swaps server-side, nothing retires on a clock, and a new Blabb version lands on your machine only when you choose to update it — with release notes you can read first. If a version works for you, it keeps working, on the same terms, for as long as you keep it.

The infrastructure bill is constant

There is no per-minute compute cost riding on your dictation, so there's no pricing pressure to quietly make it cheaper to serve. Your cost is your PC, which you already own, and a licence that costs the same regardless of how much you dictate.

04Our side of the bargain

Local removes their variables. These rules govern ours.

"It runs on your PC" isn't a quality claim by itself — local software can also ship bad models. So we hold our own model choices to a written standard, because the only variable left in your dictation is us.

Accuracy before speed — always

We will not swap the speech engine because a new one is faster or cheaper to run. The transcription model is chosen on measured accuracy — it's why Blabb ships Cohere Transcribe, which led the Open ASR Leaderboard on release at 5.42% word error rate. Speed is tuned around that choice, never traded against it.

Faithfulness before flourish

The rewrite model's job is to follow a closed list of permitted edits — fix grammar, handle spoken commands, keep your meaning intact — and nothing else. We don't adopt a rewrite model because it writes prettier; if it can't follow the rules, it doesn't ship. The whole standard is written up in how we tune for faithfulness.

Every change re-baselined, against a fixed suite

Any model or prompt change runs the gauntlet before release: a 440-scenario suite — including a doubled-length cohort and separate medical and legal suites — that scores faithfulness case by case. The bar we hold ourselves to is 95%-plus faithfulness to the original transcription, and no release regresses it. If a candidate model introduces a new class of hallucination, the suite catches it and it doesn't reach you. Those tests aren't internal theatre: the failures they pin are the ones we work through in public (see the anti-slop work).

Open about what changed, and when

When we do adopt a new model, you'll hear it from us — what changed, why, and what the suite says before and after. No silent swaps, no "we rebalanced capacity". Upgrades arrive as releases you choose to install, described in release notes you can read before you do.

05The honest edges

Local has a ceiling. It's just a stable one.

Fairness cuts both ways. On-device dictation is bounded by your hardware: a slower or older PC means slower transcription, and we publish exactly what affects speed and how to tune for lower-spec machines. If your GPU is busy running a game, dictation competes with it. That's the honest trade: local performance is a function of your machine — but it's the same function every day, and it's one you can measure, plan for, and upgrade on your own schedule.

The other honest admission: cloud tools genuinely improve over time too — server-side upgrades can make a product better without you doing anything. The difference is who decides and what verification happens before it reaches your screen. With Blabb, an improvement arrives only after it passes the suite, only in a release you choose, and only with the numbers published. Stability with a visible upgrade path, or drift with a status page: that's the real choice.

Reliability isn't a promise. It's architecture.

Two weeks free, models in the installer, dictation that runs on your PC.

Get it from Microsoft Store How on-device dictation works

FREE TIER 2,000 WORDS/DAY · 14-DAY UNLIMITED TRIAL IN-APP · CANCEL ANY TIME