Blabb Get Blabb free

Prevent AI data leakage
at the architecture

Cloud speech-to-text and LLM cleanup tools open a network egress path for your most sensitive speech and text — audio, transcripts, and prompts that can be logged, stored, or used to train a model. Run recognition and post-processing entirely on-device and that path simply isn't there. This guide maps the leakage surface to the architectural fix. Informational, not legal advice — confirm specifics with your counsel or compliance team.

01The egress path

Cloud AI dictation opens a data exit

Your speech becomes someone else's data

Speech-to-text and LLM cleanup that run in the cloud route your audio, transcript, and prompts to a vendor's servers. Once there, that content can be logged, retained, and used to train or fine-tune models. The OWASP Top 10 for LLM Applications names "Sensitive Information Disclosure" (LLM06) as a core risk of exactly this pattern, alongside "Training Data Poisoning" (LLM03) and "Model Theft" (LLM10) (OWASP). The egress path itself is the vulnerability.

Leaks are documented, not theoretical

In March 2023 a confirmed OpenAI bug let some ChatGPT users see other users' conversation titles plus names, email and billing addresses, and the last four digits of payment cards — a real disclosure of sensitive user data through an LLM service, not a hypothetical risk (Wikipedia, citing OpenAI's disclosures).

Regulators treated it as a data problem

In March 2023 Italy's data-protection authority banned ChatGPT nationwide over training-data use under the EU GDPR. In July 2023 the U.S. Federal Trade Commission opened a civil investigative demand into OpenAI's data and privacy practices (Wikipedia). NIST's AI Risk Management Framework likewise lists "Privacy-Enhanced" and "Secure and Resilient" as required trustworthiness characteristics of any AI system (NIST AI RMF).

02The architectural fix

Run it on-device and the path no longer exists

No endpoint to exfiltrate to

Blabb's speech recognition (Cohere Transcribe, openly licensed) and AI cleanup (IBM Granite — an instruction-following enterprise LLM, not a chatbot) both run on your own CPU or GPU. No audio is uploaded. No transcript is uploaded. No prompt is uploaded. There is no API endpoint the content could be sent to, so the OWASP LLM06 disclosure surface is structurally absent — not mitigated, absent.

Verify it by unplugging the network

The test for any "private" AI tool is to pull the network cable. Blabb keeps dictating: both models ship inside the installer and run offline. The only outbound call is a licence check (roughly once a day, or every 30 days) through Polar that carries no dictation content — no audio, no text, no prompts. A verified subscription runs for 30 days without ever reaching the licence server.

Local storage is encrypted, retention is yours

What Blabb keeps locally — your history — is SQLite with row-level DPAPI and HMAC-SHA256 encryption, keys tied to your Windows account, plus audit logging with a cryptographic chain of custody and PHI masking on by default. Retention is a setting you choose: don't save history at all, or 1 day, 1 week (the default), 1 month, or forever. Optional crash reports are off by default and scrub sensitive strings.

03Where it matters

Removing the path changes what you can dictate

Regulated and confidential work

For dictated PHI, client communications, and NDA-protected material, "never transmitted" is a cleaner answer than "transmitted under a signed agreement." Because nothing leaves the machine, there is no Business Associate relationship to negotiate and no vendor breach surface for the dictated content. Encyclopedic coverage of generative AI notes that running models locally is motivated precisely by "protection of privacy and intellectual property" (Wikipedia). For the full HIPAA analysis, see our HIPAA & Security page.

Source code and trade secrets

Engineering teams have learned the hard way that pasting source code, internal notes, or credentials into a cloud chatbot hands that material to the vendor. With cleanup running on-device, dictated notes, identifiers, and tokens are processed locally and never added to a third-party corpus. Blabb is also instructed to edit what you said — not add content — and preserves verbatim tokens like SHAs, keys, IPs, and phone numbers.

Locked-down environments

Because Blabb types at the cursor through selectable input methods rather than clipboard paste, it works inside RDP, Citrix, and VDI sessions where paste is blocked and where cloud dictation is often firewalled outright — see our Citrix & RDP guide. Honest limit: it cannot type into elevated or admin (UIPI-protected) windows — a Windows boundary, not a Blabb bug.

You can't leak what never left your machine.

Two weeks free. Your audio, transcripts, and prompts stay on-device.

Get it from Microsoft Store How on-device works

FREE TIER 2,000 WORDS/DAY · 14-DAY UNLIMITED TRIAL IN-APP · CANCEL ANY TIME