Local transcription runs a speech recognition model such as Whisper on your own computer, so the audio never leaves your machine; cloud transcription uploads the audio to a provider's servers and sends back text. Local is usually the better choice for confidential client material and for high volumes once you have suitable hardware, while cloud services win on convenience, speed on weak hardware, and extras such as speaker labelling. If you handle other people's recordings in the UK, the choice also affects your data protection obligations.
This guide compares the two honestly and gives you a checklist for deciding. It is practical guidance, not legal advice.
What "local" and "cloud" actually mean
Local transcription means the model files are downloaded once and run on your hardware. Examples include the open-source >openai/whisper package, community implementations such as whisper.cpp and faster-whisper, and desktop apps built on top of them. After the initial model download, no internet connection is needed.
Cloud transcription means you upload audio to a service, through a website, an app or an API, and the processing happens on the provider's computers. That includes dedicated transcription services, the automatic captions offered by video platforms, and speech-to-text APIs from large cloud providers.
Be careful with the middle ground. Some apps that look local actually upload audio in the background, and some browser tools process files entirely in your browser without uploading anything. Check the privacy policy and, if in doubt, watch whether the tool works with your network disconnected.
Side-by-side comparison
| Factor | Local | Cloud |
|---|---|---|
| Where audio goes | Stays on your machine | Uploaded to the provider (and possibly its sub-processors) |
| Setup | Install software and models; may need a GPU | Account and upload; minimal setup |
| Cost pattern | Hardware and electricity; no per-minute fee | Usually per minute, per hour or subscription |
| Speed | Depends heavily on your GPU or CPU | Generally fast, independent of your hardware |
| Works offline | Yes | No |
| Extras | Depends on the tool; diarisation often needs extra setup | Speaker labels, summaries, editing interfaces often built in |
| Data protection admin | Minimal third-party paperwork | Contracts, transfer checks, provider due diligence |
| Consistency | You control the model version | Provider may change models without notice |
Privacy: the core difference
With local transcription, the question of who else can see the audio mostly disappears. You still need to secure your own devices (encrypted disks, sensible backups, not leaving raw files in shared folders), but you are not relying on anyone else's security or policies. Our guide on protecting client footage covers the device side.
With cloud transcription, you are trusting the provider with the recording. Before uploading anything sensitive, find out:
- Retention. How long are uploaded files and transcripts kept, and can you delete them?
- Training use. Is customer audio used to improve the provider's models, and can you opt out? Terms often differ between free, paid and business or API tiers.
- Location. Where is the data processed and stored, and which other companies (sub-processors) touch it?
- Access. Can the provider's staff listen to recordings, for example for quality review?
- Security. Encryption in transit and at rest, access controls, and any independent certifications they state.
Read the provider's current terms and data processing agreement rather than relying on marketing pages, and check again periodically; terms change.
UK GDPR: what to think about with client audio
This section summarises general principles as we understand them. It is not legal advice; for your own situation, consult the Information Commissioner's Office guidance or a qualified adviser.
A recording of someone's voice, especially alongside a name, a face in video or the content of what they say, is normally personal data under UK GDPR. If a client sends you interview recordings to transcribe, there are usually three parties to think about:
- The client is typically the controller: they decide why the recording is processed.
- You are typically a processor: you handle the data on their instructions.
- A cloud transcription service you use would be a sub-processor acting for you.
Several obligations follow from that:
- Written contracts. UK GDPR Article 28 requires a written contract between a controller and processor with specific terms. Many clients will already expect one with you. If you use a cloud service, you would normally need the client's authorisation (general or specific) to engage a sub-processor, and a contract with that sub-processor offering equivalent protections. Providers usually offer a standard data processing agreement for this.
- International transfers. If the provider processes data outside the UK, that may be a "restricted transfer". It is permitted where the destination is covered by UK adequacy regulations, or where appropriate safeguards are in place, such as the ICO's International Data Transfer Agreement or the UK Addendum to the EU Standard Contractual Clauses. Ask the provider which mechanism it relies on.
- Special category data. Recordings about health, religion, sexuality, political views and similar topics may contain special category data, which needs extra justification and care. Voice data used to identify someone by their voice can count as biometric data. Ordinary transcription is not usually aimed at that, but features such as speaker recognition deserve a closer look.
- Risk assessment. For high-risk processing, the controller may need a data protection impact assessment. Even when one is not required, a short written record of why you chose a particular tool is good practice.
- Transparency and minimisation. People recorded should have been told how their data will be used. Upload only what you need; if only one section needs transcribing, trim the file first.
The >ICO website has guidance on controllers and processors, contracts, and international transfers that explains these points in more detail. Data protection law is periodically amended, so check the ICO site for current guidance rather than relying on older summaries.
The simplest way to reduce this workload is to not send the audio anywhere. Transcribing locally does not remove your responsibilities as a processor, but it removes the sub-processor contract, the transfer question and the provider due diligence.
Cost: how to compare fairly
We will not quote prices, because they change often and vary by plan. Instead, compare on the same basis:
- Cloud: find the provider's current per-minute or per-hour rate for the tier you would actually use, multiply by your realistic monthly volume, and add any seat or subscription fees.
- Local: if you already own a capable computer, the extra cost is mainly electricity and your time. If you would need a new GPU, spread its price over the months you expect to use it, and remember it also speeds up editing and rendering.
- Your time: a less accurate transcript that takes an extra half hour to correct can wipe out any saving. Test accuracy on your own audio before comparing costs.
Low, occasional volumes usually favour cloud services. Regular high volumes, such as a weekly podcast back catalogue or ongoing client interviews, often favour local processing once set up.
Speed and accuracy
Cloud services are generally quick regardless of your hardware. Local speed depends on the model and your machine: on a modern NVIDIA GPU, smaller Whisper models can transcribe much faster than real time, while running a large model on an older laptop's CPU can be slower than real time. Our guide to Whisper model sizes explains the trade-offs.
Accuracy is not automatically better in the cloud. Many services use models of similar quality to the open-source ones. The biggest differences usually come from audio quality and settings; see transcription accuracy tips.
A minimal local run that produces subtitles and never touches the network after the model download:
whisper client_interview.wav --model turbo --language en --output_format srt --output_dir transcripts
A hybrid approach
Many small teams use both:
- Local for anything client-owned, unreleased, legally sensitive or involving members of the public.
- Cloud for your own published content, where the audio is going public anyway and speed or collaboration features matter.
Write this rule down so everyone on the team follows the same approach.
Quick checklist
- Is the audio yours, or a client's or a third party's?
- Does it contain sensitive topics or identify members of the public?
- Has the client authorised use of sub-processors?
- Have you read the provider's current retention, training and location terms?
- Is there a data processing agreement and a lawful transfer mechanism in place?
- Would a local model on your hardware be accurate and fast enough?
- Have you tested accuracy on a representative sample?
- Do you delete uploads and local copies when the job is finished?
FAQ
Is local transcription automatically GDPR compliant?
No. It removes the third-party element, but you still need to secure devices, keep data only as long as needed, and act on the controller's instructions. It simply makes compliance easier to demonstrate.
Are YouTube's automatic captions a cloud service?
Yes. The audio is processed on Google's systems as part of the upload. For content you are publishing on YouTube anyway, that is usually fine; it is a different question for unreleased client material.
Does a browser-based tool upload my audio?
It depends on the tool. Some run the model inside your browser; others upload to a server. Check the privacy information, or test it with your network disconnected. See how browser file transfer works for the underlying technology.
Can I use a free cloud tier for client work?
Check the terms carefully. Free tiers sometimes have different retention or training-use terms from paid business plans, and may not offer a data processing agreement at all.