What is voice-to-text — and how do you pick an app?
Voice-to-text — also called speech-to-text — converts spoken words into written text. It is used two ways: live dictation, where words appear as you speak them in whatever application you are in, and transcription, where spoken or recorded audio becomes a text record. What separates the tools is where the audio is processed (on your device or in the cloud), whether use is limited, and which platforms are covered.
Dictation and transcription are not the same thing
Both turn speech into text, but the job differs. Dictation is live: you speak and the words appear in your document, chat, email, or code as you go — the tool is a faster keyboard. Transcription is after the fact: an existing recording — a meeting, an interview, a voice memo — is turned into a written record. A product built for one is not automatically good at the other, so the first question is which you actually need.
ClassEve's Lven Instant and Lven Cloud are live voice-to-text: dictation that types what you said at your cursor. They are not file-transcription tools — you do not hand them a recording. If you need a text record of a live conversation you are part of, dictate it; if you need to transcribe an existing audio file, that is a different category of tool.
On-device vs cloud — where your voice goes
The single choice that shapes privacy, cost, and reliability is where the speech model runs. Cloud voice-to-text streams your microphone audio to a server and returns text; it can be accurate on weak hardware, but every word transits off your machine and nothing works offline. On-device voice-to-text runs the model on your own hardware, so audio is never uploaded and dictation keeps working with no network — at the cost of a realistic minimum spec.
This is a property you can verify rather than a promise you have to trust: if the model runs locally, the audio has nowhere to go. It is why privacy-conscious users, and anyone who dictates on planes or restricted networks, reach for the on-device kind.
What to check before you pick one
Five questions decide it. Where is the audio processed — on your device or in the cloud? What does the tool do to your words after recognition — a deterministic clean-up that keeps exactly what you said, or a generative rewrite that can change it? Are there limits — daily caps or per-minute metering? Does it cover the platforms you actually use? And what does it cost — free built-in, one-time purchase, or subscription — and do you keep anything if you stop paying?
Raw accuracy, the spec people ask about first, has largely converged for everyday speech on modern engines. The five questions above now separate voice-to-text products more than accuracy does.