Speech (STT & TTS)
Pia includes local speech processing — no cloud APIs required.
Speech-to-Text (STT)
Section titled “Speech-to-Text (STT)”Pia runs speech-to-text locally on your machine and offers two engines you can choose between:
- Parakeet (the default) — multilingual, and it detects the spoken language automatically.
- Whisper — lets you pick a model size to balance speed and accuracy.
Choosing an engine
Section titled “Choosing an engine”- Open Settings → General → Speech
- Pick your engine from the Speech-to-text engine dropdown
- Click the Download button to grab the required files — you only need to do this once per engine
Parakeet
Section titled “Parakeet”Parakeet is the default engine. There’s no model size to choose — it works out of the box and detects the language you’re speaking on its own. Just click Download to set it up.
Whisper
Section titled “Whisper”With Whisper, choose a model size from the dropdown: Tiny, Base, Small, Medium, or Large. Then click Download to download that model.
Recording
Section titled “Recording”Once your engine is set up, click the microphone icon in the chat to start recording.
Text-to-Speech (TTS)
Section titled “Text-to-Speech (TTS)”Pia synthesizes speech in the app itself, offline, and fetches its voices from Pia’s own servers. There is no separate speech engine to download, and networks that blocked the third-party model host Pia used before can install a voice again.
Downloading a Voice
Section titled “Downloading a Voice”- Open Settings → General → Speech and scroll to Text-to-Speech
- Scroll through the voice list — each card shows the voice name, language, gender, and quality
- Find a voice you like and click Download
- Wait for the progress bar to finish
You can download as many voices as you want. Alba (English GB) is the voice Pia reads aloud in out of the box.
Selecting a Voice
Section titled “Selecting a Voice”A voice is selected as soon as its download finishes, so speech works without a second step. The card shows an Active badge to confirm which one is in use.
Only one voice can be active at a time. To switch, click Select on a different downloaded voice. If the saved voice is no longer on this device, Pia takes up one that is rather than staying silent.
Removing a Voice
Section titled “Removing a Voice”Remove from this device on a voice’s row deletes it and gives back the space it took. Pia asks first — “Remove the voice X from this device? You can download it again at any time.” — and if it was the voice in use, Pia moves to another installed one.
Two voices are no longer offered
Section titled “Two voices are no longer offered”The recordings Lessac and Ryan were built from are licensed for research and non-commercial use only, so Pia no longer offers them. If you had one selected, Pia moves to another installed voice or asks you to pick one.
Using TTS
Section titled “Using TTS”Once a voice is active, AI responses are read aloud automatically. Spoken answers begin sooner than they used to, and the short phrases Pia says while it is thinking are ready as soon as a voice is installed.
Voice Conversation
Section titled “Voice Conversation”Instead of typing and reading, you can just talk. Click the Voice conversation button in the message bar and Pia opens a full-screen overlay: it listens, thinks, answers out loud, then listens again. It uses whatever engine and voice you set up above.
How a turn works
Section titled “How a turn works”- Listening — speak normally. A level meter shows Pia is hearing you.
- About a second and a half of silence ends your turn automatically, so you don’t have to press anything.
- Processing — Pia transcribes what you said and works out a reply.
- Speaking — the reply is read aloud, and appears in the running transcript on screen.
- Back to listening, until you leave.
Three buttons sit at the bottom:
| Button | When it shows | What it does |
|---|---|---|
| Done | While listening | Ends your turn now, without waiting for the silence. |
| Stop | While speaking | Cuts the reply short and starts listening again. |
| End | Always | Leaves voice mode. |
Everything you say and everything Pia answers is added to the chat you were in, so you can scroll back through it afterwards like any other conversation. Replies are attributed to the active persona, same as typed ones.
Tools in voice mode
Section titled “Tools in voice mode”There’s no way to show you an approval card when you’re not looking at the screen, so voice mode is stricter than the chat window. A tool that changes something runs only if it’s already covered:
- you marked it Always allow, or
- Agent autonomy is on and it’s one of Pia’s own write tools
Anything else is refused out loud, and Pia tells you to ask again in the chat window.
Privacy
Section titled “Privacy”All speech processing happens locally on your device. Audio is never sent to external servers.