Skip to content

Speech (STT & TTS)

Pia includes local speech processing — no cloud APIs required.

Pia runs speech-to-text locally on your machine and offers two engines you can choose between:

  • Parakeet (the default) — multilingual, and it detects the spoken language automatically.
  • Whisper — lets you pick a model size to balance speed and accuracy.
  1. Open Settings → General → Speech
  2. Pick your engine from the Speech-to-text engine dropdown
  3. Click the Download button to grab the required files — you only need to do this once per engine

Parakeet is the default engine. There’s no model size to choose — it works out of the box and detects the language you’re speaking on its own. Just click Download to set it up.

With Whisper, choose a model size from the dropdown: Tiny, Base, Small, Medium, or Large. Then click Download to download that model.

Once your engine is set up, click the microphone icon in the chat to start recording.

Pia synthesizes speech in the app itself, offline, and fetches its voices from Pia’s own servers. There is no separate speech engine to download, and networks that blocked the third-party model host Pia used before can install a voice again.

  1. Open Settings → General → Speech and scroll to Text-to-Speech
  2. Scroll through the voice list — each card shows the voice name, language, gender, and quality
  3. Find a voice you like and click Download
  4. Wait for the progress bar to finish

You can download as many voices as you want. Alba (English GB) is the voice Pia reads aloud in out of the box.

A voice is selected as soon as its download finishes, so speech works without a second step. The card shows an Active badge to confirm which one is in use.

Only one voice can be active at a time. To switch, click Select on a different downloaded voice. If the saved voice is no longer on this device, Pia takes up one that is rather than staying silent.

Remove from this device on a voice’s row deletes it and gives back the space it took. Pia asks first — “Remove the voice X from this device? You can download it again at any time.” — and if it was the voice in use, Pia moves to another installed one.

The recordings Lessac and Ryan were built from are licensed for research and non-commercial use only, so Pia no longer offers them. If you had one selected, Pia moves to another installed voice or asks you to pick one.

Once a voice is active, AI responses are read aloud automatically. Spoken answers begin sooner than they used to, and the short phrases Pia says while it is thinking are ready as soon as a voice is installed.

Instead of typing and reading, you can just talk. Click the Voice conversation button in the message bar and Pia opens a full-screen overlay: it listens, thinks, answers out loud, then listens again. It uses whatever engine and voice you set up above.

  1. Listening — speak normally. A level meter shows Pia is hearing you.
  2. About a second and a half of silence ends your turn automatically, so you don’t have to press anything.
  3. Processing — Pia transcribes what you said and works out a reply.
  4. Speaking — the reply is read aloud, and appears in the running transcript on screen.
  5. Back to listening, until you leave.

Three buttons sit at the bottom:

Button When it shows What it does
Done While listening Ends your turn now, without waiting for the silence.
Stop While speaking Cuts the reply short and starts listening again.
End Always Leaves voice mode.

Everything you say and everything Pia answers is added to the chat you were in, so you can scroll back through it afterwards like any other conversation. Replies are attributed to the active persona, same as typed ones.

There’s no way to show you an approval card when you’re not looking at the screen, so voice mode is stricter than the chat window. A tool that changes something runs only if it’s already covered:

  • you marked it Always allow, or
  • Agent autonomy is on and it’s one of Pia’s own write tools

Anything else is refused out loud, and Pia tells you to ask again in the chat window.

All speech processing happens locally on your device. Audio is never sent to external servers.