The ElevenLabs steps run on your company's own ElevenLabs plan. Text to speech reads text aloud in a voice you choose and gives back an audio file — a spoken order update to send on WhatsApp, a voice message for a call line. Speech to text turns a recording into text, and can tell the speakers apart.
Who can do this
Workspace Admins and Editors, on every plan.
Before you start
An ElevenLabs connection in Connections — see Connect ElevenLabs.
Steps
Text to speech:
- Select + where the audio should be made. In the step picker, open AI, then ElevenLabs, and select Text to speech.
- In Connection, choose the connection.
- In Voice, choose a voice. The list is read from your ElevenLabs account; type to search it.
- In Model, keep the default, or choose another from the list. Multilingual models speak the language the text is in.
- In Text, write what is said, with values in
{{ }}— Hello {{ $json.name }}, your Northwind Logistics delivery {{ $json.order }} arrives tomorrow morning. - In Format, keep MP3 (the default), or choose WAV.
- Select Run this step. Output shows the audio file, with Download to listen to it.
Speech to text:
- In the step picker, under ElevenLabs, select Speech to text, and choose the Connection.
- Under File, select Add a file and drag in the recording from an earlier step. See Pass files between steps.
- In Language, keep Detect (the default), or choose the language spoken.
- Optional: switch on Tell speakers apart, so each part of the text says who spoke — speaker 1, speaker 2.
- Select Run this step.
Output
| Step | Keys |
|---|---|
| Text to speech | file (the audio), voice, characters, usage |
| Speech to text | text, language, duration (seconds of audio), segments (each with start, end, text and, with speakers on, speaker), usage |
usage holds model and seconds, with characters for Text to speech and audioSeconds for Speech to text. Output shows them as the usage strip. Every test run spends them again — which is what Pin is for.
The audio file travels with the item: drag it into a later step's File or Attachments, or use {{ $json.file }}. It is kept with the run, as every run file is.
Good to know
- Each item is a separate call. Twenty items make twenty audio files, and use twenty items' worth of characters.
- Text to speech counts characters against your ElevenLabs plan, spaces and punctuation included.
- Text up to 5,000 characters a step; split longer text before the step.
- Dubbing — translating speech into another language in the same voice — is not offered yet.
If something goes wrong
| What the run says | Why | What to do |
|---|---|---|
| ElevenLabs refused: … Check the connection's details on Connections. | The key is wrong, revoked, or lacks this step's permission. | Check the key in ElevenLabs, or replace it on Connections. |
| ElevenLabs refused: … | ElevenLabs refused the request — a voice that was deleted, too many requests for the plan, or no characters left. It is not tried again. | Read ElevenLabs's reason. |
| File is empty. Choose a file from an earlier step. | File has no file. | Drag in a file from Data from earlier steps. |
| Attachment 1 is not a file when the step ran — … may not have brought one. | The earlier step brought no file in that run. | Check the earlier step gives a file. |
| Text is … characters; one step speaks up to 5,000. Split it first. | The text is too long. | Split it into several steps or items. |
| Could not reach ElevenLabs: … / ElevenLabs did not answer within 5 minutes. | ElevenLabs did not answer. | Try again; set If this step fails to try again. |