The ElevenLabs steps run on your company's own ElevenLabs plan. Text to speech reads text aloud in a voice you choose and gives back an audio file — a spoken order update to send on WhatsApp, a voice message for a call line. Speech to text turns a recording into text, and can tell the speakers apart.

Who can do this

Workspace Admins and Editors, on every plan.

Before you start

An ElevenLabs connection in Connections — see Connect ElevenLabs.

Steps

Text to speech:

  1. Select + where the audio should be made. In the step picker, open AI, then ElevenLabs, and select Text to speech.
  2. In Connection, choose the connection.
  3. In Voice, choose a voice. The list is read from your ElevenLabs account; type to search it.
  4. In Model, keep the default, or choose another from the list. Multilingual models speak the language the text is in.
  5. In Text, write what is said, with values in {{ }} — Hello {{ $json.name }}, your Northwind Logistics delivery {{ $json.order }} arrives tomorrow morning.
  6. In Format, keep MP3 (the default), or choose WAV.
  7. Select Run this step. Output shows the audio file, with Download to listen to it.

Speech to text:

  1. In the step picker, under ElevenLabs, select Speech to text, and choose the Connection.
  2. Under File, select Add a file and drag in the recording from an earlier step. See Pass files between steps.
  3. In Language, keep Detect (the default), or choose the language spoken.
  4. Optional: switch on Tell speakers apart, so each part of the text says who spoke — speaker 1, speaker 2.
  5. Select Run this step.

Output

Step Keys
Text to speech file (the audio), voice, characters, usage
Speech to text text, language, duration (seconds of audio), segments (each with start, end, text and, with speakers on, speaker), usage

usage holds model and seconds, with characters for Text to speech and audioSeconds for Speech to text. Output shows them as the usage strip. Every test run spends them again — which is what Pin is for.

The audio file travels with the item: drag it into a later step's File or Attachments, or use {{ $json.file }}. It is kept with the run, as every run file is.

Good to know

  • Each item is a separate call. Twenty items make twenty audio files, and use twenty items' worth of characters.
  • Text to speech counts characters against your ElevenLabs plan, spaces and punctuation included.
  • Text up to 5,000 characters a step; split longer text before the step.
  • Dubbing — translating speech into another language in the same voice — is not offered yet.

If something goes wrong

What the run says Why What to do
ElevenLabs refused: … Check the connection's details on Connections. The key is wrong, revoked, or lacks this step's permission. Check the key in ElevenLabs, or replace it on Connections.
ElevenLabs refused: … ElevenLabs refused the request — a voice that was deleted, too many requests for the plan, or no characters left. It is not tried again. Read ElevenLabs's reason.
File is empty. Choose a file from an earlier step. File has no file. Drag in a file from Data from earlier steps.
Attachment 1 is not a file when the step ran — … may not have brought one. The earlier step brought no file in that run. Check the earlier step gives a file.
Text is … characters; one step speaks up to 5,000. Split it first. The text is too long. Split it into several steps or items.
Could not reach ElevenLabs: … / ElevenLabs did not answer within 5 minutes. ElevenLabs did not answer. Try again; set If this step fails to try again.