Two OpenAI steps work with sound on your company's own OpenAI key. Transcribe audio — Turns a recording into text. Text to speech — Turns text into speech in one of OpenAI's voices, and gives back an audio file that later steps can attach, upload or send.
Who can do this
Workspace Admins and Editors, on every plan.
Before you start
An OpenAI connection in Connections — see Connect OpenAI.
Steps
Transcribe audio:
- Select + where the recording should be turned into text. In the step picker, open AI, then OpenAI, and select Transcribe audio.
- In Connection, choose the connection with your key.
- Under File, select Add a file and drag in the recording from an earlier step — MP3, M4A, WAV, OGG or WebM, up to 25 MB. See Pass files between steps.
- Optional: in Model, choose one from the list read from OpenAI. Left empty, the step uses OpenAI's current transcription model. Choose a Whisper model to get the parts of the recording with their times as well.
- Optional: in Language, keep Detect, or choose the language the recording is in — it helps with short recordings.
- Optional: in Prompt, list words the recording uses that a model might mishear — Consulace, Northwind Logistics, NW-20417.
- Select Run this step.
Text to speech:
- In the step picker, under OpenAI, select Text to speech, and choose the Connection.
- In Voice, keep Alloy, or choose another of OpenAI's voices.
- Optional: in Model, choose one from the list. Left empty, the step uses OpenAI's current speech model.
- In Text, type what to say, with values in
{{ }}— Hello {{ $json.name }}, your Northwind Logistics delivery arrives tomorrow morning. Up to 4,096 characters. - In Format, keep MP3, or choose WAV.
- Select Run this step. Output shows the audio, with Download.
Output
Transcribe audio gives text, language (a two-letter code), duration (seconds, when OpenAI says), segments (each with start, end and text — Whisper models only) and usage — model, audioSeconds or tokensIn and tokensOut, and seconds.
Text to speech gives file (the audio), voice, characters and usage — model, characters and seconds. The audio travels with the item: drag it into a later step's File or Attachments from Data from earlier steps.
Output shows usage as the usage strip.
Good to know
- Each item is a separate call, and a separate cost on your OpenAI account. Pin the output while you build — see Pin a step's output.
- A step may take up to five minutes; a long recording takes longest.
- Longer recordings must be split first, or sent compressed — an MP3 or M4A is far smaller than a WAV.
If something goes wrong
| What the run says | Why | What to do |
|---|---|---|
| OpenAI refused: … Check the connection's details on Connections. | The key is wrong or revoked. | Replace the key on Connections. |
| OpenAI refused: … | OpenAI refused the request — an audio format it cannot read, a rate limit, or no credit left. It is not tried again. | Read OpenAI's reason. |
| File is empty. Choose a file from an earlier step. | File has no file. | Drag in a recording from Data from earlier steps. |
| … is over 25 MB, the most OpenAI transcribes in one go. Shorten it, or send it compressed (MP3 or M4A). | The recording is too large. | Shorten or compress it first. |
| Text is … characters; one step speaks up to 4,096. Split it first. | Text is too long. | Split the text across items. |
| Could not reach OpenAI: … / OpenAI did not answer within 5 minutes. | OpenAI did not answer. | Try again; set If this step fails to try again. |