The Deepgram steps run on your company's own Deepgram key. Transcribe audio turns a recording into text, with who spoke when and, if you ask, a summary and the topics discussed — useful for sales calls and support calls. Text to speech reads text aloud in one of Deepgram's voices and gives back an audio file.

Who can do this

Workspace Admins and Editors, on every plan.

Before you start

A Deepgram connection in Connections — see Connect Deepgram.

Steps

Transcribe audio:

  1. Select + where the recording should be transcribed. In the step picker, open AI, then Deepgram, and select Transcribe audio.
  2. In Connection, choose the connection.
  3. In Audio, choose where it comes from:
    • A file — select Add a file and drag in the recording from an earlier step. See Pass files between steps.
    • An address — an https address Deepgram can open without signing in, such as a call recording link from your phone system.
  4. In Model, keep Deepgram's general model (the default), or choose another from the list — there are models for phone calls and meetings.
  5. In Language, keep Detect (the default), or choose the language spoken.
  6. Choose what to add. Punctuation is on; the others are off:
    • Speaker labels — each part says which speaker it was.
    • Summary — a short summary of the whole recording.
    • Topics — what was talked about.
  7. Select Run this step.

Text to speech:

  1. In the step picker, under Deepgram, select Text to speech, and choose the Connection.
  2. In Voice, choose one of Deepgram's voices from the list.
  3. In Text, write what is said, with values in {{ }}.
  4. In Format, keep MP3 (the default), or choose WAV.
  5. Select Run this step. Output shows the audio file, with Download to listen to it.

Output

Step Keys
Transcribe audio text, language, duration (seconds of audio), segments (each with start, end, text and, with speaker labels, speaker), words (each with word, start, end, confidence), summary and topics when asked for, usage
Text to speech file (the audio), voice, characters, usage

usage holds model and seconds, with audioSeconds for Transcribe audio and characters for Text to speech. Output shows them as the usage strip.

Good to know

  • Each item is a separate call, and a separate cost on your Deepgram project. Summary and topics cost extra in Deepgram.
  • An address must be public. Deepgram fetches it itself; an address that needs a sign-in fails.
  • Long recordings take time. The step waits up to five minutes for Deepgram's answer.
  • Summary and topics work best on English recordings; Deepgram may give none for other languages.

If something goes wrong

What the run says Why What to do
Deepgram refused: … Check the connection's details on Connections. The key is wrong, revoked, or lacks this step's permission. Check the key in Deepgram, or replace it on Connections.
Deepgram refused: … Deepgram refused the request — audio it cannot read, an address it cannot open, or no balance left. It is not tried again. Read Deepgram's reason.
File is empty. Choose a file from an earlier step. File has no file. Drag in a file from Data from earlier steps.
Attachment 1 is not a file when the step ran — … may not have brought one. The earlier step brought no file in that run. Check the earlier step gives a file.
Address must be an https address Deepgram can open without signing in. An address was chosen, and it is not https. Give the recording's https address, or choose A file.
Could not reach Deepgram: … / Deepgram did not answer within 5 minutes. Deepgram did not answer. Try again; set If this step fails to try again.