Batch Transcription resource
Legal notice and public beta
Batch Transcription Configurations use artificial intelligence or machine learning technologies. By enabling or using any of these features or functionalities within Batch Transcription Configurations, you acknowledge and agree that your use of these features or functionalities is subject to the terms of the Predictive and Generative AI/ML Features Addendum.
Batch Transcription Configurations is currently available as a Public Beta release and the information contained in this document is subject to change. Some features are not yet implemented and others may be changed before the product is declared as Generally Available. Public Beta products are not covered by the Twilio Support Terms or Twilio Service Level Agreement.
Batch Transcription Configurations is not PCI compliant or a HIPAA Eligible Service and should not be used in workflows that are subject to HIPAA or PCI.
A Batch Transcription resource represents an asynchronous transcription job for a recorded conversation. To submit audio for transcription, call the Create a Transcription endpoint. You can transcribe Twilio Recordings using a Recording SID, or provide a direct URL to an externally hosted audio file.
Batch Transcription supports several audio formats, each suited for different needs:
- Stereo: two channels that provide spatial sound but don't separate speakers.
- Dual-channel: two distinct audio tracks in the same file, ideal for differentiating speakers such as agents and customers in call recordings. This format improves transcription accuracy and participant differentiation.
For better transcription accuracy, use dual-channel recordings, especially when speaker differentiation is important.
To transcribe a Twilio Recording, provide the Recording SID in the sourceId parameter.
- The recording file size must not exceed 3 GB.
- Audio duration can't exceed eight hours.
- Recordings shorter than two seconds aren't transcribed.
- To transcribe Twilio Recordings stored in external storage, use the
mediaUrlparameter. ThesourceIdparameter isn't supported for externally stored Twilio Recordings. - You can't transcribe encrypted Voice Recordings. Move those recordings to your own external storage, generate pre-signed URLs for the decrypted files, and use the
mediaUrlparameter instead.
To transcribe a recording stored externally, provide the recording's URL in the mediaUrl parameter.
Batch Transcription supports stereo audio for the following formats:
- WAV (PCM-encoded)
- MP3
- FLAC
- The maximum file size allowed is 3 GB.
- The maximum audio length is eight hours.
- The minimum sample rate required is 8 kHz (telephony grade). For best results, use 16 kHz.
Warning
You must provide either sourceId or mediaUrl, but not both.
Optionally, provide a participants array to identify who is on each audio channel. Each participant requires an audioChannelIndex (1 or 2) and can include a type (CUSTOMER, HUMAN_AGENT, or AI_AGENT), address, and name.
When sourceId is provided, participant data is inferred from the call metadata.
When mediaUrl is provided, Twilio has no call metadata to infer participants from, so any participant data must come from the participants array. participants is optional unless you also set conversationConfigurationId to store the transcript in a conversation. In that case, participants is required so that Twilio can attribute the transcript to the right conversation participants.
Info
A transcriptionConfigurationId is required to create a Transcription. This ID identifies the configuration that controls transcription behavior such as engine, language, and callbacks. See the Transcription Configuration resource for details.
To submit a Twilio Recording for transcription, provide the Recording SID in sourceId.
When sourceId is provided, participant data is inferred from the call metadata.
1// Download the helper library from https://www.twilio.com/docs/node/install2const twilio = require("twilio"); // Or, for ESM: import twilio from "twilio";34// Find your Account SID at twilio.com/console5// Provision API Keys at twilio.com/console/runtime/api-keys6// and set the environment variables. See http://twil.io/secure7// For local testing, you can use your Account SID and Auth token8const accountSid = process.env.TWILIO_ACCOUNT_SID;9const apiKey = process.env.TWILIO_API_KEY;10const apiSecret = process.env.TWILIO_API_SECRET;11const client = twilio(apiKey, apiSecret, { accountSid: accountSid });1213async function createV3Transcriptions() {14const transcription = await client.voice.v3.transcriptions.create({15transcriptionConfigurationId:16"voice_transcriptionconfiguration_5pe8jw3ahdmsh7zr06yh4d45x1",17sourceId: "RExxxxxxxxxxxxxxxxxxxxxxxxxxxxx",18participants: [19{20type: "CUSTOMER",21address: "+15558675310",22name: "Dana A.",23audioChannelIndex: 1,24},25{26type: "HUMAN_AGENT",27address: "+15017122661",28name: "Quinn N.",29audioChannelIndex: 2,30},31],32});3334console.log(transcription.status);35}3637createV3Transcriptions();
1{2"status": "PENDING",3"statusUrl": "https://voice.twilio.com/v3/Transcriptions/voice_transcription_7n7hnd7sf68yfv5re42nvvz7aa",4"transcription": {5"id": "voice_transcription_7n7hnd7sf68yfv5re42nvvz7aa",6"accountId": "ACXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX",7"status": "PENDING",8"transcriptionConfigurationId": "voice_transcriptionconfiguration_5pe8jw3ahdmsh7zr06yh4d45x1",9"sourceId": "RExxxxxxxxxxxxxxxxxxxxxxxxxxxxx",10"mediaUrl": null,11"audioStartedAt": "2026-03-10T19:42:16Z",12"conversationId": null,13"participants": [14{ "audioChannelIndex": 1, "type": "CUSTOMER", "address": "+15558675310", "name": "Dana A." },15{ "audioChannelIndex": 2, "type": "HUMAN_AGENT", "address": "+15017122661", "name": "Quinn N." }16],17"duration": null,18"resolvedConfiguration": {19"transcriptionEngine": "deepgram",20"speechModel": "nova-3",21"language": "en-US",22"transcriptionStatusCallback": {23"url": "https://example.com/transcription/callback",24"method": "POST",25"events": null26},27"conversationConfigurationId": "conv_configuration_5pe8jw3ahdmsh7zr06yh4d45x1",28"participantDefaults": [29{ "audioChannelIndex": 1, "type": "CUSTOMER" },30{ "audioChannelIndex": 2, "type": "HUMAN_AGENT" }31]32},33"createdAt": "2026-04-02T19:25:20Z",34"updatedAt": "2026-04-02T19:25:20Z",35"url": "https://voice.twilio.com/v3/Transcriptions/voice_transcription_7n7hnd7sf68yfv5re42nvvz7aa"36}37}
To submit an externally hosted audio file for transcription, provide its URL in mediaUrl.
If you also set conversationConfigurationId, you must include participants. See Specify participant information.
1curl -X POST https://voice.twilio.com/v3/Transcriptions \2-u "$TWILIO_API_KEY:$TWILIO_API_SECRET" \3-H "Content-Type: application/json" \4-d '{5"transcriptionConfigurationId": "voice_transcriptionconfiguration_5pe8jw3ahdmsh7zr06yh4d45x1",6"mediaUrl": "https://example.com/audio/recording.wav",7"audioStartedAt": "2026-03-10T19:42:16Z",8"participants": [9{ "audioChannelIndex": 1, "type": "CUSTOMER", "address": "+15558675310", "name": "Dana A." },10{ "audioChannelIndex": 2, "type": "HUMAN_AGENT", "address": "+15017122661", "name": "Quinn N." }11]12}'
1{2"status": "PENDING",3"statusUrl": "https://voice.twilio.com/v3/Transcriptions/voice_transcription_4abcde1234567890abcde12345",4"transcription": {5"id": "voice_transcription_4abcde1234567890abcde12345",6"accountId": "ACXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX",7"status": "PENDING",8"transcriptionConfigurationId": "voice_transcriptionconfiguration_5pe8jw3ahdmsh7zr06yh4d45x1",9"sourceId": null,10"mediaUrl": "https://example.com/audio/recording.wav",11"audioStartedAt": "2026-03-10T19:42:16Z",12"conversationId": null,13"participants": [14{ "audioChannelIndex": 1, "type": "CUSTOMER", "address": "+15558675310", "name": "Dana A." },15{ "audioChannelIndex": 2, "type": "HUMAN_AGENT", "address": "+15017122661", "name": "Quinn N." }16],17"duration": null,18"resolvedConfiguration": {19"transcriptionEngine": "deepgram",20"speechModel": "nova-3",21"language": "en-US",22"transcriptionStatusCallback": {23"url": "https://example.com/transcription/callback",24"method": "POST",25"events": null26},27"conversationConfigurationId": "conv_configuration_5pe8jw3ahdmsh7zr06yh4d45x1",28"participantDefaults": [29{ "audioChannelIndex": 1, "type": "CUSTOMER" },30{ "audioChannelIndex": 2, "type": "HUMAN_AGENT" }31]32},33"createdAt": "2026-04-02T19:25:20Z",34"updatedAt": "2026-04-02T19:25:20Z",35"url": "https://voice.twilio.com/v3/Transcriptions/voice_transcription_4abcde1234567890abcde12345"36}37}
To receive transcription results to a webhook, set the transcriptionStatusCallback value in your transcription configuration to a URL that you control where you can receive POST requests. When transcription processing completes, Twilio sends a POST request with the full transcript to your callback URL.
Info
A transcription is marked completed when processing succeeds, regardless of webhook delivery status. If webhook delivery ultimately fails and you haven't set a conversationConfigurationId, the transcript content is unrecoverable. For critical workloads, set a conversationConfigurationId on your Transcription Configuration as a backup. This backup stores results in a conversation configuration even if webhook delivery fails.
The webhook request body contains the complete transcription result. Properties with a null value are omitted from the payload, so fields such as mediaUrl and conversationId don't appear in every webhook.
The following table describes the top-level properties:
| Property | Type | Description |
|---|---|---|
id | string | The ID of the transcription job. Begins with voice_transcription_. |
accountId | string | Your Twilio Account SID. |
status | string | The status of the transcription. Either completed or failed. |
transcriptionConfigurationId | string | The ID of the Transcription Configuration used for this transcription. |
sourceId | string or null | The SID of the Twilio Recording, if you're transcribing a recording. |
mediaUrl | string or null | The URL of the external audio file, if you're transcribing third-party media. |
conversationId | string or null | The ID of the conversation, if Twilio stored the transcript. Omitted for webhook-only delivery. |
audioStartedAt | datetime | The start time of the audio recording, in ISO 8601 format. |
duration | integer | The audio duration, in seconds. |
participants | array | An array of resolved participants, each with a type, address, name, and channel index. |
resolvedConfiguration | object | The engine, model, language, and settings applied to this transcription. |
sentences | array | The full transcript content. See the following sentence properties. |
error | object or null | Present when status is failed. See the following error properties. |
createdAt | datetime | The date and time the transcription job was created, in ISO 8601 format. |
updatedAt | datetime | The date and time the status was last updated, in ISO 8601 format. |
The following table describes the properties of each object in the sentences array:
| Property | Type | Description |
|---|---|---|
audioChannelIndex | integer | The audio channel. For dual-channel audio, channel 1 is typically the customer and 2 the agent. |
sentenceIndex | integer | The position within the transcript, starting at 1. |
participantIndex | integer or null | The index into the participants array that identifies the speaker. |
languageCode | string or null | The detected language for this sentence, in BCP-47 format. Present when you set "language": "multi". |
startTimeSeconds | number | When the sentence begins, in seconds from the audio start. |
endTimeSeconds | number | When the sentence ends, in seconds from the audio start. |
text | string | The transcribed text. |
confidence | number | The confidence score, from 0.0 to 1.0. |
words | array | Word-level detail. See the following word properties. |
The following table describes the properties of each object in the words array:
| Property | Type | Description |
|---|---|---|
word | string | The transcribed word. |
startTimeSeconds | number | When the word begins, in seconds from the audio start. |
endTimeSeconds | number | When the word ends, in seconds from the audio start. |
The following table describes the error object, which is present only when status is failed:
| Property | Type | Description |
|---|---|---|
level | string | The severity of the error. |
message | string | A description of what went wrong. |
code | integer | The error code identifying the failure. |
The following example shows a webhook payload for a completed transcription:
1{2"id": "voice_transcription_7n7hnd7sf68yfv5re42nvvz7aa",3"accountId": "ACXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX",4"status": "completed",5"transcriptionConfigurationId": "voice_transcriptionconfiguration_5pe8jw3ahdmsh7zr06yh4d45x1",6"sourceId": "RExxxxxxxxxxxxxxxxxxxxxxxxxxxxx",7"audioStartedAt": "2026-06-15T19:42:16Z",8"duration": 245,9"participants": [10{11"type": "CUSTOMER",12"address": "+18005550100",13"name": "Dana A.",14"audioChannelIndex": 115},16{17"type": "HUMAN_AGENT",18"address": "+18005550101",19"name": "Quinn N.",20"audioChannelIndex": 221}22],23"resolvedConfiguration": {24"transcriptionEngine": "deepgram",25"speechModel": "nova-3",26"language": "en-US",27"transcriptionStatusCallback": {28"url": "https://example.com/transcription-webhook",29"method": "POST"30}31},32"sentences": [33{34"audioChannelIndex": 1,35"sentenceIndex": 1,36"participantIndex": 0,37"startTimeSeconds": 0.08,38"endTimeSeconds": 3.2,39"text": "Hi, I'm calling about my recent order.",40"confidence": 0.97,41"words": [42{ "word": "Hi,", "startTimeSeconds": 0.08, "endTimeSeconds": 0.3 },43{ "word": "I'm", "startTimeSeconds": 0.48, "endTimeSeconds": 0.72 },44{ "word": "calling", "startTimeSeconds": 0.72, "endTimeSeconds": 1.12 },45{ "word": "about", "startTimeSeconds": 1.12, "endTimeSeconds": 1.36 },46{ "word": "my", "startTimeSeconds": 1.36, "endTimeSeconds": 1.52 },47{ "word": "recent", "startTimeSeconds": 1.52, "endTimeSeconds": 1.92 },48{ "word": "order.", "startTimeSeconds": 1.92, "endTimeSeconds": 3.2 }49]50}51],52"createdAt": "2026-06-15T19:42:16Z",53"updatedAt": "2026-06-15T19:42:45Z"54}
The following example shows a webhook payload for a failed transcription, which includes the error object instead of sentences:
1{2"id": "voice_transcription_4abcde1234567890abcde12345",3"accountId": "ACXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX",4"status": "failed",5"transcriptionConfigurationId": "voice_transcriptionconfiguration_5pe8jw3ahdmsh7zr06yh4d45x1",6"mediaUrl": "https://example.com/audio/recording.wav",7"audioStartedAt": "2026-06-15T18:10:05Z",8"participants": [],9"resolvedConfiguration": {10"transcriptionEngine": "deepgram",11"speechModel": "nova-3",12"language": "en-US",13"transcriptionStatusCallback": {14"url": "https://example.com/transcription-webhook",15"method": "POST"16}17},18"error": {19"code": 95111,20"message": "Failed to download media file: unauthorized",21"level": "ERROR"22},23"createdAt": "2026-06-15T18:10:05Z",24"updatedAt": "2026-06-15T18:10:09Z"25}
Before your application acts on a webhook, verify that Twilio sent it. Twilio signs each request with an X-Twilio-Signature header that you can validate using the server-side Twilio helper libraries. For the steps and code samples, see Webhooks security.
If your endpoint is unavailable or returns an error, Twilio retries the request based on the retry count and retry policy configured for your account. To control how many times Twilio retries, which failures it retries on, and the total timeout, see Webhooks connection overrides.
Info
Use the statusUrl from the response to poll for progress. The response includes a Retry-After header with the recommended number of seconds to wait before polling again.
1// Download the helper library from https://www.twilio.com/docs/node/install2const twilio = require("twilio"); // Or, for ESM: import twilio from "twilio";34// Find your Account SID at twilio.com/console5// Provision API Keys at twilio.com/console/runtime/api-keys6// and set the environment variables. See http://twil.io/secure7// For local testing, you can use your Account SID and Auth token8const accountSid = process.env.TWILIO_ACCOUNT_SID;9const apiKey = process.env.TWILIO_API_KEY;10const apiSecret = process.env.TWILIO_API_SECRET;11const client = twilio(apiKey, apiSecret, { accountSid: accountSid });1213async function fetchTranscription() {14const transcription = await client.voice.v315.transcriptions("voice_transcription_7n7hnd7sf68yfv5re42nvvz7aa")16.fetch();1718console.log(transcription.operationId);19}2021fetchTranscription();
1{2"operationId": "voice_transcription_7n7hnd7sf68yfv5re42nvvz7aa",3"status": "COMPLETED",4"statusUrl": "https://voice.twilio.com/v3/Transcriptions/voice_transcription_7n7hnd7sf68yfv5re42nvvz7aa",5"transcription": {6"id": "voice_transcription_7n7hnd7sf68yfv5re42nvvz7aa",7"accountId": "ACXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX",8"status": "COMPLETED",9"transcriptionConfigurationId": "voice_transcriptionconfiguration_5pe8jw3ahdmsh7zr06yh4d45x1",10"sourceId": "RExxxxxxxxxxxxxxxxxxxxxxxxxxxxx",11"mediaUrl": null,12"audioStartedAt": "2026-03-10T19:42:16Z",13"conversationId": "conv_conversation_01k1etx3jbfx88476ccja0889c",14"participants": [15{ "audioChannelIndex": 1, "type": "CUSTOMER", "address": "+18005550100", "name": "Dana A." },16{ "audioChannelIndex": 2, "type": "HUMAN_AGENT", "address": "+18005550101", "name": "Quinn N." }17],18"duration": 120,19"resolvedConfiguration": {20"transcriptionEngine": "deepgram",21"speechModel": "nova-3",22"language": "en-US",23"transcriptionStatusCallback": {24"url": "https://example.com/transcription/callback",25"method": "POST",26"events": null27},28"conversationConfigurationId": "conv_configuration_5pe8jw3ahdmsh7zr06yh4d45x1",29"participantDefaults": [30{ "audioChannelIndex": 1, "type": "CUSTOMER" },31{ "audioChannelIndex": 2, "type": "HUMAN_AGENT" }32]33},34"createdAt": "2026-04-02T19:25:20Z",35"updatedAt": "2026-04-02T19:25:20Z",36"url": "https://voice.twilio.com/v3/Transcriptions/voice_transcription_7n7hnd7sf68yfv5re42nvvz7aa"37}38}
Retrieve a paginated list of the account's transcriptions, newest first. All filters are optional; if you provide more than one, only transcriptions that match every filter are returned.
| Parameter | Type | Description |
|---|---|---|
status | string | Only return transcriptions in this status: PENDING, RUNNING, COMPLETED, or FAILED. |
sourceId | string | Only return transcriptions for this Recording SID (RE followed by 32 lowercase hexadecimal characters). |
languageCode | string | Only return transcriptions whose resolved language matches this value exactly. Case sensitive, for example en-US. |
createdAfter | string | Only return transcriptions created at or after this time (inclusive). ISO 8601. |
createdBefore | string | Only return transcriptions created strictly before this time (exclusive). ISO 8601. |
pageSize | integer | Number of results per page, from 1 to 100. Defaults to 50. |
pageToken | string | Opaque cursor from a previous response's meta.nextToken or meta.previousToken. Only valid for the same filter set that produced it. |
1// Download the helper library from https://www.twilio.com/docs/node/install2const twilio = require("twilio"); // Or, for ESM: import twilio from "twilio";34// Find your Account SID at twilio.com/console5// Provision API Keys at twilio.com/console/runtime/api-keys6// and set the environment variables. See http://twil.io/secure7// For local testing, you can use your Account SID and Auth token8const accountSid = process.env.TWILIO_ACCOUNT_SID;9const apiKey = process.env.TWILIO_API_KEY;10const apiSecret = process.env.TWILIO_API_SECRET;11const client = twilio(apiKey, apiSecret, { accountSid: accountSid });1213async function listV3Transcriptions() {14const transcriptions = await client.voice.v3.transcriptions.list({15status: "COMPLETED",16limit: 20,17});1819transcriptions.forEach((t) => console.log(t.id));20}2122listV3Transcriptions();
1{2"transcriptions": [3{4"id": "voice_transcription_72c3dv9f7r97j85angc18z4bjv",5"accountId": "ACXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX",6"status": "COMPLETED",7"transcriptionConfigurationId": "voice_transcriptionconfiguration_0tce7fytzhr04xvxzvxntwashk",8"sourceId": null,9"mediaUrl": "https://example.com/audio/it-IT.flac",10"audioStartedAt": "2026-08-14T18:19:02Z",11"conversationId": "conv_conversation_01kq0b2vqceb8vy5smkrwvpx5d",12"conversationUrl": "https://maestro.twilio.com/v1/Conversations/conv_conversation_01kq0b2vqceb8vy5smkrwvpx5d",13"participants": [14{ "audioChannelIndex": 1, "type": "CUSTOMER", "address": "+18005550100", "name": "Dana A." },15{ "audioChannelIndex": 2, "type": "AI_AGENT", "address": "+18005550101", "name": "AI Assistant" }16],17"duration": 47,18"resolvedConfiguration": {19"transcriptionEngine": "google",20"speechModel": "chirp_2",21"language": "it-IT",22"conversationConfigurationId": "conv_configuration_5pe8jw3ahdmsh7zr06yh4d45x1",23"participantDefaults": [24{ "audioChannelIndex": 1, "type": "CUSTOMER" },25{ "audioChannelIndex": 2, "type": "AI_AGENT" }26]27},28"createdAt": "2026-08-14T18:19:02Z",29"updatedAt": "2026-08-14T18:19:10Z",30"url": "https://voice.twilio.com/v3/Transcriptions/voice_transcription_72c3dv9f7r97j85angc18z4bjv"31},32{33"id": "voice_transcription_05b5208fvx8dft0n5f85v1bdyz",34"accountId": "ACXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX",35"status": "COMPLETED",36"transcriptionConfigurationId": "voice_transcriptionconfiguration_0tce7fytzhr04xvxzvxntwashk",37"sourceId": null,38"mediaUrl": "https://example.com/audio/en-US.flac",39"audioStartedAt": "2026-08-14T18:19:02Z",40"conversationId": "conv_conversation_01kq0c7htg4m2n9r5x8vw3pqad",41"conversationUrl": "https://maestro.twilio.com/v1/Conversations/conv_conversation_01kq0c7htg4m2n9r5x8vw3pqad",42"participants": [43{ "audioChannelIndex": 1, "type": "CUSTOMER", "address": "+18005550100", "name": "Ravi P." },44{ "audioChannelIndex": 2, "type": "AI_AGENT", "address": "+18005550101", "name": "AI Assistant" }45],46"duration": 63,47"resolvedConfiguration": {48"transcriptionEngine": "google",49"speechModel": "chirp_2",50"language": "en-US",51"conversationConfigurationId": "conv_configuration_5pe8jw3ahdmsh7zr06yh4d45x1",52"participantDefaults": [53{ "audioChannelIndex": 1, "type": "CUSTOMER" },54{ "audioChannelIndex": 2, "type": "AI_AGENT" }55]56},57"createdAt": "2026-08-14T18:19:02Z",58"updatedAt": "2026-08-14T18:19:10Z",59"url": "https://voice.twilio.com/v3/Transcriptions/voice_transcription_05b5208fvx8dft0n5f85v1bdyz"60}61],62"meta": {63"key": "transcriptions",64"pageSize": 2,65"nextToken": "Y3JlYXRlZEF0PTIwMjYtMDgtMTRUMTg6MTk6MDIuNjkyWiZpZD12b2ljZV90cmFuc2NyaXB0aW9uXzA1YjUyMDhmdng4ZGZ0MG41Zjg1djFiZHl6Jl9fc2Y9MA",66"previousToken": null67}68}
Info
Pagination uses opaque cursors, not page numbers. To fetch the next page, send the request again with pageToken set to the meta.nextToken from the previous response; use meta.previousToken to page backward. A token is null when there is no page in that direction.