> ## Documentation Index
> Fetch the complete documentation index at: https://api-tools.memories.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# TikTok Video Caption

> Get TikTok video caption.

<Info>
  **Product**: Visual Intelligence — Social Media Scraping
  **Use case**: Fetch video metadata, transcripts, captions, and comments from YouTube, Instagram, TikTok, and Twitter/X
  **Host**: `https://mavi-backend.memories.ai/serve/api/v2`
  **Auth**: `Authorization: sk-mavi-...` (no `Bearer` prefix)
</Info>

This endpoint allows you to generate visual captions for a TikTok video.

<Note>
  Channel routing guide: see [Social Media Scraping Overview](/visual-intelligence/social-media-scraping-overview). Endpoints with a `channel` request field let you choose `apify`, `rapid`, or `memories.ai`; endpoints without this field use managed routing.
</Note>

<Note>
  **Pricing:** Base fee ($0.01/video) + Input tokens ($0.45/1M) + Output tokens ($3.75/1M) + Duration ($0.0001/sec)
</Note>

<Warning>
  This is an **asynchronous** endpoint. It returns a `task_id` immediately. You must configure a [Webhook](/visual-intelligence/getting-started/webhooks) to receive the processing results.

  If no webhook is configured on the account, the call is **rejected at submit time** with `HTTP 400`:

  ```json theme={null}
  {"code": 400, "msg": "An async request requires at least one webhook.", "data": null}
  ```
</Warning>

### Code Example

<CodeGroup>
  ```python Python theme={null}
  import requests

  BASE_URL = "https://mavi-backend.memories.ai/serve/api/v2"
  API_KEY = "sk-mavi-..."
  HEADERS = {
      "Authorization": f"{API_KEY}"
  }

  def tiktok_video_caption(video_url: str):
      url = f"{BASE_URL}/tiktok/video/mai/transcript"
      data = {"video_url": video_url}
      resp = requests.post(url, headers=HEADERS, json=data)
      return resp.json()

  # Usage example
  result = tiktok_video_caption("https://www.tiktok.com/@cutshall73/video/7543017294226558221")
  print(result)
  ```
</CodeGroup>

### Response

Returns the caption task information.

<ResponseExample>
  ```json Response theme={null}
  {
    "code": 200,
    "msg": "success",
    "data": {
      "task_id": "1cd78354af824c8eb1dafe4ed2435720"
    },
    "failed": false,
    "success": true
  }
  ```

  ```json Callback Response theme={null}
  {
    "code": 200,
    "message": "SUCCESS",
    "data": {
      "videoTranscript": {
        "data": {
          "data": [
            {
              "end_time": 16.0,
              "start_time": 0.0,
              "transcript": "A woman with blonde hair and glasses, wearing a leopard print dress, stands behind a young boy with blonde hair who is sitting in a high chair. The boy is wearing a white t-shirt with a black and red graphic that says SEE YOU LATER. He has food on his face and is holding a small, clear object in his left hand. The woman leans forward, talking to the boy, and then reaches down to adjust something on the high chair tray. The boy smiles and laughs, looking up at the woman, who then adjusts her hair."
            },
            {
              "end_time": 30.0,
              "start_time": 16.0,
              "transcript": "The woman continues to talk to the boy, who is still in the high chair. She points to something on the tray and then reaches down to pick up a small, dark object. The boy watches her, then smiles and laughs again. Another child's blonde hair briefly appears in the bottom left corner of the frame as the woman continues to interact with the boy."
            },
            {
              "end_time": 31.0,
              "start_time": 30.0,
              "transcript": "A young child with blonde hair leans into the frame, looking directly at the viewer. Behind the child, a woman with glasses and blonde hair, wearing a leopard print top, smiles and gestures with her right hand. Another child, a toddler, sits in a high chair in the background, partially obscured."
            }
          ],
          "error_rate": 0.0,
          "usage_metadata": {
            "duration": 0.0,
            "model": "gemini-2.5-flash",
            "output_tokens": 899,
            "prompt_tokens": 13760
          }
        },
        "msg": "Video transcription completed successfully",
        "success": true
      },
      "audioTranscript": {
        "data": {
          "data": [
            {
              "end_time": 0.74,
              "speaker": null,
              "start_time": 0.0,
              "text": " Say hi!"
            },
            {
              "end_time": 1.42,
              "speaker": null,
              "start_time": 1.12,
              "text": " Hi!"
            },
            {
              "end_time": 3.26,
              "speaker": null,
              "start_time": 1.58,
              "text": " Say I'm alive!"
            }
          ],
          "usage_metadata": {
            "duration": 36.53898,
            "model": "whisper-1",
            "output_tokens": 0,
            "prompt_tokens": 0
          }
        },
        "msg": "ASR transcription completed successfully",
        "success": true
      }
    },
    "task_id": "fe8587b278534ed5af6b0e9bcba9ed11"
  }
  ```
</ResponseExample>

### Response Parameters

| Parameter     | Type    | Description                                      |
| ------------- | ------- | ------------------------------------------------ |
| code          | string  | Response code indicating the result status       |
| msg           | string  | Response message describing the operation result |
| data          | object  | Response data object containing task information |
| data.task\_id | string  | Unique identifier of the caption task            |
| success       | boolean | Indicates whether the operation was successful   |
| failed        | boolean | Indicates whether the operation failed           |

### Callback Response Parameters

When the TikTok video caption generation is complete, a callback will be sent to your configured webhook URL.

| Parameter                                                | Type           | Description                                                                     |
| -------------------------------------------------------- | -------------- | ------------------------------------------------------------------------------- |
| code                                                     | string         | Response code (200 indicates success)                                           |
| message                                                  | string         | Status message (e.g., "SUCCESS")                                                |
| data                                                     | object         | Response data object containing both video and audio transcription results      |
| data.videoTranscript                                     | object         | Video transcription result object                                               |
| data.videoTranscript.data                                | object         | Inner data object containing video caption segments and usage information       |
| data.videoTranscript.data.data                           | array          | Array of video caption segments with timestamps                                 |
| data.videoTranscript.data.data\[].start\_time            | number         | Start time of the video segment in seconds                                      |
| data.videoTranscript.data.data\[].end\_time              | number         | End time of the video segment in seconds                                        |
| data.videoTranscript.data.data\[].transcript             | string         | Video transcription text describing the visual content                          |
| data.videoTranscript.data.error\_rate                    | number         | Error rate of the video caption (0.0 means no errors)                           |
| data.videoTranscript.data.usage\_metadata                | object         | Usage statistics for the video caption                                          |
| data.videoTranscript.data.usage\_metadata.duration       | number         | Processing duration in seconds                                                  |
| data.videoTranscript.data.usage\_metadata.model          | string         | The AI model used for video caption (e.g., "gemini-2.5-flash")                  |
| data.videoTranscript.data.usage\_metadata.output\_tokens | integer        | Number of tokens in the generated video caption                                 |
| data.videoTranscript.data.usage\_metadata.prompt\_tokens | integer        | Number of tokens in the input prompt                                            |
| data.videoTranscript.msg                                 | string         | Detailed message about the video caption result                                 |
| data.videoTranscript.success                             | boolean        | Indicates whether the video caption was successful                              |
| data.audioTranscript                                     | object         | Audio transcription result object                                               |
| data.audioTranscript.data                                | object         | Inner data object containing audio transcription segments and usage information |
| data.audioTranscript.data.data                           | array          | Array of audio transcription segments with timestamps                           |
| data.audioTranscript.data.data\[].start\_time            | number         | Start time of the audio segment in seconds                                      |
| data.audioTranscript.data.data\[].end\_time              | number         | End time of the audio segment in seconds                                        |
| data.audioTranscript.data.data\[].text                   | string         | Audio transcription text for this segment                                       |
| data.audioTranscript.data.data\[].speaker                | string \| null | Speaker identifier (null if speaker identification not enabled)                 |
| data.audioTranscript.data.usage\_metadata                | object         | Usage statistics for the audio transcription                                    |
| data.audioTranscript.data.usage\_metadata.duration       | number         | Audio duration in seconds                                                       |
| data.audioTranscript.data.usage\_metadata.model          | string         | The model used for audio transcription (e.g., "whisper-1")                      |
| data.audioTranscript.data.usage\_metadata.output\_tokens | integer        | Number of output tokens (0 for audio transcription)                             |
| data.audioTranscript.data.usage\_metadata.prompt\_tokens | integer        | Number of prompt tokens (0 for audio transcription)                             |
| data.audioTranscript.msg                                 | string         | Detailed message about the audio transcription result                           |
| data.audioTranscript.success                             | boolean        | Indicates whether the audio transcription was successful                        |
| task\_id                                                 | string         | The task ID associated with this transcription request                          |


## OpenAPI

````yaml POST /tiktok/video/mai/transcript
openapi: 3.1.0
info:
  title: Video Metadata & Transcript API Reference
  description: >-
    REST APIs for retrieving video metadata and transcripts from YouTube,
    Instagram, TikTok, and Twitter
  version: v1.0.1
servers:
  - url: https://mavi-backend.memories.ai/serve/api/v2
security:
  - ApiKeyAuth: []
paths:
  /tiktok/video/mai/transcript:
    post:
      summary: TikTok Video Caption
      description: Get TikTok video caption.
      operationId: tiktok_video_mai_transcript
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                video_url:
                  type: string
                  description: The TikTok video URL
                  example: https://www.tiktok.com/@cutshall73/video/7543017294226558221
              required:
                - video_url
      responses:
        '200':
          description: Video transcript
          content:
            application/json:
              schema:
                type: object
                properties:
                  transcript:
                    type: string
                    example: Hello, this is the transcript text...
                  status:
                    type: string
                    example: success
components:
  securitySchemes:
    ApiKeyAuth:
      type: apiKey
      in: header
      name: Authorization

````