Voice cloning | agent view
The same page content and links, with explicit tool capabilities. Tools marked with an API can run directly through HTTP. Other tools require the browser interface or are not connected yet.
Tool capabilities
{
"id": "voice-cloning",
"category": "speech-tools",
"built": true,
"requirements": "Evaluate consented own-voice synthesis as a separate future workflow with revocable voice profiles and clear disclosure. Do not include celebrity voice presets; rights and model quality require review.",
"implementation": {
"fields": [
{
"key": "consent",
"label": "This is my own voice, or I have permission to synthesize it",
"type": "checkbox",
"value": false
}
],
"files": true,
"accept": ".wav,.mp3,.m4a,.flac,.mp4,.mov,.webm",
"multiple": false,
"input": true,
"processing": "Server required",
"native": true,
"ai": false,
"maxFileBytes": 20971520,
"note": "Create a short English speech sample from a voice you are permitted to use. Upload 3 to 15 seconds of clear speech and enter up to 160 characters. Local Chatterbox processing, built-in watermark, and no saved voice profile. Disclose synthetic audio. Review similarity and pronunciation; results may vary. Allow up to three minutes for this experimental local voice model.",
"extra": true,
"timeoutSeconds": 180,
"experience": {
"automatic": false,
"single": false,
"label": "Your text",
"placeholder": "Type or paste your input…",
"help": "Up to 1,000,000 characters. Review the settings before running.",
"action": "Generate synthetic speech"
},
"mode": "explicit",
"interactive": false
},
"execution": "POST /api/v1/tools/voice-cloning/run",
"api": {
"id": "voice-cloning",
"name": "Voice cloning",
"category": "speech-tools",
"description": "Create a short English speech sample from a voice you are permitted to use.",
"notes": "Create a short English speech sample from a voice you are permitted to use. Upload 3 to 15 seconds of clear speech and enter up to 160 characters. Local Chatterbox processing, built-in watermark, and no saved voice profile. Disclose synthetic audio. Review similarity and pronunciation; results may vary. Allow up to three minutes for this experimental local voice model.",
"human_path": "/en/speech-tools/voice-cloning/",
"method": "POST",
"endpoint": "/api/v1/tools/voice-cloning/run",
"schema_url": "/api/v1/tools/voice-cloning",
"input_schema": {
"type": "object",
"additionalProperties": false,
"properties": {
"input": {
"type": "string",
"maxLength": 1000000,
"default": "",
"description": "Plain text input. File tools use files instead unless otherwise documented."
},
"options": {
"type": "object",
"additionalProperties": false,
"properties": {
"consent": {
"type": "boolean",
"description": "This is my own voice, or I have permission to synthesize it",
"default": false
}
}
},
"files": {
"type": "array",
"maxItems": 1,
"items": {
"type": "object",
"required": [
"name",
"base64"
],
"additionalProperties": false,
"properties": {
"name": {
"type": "string",
"maxLength": 200,
"description": "Filename only, no path."
},
"mime": {
"type": "string",
"maxLength": 150
},
"bytes": {
"type": "integer",
"minimum": 0,
"description": "Optional decoded byte count; must match content if supplied."
},
"base64": {
"type": "string",
"contentEncoding": "base64",
"description": "File bytes as padded base64. Use the tool-specific limits.files_bytes value for the total decoded input size."
}
}
}
}
}
},
"output_schema": {
"type": "object",
"required": [
"tool",
"result"
],
"properties": {
"tool": {
"type": "string"
},
"result": {
"type": "object",
"required": [
"text",
"files"
],
"properties": {
"text": {
"type": [
"string",
"null"
]
},
"data": {
"description": "Parsed JSON when the textual result is JSON."
},
"name": {
"type": "string"
},
"mime": {
"type": "string"
},
"files": {
"type": "array",
"items": {
"type": "object",
"required": [
"name",
"mime",
"bytes",
"base64"
],
"properties": {
"name": {
"type": "string"
},
"mime": {
"type": "string"
},
"bytes": {
"type": "integer"
},
"base64": {
"type": "string",
"contentEncoding": "base64"
}
}
}
}
}
}
}
},
"processing": "Server; self-hosted native engine, no external conversion service",
"limits": {
"request_bytes": 31457280,
"input_characters": 1000000,
"files_bytes": 20971520,
"max_files": 1,
"timeout_seconds": 180
}
},
"agent_processing": "Server; self-hosted native engine, no external conversion service",
"browser_agent": {
"supported": true,
"name": "run_current_tool",
"discovery": "WebMCP on the human page in a supporting browser",
"verification": "See audit/API-AUDIT.md; availability is not verification"
}
}Page guide
Create a short English speech sample from a voice you are permitted to use.
How to use Voice cloning
- Choose a file in WAV, MP3, M4A, FLAC, MP4, MOV or WEBM.
- Review the settings, then choose Generate synthetic speech.
- Review the result, then use Copy result or a download link when available.
What to expect
Create a short English speech sample from a voice you are permitted to use. Create a short English speech sample from a voice you are permitted to use. Upload 3 to 15 seconds of clear speech and enter up to 160 characters. Local Chatterbox processing, built-in watermark, and no saved voice profile. Disclose synthetic audio. Review similarity and pronunciation; results may vary. Allow up to three minutes for this experimental local voice model.