Speaker diarization | agent view
The same page content and links, with explicit tool capabilities. Tools marked with an API can run directly through HTTP. Other tools require the browser interface or are not connected yet.
Tool capabilities
{
"id": "speaker-diarization",
"category": "speech-tools",
"built": true,
"requirements": "Label speaker turns as Speaker 1 and Speaker 2; let users merge and rename speakers; review overlap; export timestamped JSON. Does not identify people by voice.",
"implementation": {
"fields": [
{
"key": "speakers",
"label": "Number of speakers (0 to estimate)",
"type": "number",
"value": 0,
"min": 0,
"max": 6,
"step": 1
}
],
"files": true,
"accept": ".wav,.mp3,.m4a,.flac,.mp4,.mov,.webm",
"multiple": false,
"input": false,
"processing": "Server required",
"native": true,
"ai": false,
"maxFileBytes": 20971520,
"note": "Group non-overlapping voices using local ECAPA speaker embeddings. Up to one minute. Returns speaker labels and approximate 250 ms timestamps. It does not identify people or separate overlapping speech. Review short turns and noisy recordings.",
"extra": true,
"experience": {
"automatic": false,
"single": false,
"label": "Your text",
"placeholder": "Type or paste your input…",
"help": "Up to 1,000,000 characters. Review the settings before running.",
"action": null
},
"mode": "explicit",
"interactive": false
},
"execution": "POST /api/v1/tools/speaker-diarization/run",
"api": {
"id": "speaker-diarization",
"name": "Speaker diarization",
"category": "speech-tools",
"description": "Group non-overlapping voices using local ECAPA speaker embeddings.",
"notes": "Group non-overlapping voices using local ECAPA speaker embeddings. Up to one minute. Returns speaker labels and approximate 250 ms timestamps. It does not identify people or separate overlapping speech. Review short turns and noisy recordings.",
"human_path": "/en/speech-tools/speaker-diarization/",
"method": "POST",
"endpoint": "/api/v1/tools/speaker-diarization/run",
"schema_url": "/api/v1/tools/speaker-diarization",
"input_schema": {
"type": "object",
"additionalProperties": false,
"properties": {
"input": {
"type": "string",
"maxLength": 1000000,
"default": "",
"description": "Plain text input. File tools use files instead unless otherwise documented."
},
"options": {
"type": "object",
"additionalProperties": false,
"properties": {
"speakers": {
"type": "number",
"description": "Number of speakers (0 to estimate)",
"default": 0.0,
"minimum": 0,
"maximum": 6,
"multipleOf": 1
}
}
},
"files": {
"type": "array",
"maxItems": 1,
"items": {
"type": "object",
"required": [
"name",
"base64"
],
"additionalProperties": false,
"properties": {
"name": {
"type": "string",
"maxLength": 200,
"description": "Filename only, no path."
},
"mime": {
"type": "string",
"maxLength": 150
},
"bytes": {
"type": "integer",
"minimum": 0,
"description": "Optional decoded byte count; must match content if supplied."
},
"base64": {
"type": "string",
"contentEncoding": "base64",
"description": "File bytes as padded base64. Use the tool-specific limits.files_bytes value for the total decoded input size."
}
}
}
}
}
},
"output_schema": {
"type": "object",
"required": [
"tool",
"result"
],
"properties": {
"tool": {
"type": "string"
},
"result": {
"type": "object",
"required": [
"text",
"files"
],
"properties": {
"text": {
"type": [
"string",
"null"
]
},
"data": {
"description": "Parsed JSON when the textual result is JSON."
},
"name": {
"type": "string"
},
"mime": {
"type": "string"
},
"files": {
"type": "array",
"items": {
"type": "object",
"required": [
"name",
"mime",
"bytes",
"base64"
],
"properties": {
"name": {
"type": "string"
},
"mime": {
"type": "string"
},
"bytes": {
"type": "integer"
},
"base64": {
"type": "string",
"contentEncoding": "base64"
}
}
}
}
}
}
}
},
"processing": "Server; self-hosted native engine, no external conversion service",
"limits": {
"request_bytes": 31457280,
"input_characters": 1000000,
"files_bytes": 20971520,
"max_files": 1,
"timeout_seconds": 90
}
},
"agent_processing": "Server; self-hosted native engine, no external conversion service",
"browser_agent": {
"supported": true,
"name": "run_current_tool",
"discovery": "WebMCP on the human page in a supporting browser",
"verification": "See audit/API-AUDIT.md; availability is not verification"
}
}Page guide
Group non-overlapping voices using local ECAPA speaker embeddings.
How to use Speaker diarization
- Choose a file in WAV, MP3, M4A, FLAC, MP4, MOV or WEBM.
- Review the settings, then choose Run Speaker diarization.
- Review the result, then use Copy result or a download link when available.
What to expect
Group non-overlapping voices using local ECAPA speaker embeddings. Group non-overlapping voices using local ECAPA speaker embeddings. Up to one minute. Returns speaker labels and approximate 250 ms timestamps. It does not identify people or separate overlapping speech. Review short turns and noisy recordings.