Semantic deduplication | agent view
The same page content and links, with explicit tool capabilities. Tools marked with an API can run directly through HTTP. Other tools require the browser interface or are not connected yet.
Tool capabilities
{
"id": "semantic-deduplication",
"category": "data-tools",
"built": true,
"requirements": "Group near-duplicate text rows in uploaded CSV; set similarity threshold; review proposed groups before export; never automatically discard records based only on embeddings.",
"implementation": {
"fields": [
{
"key": "column",
"label": "Text column name",
"type": "text",
"value": "text"
},
{
"key": "threshold",
"label": "Similarity threshold",
"type": "number",
"value": 0.85,
"min": 0.5,
"max": 1
}
],
"files": true,
"multiple": false,
"input": false,
"processing": "Server required",
"native": true,
"extra": true,
"maxFileBytes": 20971520,
"accept": ".csv",
"note": "Local MiniLM embeddings propose similarity groups. At most 500 rows with 1,000 characters per text value. All rows are retained for review; no automatic deletion.",
"experience": {
"automatic": false,
"single": false,
"label": "Your text",
"placeholder": "Type or paste your input…",
"help": "Up to 1,000,000 characters. Review the settings before running.",
"action": null
},
"mode": "explicit",
"interactive": false
},
"execution": "POST /api/v1/tools/semantic-deduplication/run",
"api": {
"id": "semantic-deduplication",
"name": "Semantic deduplication",
"category": "data-tools",
"description": "Local MiniLM embeddings propose similarity groups.",
"notes": "Local MiniLM embeddings propose similarity groups. At most 500 rows with 1,000 characters per text value. All rows are retained for review; no automatic deletion.",
"human_path": "/en/data-tools/semantic-deduplication/",
"method": "POST",
"endpoint": "/api/v1/tools/semantic-deduplication/run",
"schema_url": "/api/v1/tools/semantic-deduplication",
"input_schema": {
"type": "object",
"additionalProperties": false,
"properties": {
"input": {
"type": "string",
"maxLength": 1000000,
"default": "",
"description": "Plain text input. File tools use files instead unless otherwise documented."
},
"options": {
"type": "object",
"additionalProperties": false,
"properties": {
"column": {
"type": "string",
"description": "Text column name",
"default": "text"
},
"threshold": {
"type": "number",
"description": "Similarity threshold",
"default": 0.85,
"minimum": 0.5,
"maximum": 1
}
}
},
"files": {
"type": "array",
"maxItems": 1,
"items": {
"type": "object",
"required": [
"name",
"base64"
],
"additionalProperties": false,
"properties": {
"name": {
"type": "string",
"maxLength": 200,
"description": "Filename only, no path."
},
"mime": {
"type": "string",
"maxLength": 150
},
"bytes": {
"type": "integer",
"minimum": 0,
"description": "Optional decoded byte count; must match content if supplied."
},
"base64": {
"type": "string",
"contentEncoding": "base64",
"description": "File bytes as padded base64. Use the tool-specific limits.files_bytes value for the total decoded input size."
}
}
}
}
}
},
"output_schema": {
"type": "object",
"required": [
"tool",
"result"
],
"properties": {
"tool": {
"type": "string"
},
"result": {
"type": "object",
"required": [
"text",
"files"
],
"properties": {
"text": {
"type": [
"string",
"null"
]
},
"data": {
"description": "Parsed JSON when the textual result is JSON."
},
"name": {
"type": "string"
},
"mime": {
"type": "string"
},
"files": {
"type": "array",
"items": {
"type": "object",
"required": [
"name",
"mime",
"bytes",
"base64"
],
"properties": {
"name": {
"type": "string"
},
"mime": {
"type": "string"
},
"bytes": {
"type": "integer"
},
"base64": {
"type": "string",
"contentEncoding": "base64"
}
}
}
}
}
}
}
},
"processing": "Server; self-hosted native engine, no external conversion service",
"limits": {
"request_bytes": 31457280,
"input_characters": 1000000,
"files_bytes": 20971520,
"max_files": 1,
"timeout_seconds": 90
}
},
"agent_processing": "Server; self-hosted native engine, no external conversion service",
"browser_agent": {
"supported": true,
"name": "run_current_tool",
"discovery": "WebMCP on the human page in a supporting browser",
"verification": "See audit/API-AUDIT.md; availability is not verification"
}
}Page guide
Local MiniLM embeddings propose similarity groups.
How to use Semantic deduplication
- Choose a file in CSV.
- Review the settings, then choose Run Semantic deduplication.
- Review the result, then use Copy result or a download link when available.
What to expect
Local MiniLM embeddings propose similarity groups. Local MiniLM embeddings propose similarity groups. At most 500 rows with 1,000 characters per text value. All rows are retained for review; no automatic deletion.