# PDF to markdown

- Human page: https://happytails.ai/en/pdf-tools/pdf-to-markdown/
- Markdown: https://happytails.ai/en/pdf-tools/pdf-to-markdown.md
- Structured JSON: https://happytails.ai/agents/en/pdf-tools/pdf-to-markdown/index.json
- Visibility: public

## Tool capabilities

| Property | Value |
| --- | --- |
| Tool ID | `pdf-to-markdown` |
| Category | pdf-tools |
| Implementation | Registered implementation. Registration alone is not a production-quality guarantee. |
| Browser execution | Use the visible action or interactive controls. |
| Browser processing | Browser; see tool notes for network behavior. |
| HTTP API | Not callable through HTTP. |
| Browser-agent interface | run_current_tool |

### Implementation notes

Text extraction requires selectable text; scanned pages need OCR. Markdown export contains extracted paragraphs, not reconstructed tables or heading semantics.

### Inputs and settings

No free-text input is required. Use the file inputs and/or settings below.

Accepted browser files: `.pdf`. Choose one file.

| Setting | Meaning | Control type | Default | Constraints |
| --- | --- | --- | --- | --- |
| `pages` | Pages (blank = all) | text | "" |  |

### Browser-agent workflow

1. Open the human page in a browser that supports WebMCP.
2. Discover the interface exposed by that page; use its actual schema.
3. Select files on the page first when the tool requires files.
4. Call `run_current_tool` with valid input and settings.
5. Inspect the returned result or error and the displayed output. Retrieve files from the displayed download links.

Calls do not automatically copy or download results. Browser permissions still apply. Interface availability alone is not proof of successful execution.

### Planning specification

Headings, tables and reading order require fidelity tests; distinguish text extraction from OCR.

This describes intended requirements. Use the implementation notes, interface schema and observed output to determine current support.

### Complete machine-readable capability record

```json
{
  "id": "pdf-to-markdown",
  "category": "pdf-tools",
  "built": true,
  "requirements": "Headings, tables and reading order require fidelity tests; distinguish text extraction from OCR.",
  "implementation": {
    "fields": [
      {
        "key": "pages",
        "label": "Pages (blank = all)",
        "value": "",
        "type": "text"
      }
    ],
    "files": true,
    "input": false,
    "accept": ".pdf",
    "note": "Text extraction requires selectable text; scanned pages need OCR. Markdown export contains extracted paragraphs, not reconstructed tables or heading semantics.",
    "experience": {
      "automatic": false,
      "single": false,
      "label": "Your text",
      "placeholder": "Type or paste your input…",
      "help": "Up to 1,000,000 characters. Review the settings before running.",
      "action": "Convert file"
    },
    "mode": "explicit",
    "interactive": false
  },
  "execution": "Browser interface",
  "api": null,
  "browser_agent": {
    "supported": true,
    "name": "run_current_tool",
    "discovery": "WebMCP on the human page in a supporting browser",
    "verification": "See audit/API-AUDIT.md; availability is not verification"
  }
}
```

## Page guide

Save extracted PDF text as a Markdown file.

## Useful next steps

- [PDF to text](https://happytails.ai/en/pdf-tools/pdf-to-text/): Extract selectable text from PDF pages. ([Markdown](https://happytails.ai/en/pdf-tools/pdf-to-text.md))

- [Markdown to HTML](https://happytails.ai/en/text-tools/markdown-to-html/): Convert Markdown into HTML source. ([Markdown](https://happytails.ai/en/text-tools/markdown-to-html.md))

- [Markdown to PDF](https://happytails.ai/en/document-tools/markdown-to-pdf/): Render defined Markdown dialect with print CSS, fonts and page breaks. ([Markdown](https://happytails.ai/en/document-tools/markdown-to-pdf.md))

## How to use PDF to markdown

1. Choose a file in PDF.
2. Review the settings, then choose Convert file.
3. Review the result, then use Copy result or a download link when available.

## What to expect

Save extracted PDF text as a Markdown file. The export contains extracted paragraphs. It does not reconstruct heading levels, tables, or images. Scanned image-only pages need a separate OCR tool.
