# PDF to text

- Human page: https://happytails.ai/en/pdf-tools/pdf-to-text/
- Markdown: https://happytails.ai/en/pdf-tools/pdf-to-text.md
- Structured JSON: https://happytails.ai/agents/en/pdf-tools/pdf-to-text/index.json
- Visibility: public

## Tool capabilities

| Property | Value |
| --- | --- |
| Tool ID | `pdf-to-text` |
| Category | pdf-tools |
| Implementation | Registered implementation. Registration alone is not a production-quality guarantee. |
| Browser execution | Use the visible action or interactive controls. |
| Browser processing | Browser; see tool notes for network behavior. |
| HTTP API | Not callable through HTTP. |
| Browser-agent interface | run_current_tool |

### Implementation notes

Text extraction requires selectable text; scanned pages need OCR. Markdown export contains extracted paragraphs, not reconstructed tables or heading semantics.

### Inputs and settings

No free-text input is required. Use the file inputs and/or settings below.

Accepted browser files: `.pdf`. Choose one file.

| Setting | Meaning | Control type | Default | Constraints |
| --- | --- | --- | --- | --- |
| `pages` | Pages (blank = all) | text | "" |  |

### Browser-agent workflow

1. Open the human page in a browser that supports WebMCP.
2. Discover the interface exposed by that page; use its actual schema.
3. Select files on the page first when the tool requires files.
4. Call `run_current_tool` with valid input and settings.
5. Inspect the returned result or error and the displayed output. Retrieve files from the displayed download links.

Calls do not automatically copy or download results. Browser permissions still apply. Interface availability alone is not proof of successful execution.

### Planning specification

Text PDFs only initially; preserve reading order as best effort; detect scanned-only pages.

This describes intended requirements. Use the implementation notes, interface schema and observed output to determine current support.

### Complete machine-readable capability record

```json
{
  "id": "pdf-to-text",
  "category": "pdf-tools",
  "built": true,
  "requirements": "Text PDFs only initially; preserve reading order as best effort; detect scanned-only pages.",
  "implementation": {
    "fields": [
      {
        "key": "pages",
        "label": "Pages (blank = all)",
        "value": "",
        "type": "text"
      }
    ],
    "files": true,
    "input": false,
    "accept": ".pdf",
    "note": "Text extraction requires selectable text; scanned pages need OCR. Markdown export contains extracted paragraphs, not reconstructed tables or heading semantics.",
    "experience": {
      "automatic": false,
      "single": false,
      "label": "Your text",
      "placeholder": "Type or paste your input…",
      "help": "Up to 1,000,000 characters. Review the settings before running.",
      "action": "Convert file"
    },
    "mode": "explicit",
    "interactive": false
  },
  "execution": "Browser interface",
  "api": null,
  "browser_agent": {
    "supported": true,
    "name": "run_current_tool",
    "discovery": "WebMCP on the human page in a supporting browser",
    "verification": "See audit/API-AUDIT.md; availability is not verification"
  }
}
```

## Page guide

Extract selectable text from PDF pages.

## Useful next steps

- [PDF to markdown](https://happytails.ai/en/pdf-tools/pdf-to-markdown/): Save extracted PDF text as a Markdown file. ([Markdown](https://happytails.ai/en/pdf-tools/pdf-to-markdown.md))

- [Word counter](https://happytails.ai/en/text-tools/word-counter/): Count words, characters and lines as you type. ([Markdown](https://happytails.ai/en/text-tools/word-counter.md))

- [Compare text](https://happytails.ai/en/text-tools/compare-text/): Compare two versions of text, line by line. ([Markdown](https://happytails.ai/en/text-tools/compare-text.md))

## How to use PDF to text

1. Choose a file in PDF.
2. Review the settings, then choose Convert file.
3. Review the result, then use Copy result or a download link when available.

## What to expect

Extract selectable text from PDF pages. Choose the pages to read. Scanned image-only pages need OCR, which this tool does not perform. Review reading order in documents with columns or complex layouts.
