PDF, scanned documents, and OCR

Convert PDF to images, clean scanned documents, and run OCR online

Render PDF pages as JPG/PNG/WEBP, remove gray paper, deskew, correct perspective, redact details, extract signatures/stamps, and build searchable PDFs on your device.

PDF → JPG/PNG/WEBPTool Office
Deskew + perspectiveTool Office
Local redactionTool Office
12-language OCRTool Office

Document Image Tools

One canonical engine covers the scan lifecycle: page extraction, cleanup, geometry correction, privacy, OCR, and result packaging without separate pages for every filter or language.

Loading tool

Choose the correct PDF and scan workflow

Start from the required output. Do not run OCR when page images are enough, and verify redaction on the exported file rather than relying only on the preview.

PDF to image

Select DPI and format to extract up to 30 pages, then download one image or a ZIP in source order.

Scan cleanup

Use grayscale, thresholding, gray-background removal, rotation, deskew, and perspective correction before OCR or archiving.

OCR and privacy

Redact visible data, recognize one of 12 supported languages, export text/basic TSV, and optionally create a searchable PDF.

Convert every PDF page to an image

PDF.js renders pages at a selected DPI instead of taking screenshots.

  1. Choose one PDF up to 25 MB and 30 pages.
  2. Select JPG, PNG, or WEBP and 150, 200, or 300 DPI.
  3. Process the file and inspect the first output page.
  4. Download one image or a ZIP containing all pages in PDF order.
  • 150 DPI suits screen viewing; 300 DPI creates larger files for print/OCR.
  • Password-protected or corrupt PDFs may not render.

Clean a gray or skewed scan

Document filtering separates paper from ink but can remove very fine strokes.

  1. Select Enhance and filter scanned documents.
  2. Try grayscale first, then add moderate contrast.
  3. Use Whiten paper or threshold black and white and adjust the threshold.
  4. Enable automatic deskew or set an angle manually, then inspect small text.
  • Keep the original color scan for comparison.
  • Automatic deskew estimates only ±5° and does not replace perspective correction for strongly skewed photos.

Correct a photographed document's perspective

Four source corners are mapped to a new rectangular image.

  1. Select the page you want to correct.
  2. Estimate the border for an initial rectangular crop.
  3. Adjust all eight X/Y controls until corners match the paper edges.
  4. Process and inspect text proportions, page edges, and cropped content.
  • A background that contrasts with the paper helps border estimation.
  • Current auto-detection estimates a rectangle; manually adjust trapezoid-shaped pages.

Redact information before sharing

Redaction creates a new rasterized file with the selected region covered.

  1. Select Redact and the page to edit.
  2. Adjust the region's position and size on the preview.
  3. Prefer solid black for removal; blur/pixelation may be unsuitable for sensitive data.
  4. Export, reopen, and zoom in to verify the detail cannot be read.
  • The current workflow applies one configured region per page in a run.
  • Do not share the source file when only the exported copy has been redacted.

Extract a signature or stamp

Signature mode uses darkness; stamp mode identifies dominant red ink and creates alpha.

  1. Use an evenly lit scan with a white background.
  2. Choose Signature/dark ink or Stamp/red ink.
  3. Adjust the threshold to keep strokes while removing paper.
  4. Export PNG to preserve transparency and inspect stroke edges.
  • Do not use extracted marks to forge signatures or documents.
  • Paper shadows or colors close to the ink reduce extraction accuracy.

OCR an image or scanned PDF

PDF pages are rendered first, then each Canvas is recognized by a local OCR Worker.

  1. Choose the matching OCR language and document type.
  2. Enable searchable PDF if a text layer is required.
  3. Start OCR and allow the language model to download on first use.
  4. Proofread text, numbers, amounts, language-specific characters, and TSV output before use.
  • OCR accuracy is not guaranteed; compare the result with the source.
  • Basic TSV does not perfectly reconstruct merged cells or complex table layouts.

Process private documents in the browser

PDF/images are read only by PDF.js, Canvas, and an OCR Worker on the device; heavy libraries and models load only for the selected workflow.

No document upload

There is no API sending documents to Tool Office servers; results are created as local Blobs.

Verifiable pipeline

Rendering, filtering, cropping, redaction, and OCR expose preview/output for user review.

Safety limits

Up to 30 pages/images, 25 MB per file, and shared canvas limits reduce tab-crash risk.

PDF, scan, and OCR FAQ

Are PDFs and scans uploaded?

No. Files are processed in the browser. Tesseract may download code/language models from a CDN, but it does not send your images to an OCR service.

Which DPI should I use for PDF to image?

Use 150 DPI for light previews, 200 DPI for balance, and 300 DPI for print or small-text OCR. Higher DPI uses more memory and creates larger files.

How are deskew and perspective correction different?

Deskew rotates a few degrees to level text lines. Perspective correction maps four corners to flatten a photographed trapezoid or angled page.

Is blur a safe redaction method?

Do not assume blur is sufficient for sensitive data. A rasterized solid black box is clearer, but always reopen and verify the exported file.

Is Vietnamese OCR perfectly accurate?

No. Accuracy depends on resolution, font, accents, lighting, skew, and scan quality. Always verify names, numbers, dates, and amounts.

How does scanned PDF OCR work?

PDF.js renders each page to Canvas, then a Tesseract Worker recognizes the page image. Tesseract.js does not read PDFs directly.

What is a searchable PDF?

It preserves the page image and adds an OCR text layer for search/selection. The layer can contain the same recognition errors as OCR text.

Does the tool handle password-protected PDFs?

Not reliably. Encrypted PDFs may request a password or fail to render; this tool does not crack or remove PDF passwords.

Related tools