pdf2docx

pdf2docx

Artifex
+
+

Related Products

  • Nutrient SDK
    113 Ratings
    Visit Website
  • Apryse PDF SDK
    172 Ratings
    Visit Website
  • ARGOS Identity
    8 Ratings
    Visit Website
  • Foxit Document Workflow APIs
    8 Ratings
    Visit Website
  • MyQ
    197 Ratings
    Visit Website
  • PackageX OCR Scanning
    48 Ratings
    Visit Website
  • MobiPDF
    7,811 Ratings
    Visit Website
  • MobiOffice
    15,853 Ratings
    Visit Website
  • CirrusPrint
    2 Ratings
    Visit Website
  • Gaffa
    5 Ratings
    Visit Website

About

Docling is an easy-to-use, self-contained, MIT-licensed open source toolkit for converting messy documents into structured data and simplifying downstream document and AI processing. It can parse many popular document formats into a unified and richly structured Docling Document, including PDF, DOCX, PPTX, XLSX, HTML, Markdown, AsciiDoc, CSV, images, audio, and scanned pages through an OCR engine of the user’s choice. Docling detects tables, formulas, reading order, chunks, bounding boxes, page headers and footers, pictures, captions, code, list items, paragraphs, cells, and document structure, making extracted content easier to process, search, and ingest into AI, RAG, and agentic systems. It can export parsed documents to JSON, text, Markdown, HTML, and Doctags, giving developers flexible outputs for pipelines and applications. Docling stores and traverses components according to reading order, partitions documents into bite-sized contiguous text chunks.

About

pdf2docx is a Python library that uses PyMuPDF to extract data from PDF files, parse their layouts according to rules, and generate corresponding .docx files via python-docx. It supports conversion of text, images, tables, and other structural elements; it includes tools to extract tables, handle formatting, and preserve layout as much as possible. It offers both a command-line interface and a graphical user interface. The internal architecture is modular; it includes packages for handling pages, layout, tables, images, shape paths, text spans/blocks, and other elements, enabling fine control over how PDF content is mapped into Word documents. Developers can use the API for batch conversions or integrate it into workflows; there's documentation on installation (from PyPI or source), usage, and technical details of layout-parsing, table extraction, and internal modules. The project is open source, hosted on GitHub, and made available under its license with no warranty.

Platforms Supported

Windows Supported
Mac Supported
Linux Supported
Cloud Not Supported
On-Premises Not Supported
iPhone Not Supported
iPad Not Supported
Android Not Supported
Chromebook Not Supported

Platforms Supported

Windows Not Supported
Mac Not Supported
Linux Not Supported
Cloud Supported
On-Premises Not Supported
iPhone Not Supported
iPad Not Supported
Android Not Supported
Chromebook Not Supported

Audience

AI teams that want to convert complex documents into structured data for document processing, RAG, and agentic applications

Audience

Technical users seeking a solution to convert PDF documents into Word format programmatically while preserving layout, tables, images, and text structure

Support

Phone Support Not Supported
24/7 Live Support Not Supported
Online Supported

Support

Phone Support Supported
24/7 Live Support Not Supported
Online Supported

API

Offers API Not Supported

API

Offers API Supported

Screenshots and Videos

Screenshots and Videos

Pricing

Free
Free Version Supported
Free Trial Not Supported

Pricing

Free
Free Version Supported
Free Trial Not Supported

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation Supported
Webinars Not Supported
Live Online Not Supported
In Person Not Supported

Training

Documentation Supported
Webinars Not Supported
Live Online Not Supported
In Person Supported

Company Information

Docling
United States
www.docling.ai/

Company Information

Artifex
Founded: 1993
United States
pdf2docx.readthedocs.io/en/latest/

Alternatives

DeepSeek-OCR

DeepSeek-OCR

DeepSeek

Alternatives

PDF Conversa

PDF Conversa

ASCOMP Software
Mistral OCR 3

Mistral OCR 3

Mistral AI
AnyParser

AnyParser

CambioML
Mistral OCR 4

Mistral OCR 4

Mistral AI
Unsiloed

Unsiloed

Unsiloed.ai
PDF.co

PDF.co

ByteScout
LlamaParse

LlamaParse

LlamaIndex

Categories

OCR Supported

Categories

PDF Supported

Integrations

Python Supported
GitHub Not Supported
Google Sheets Supported
HTML Supported
JSON Supported
Markdown Supported
Microsoft Excel Supported
Microsoft Word Not Supported
Model Context Protocol (MCP) Supported
PyMuPDF Not Supported
PyPI Not Supported

Integrations

Python Supported
GitHub Supported
Google Sheets Not Supported
HTML Not Supported
JSON Not Supported
Markdown Not Supported
Microsoft Excel Not Supported
Microsoft Word Supported
Model Context Protocol (MCP) Not Supported
PyMuPDF Supported
PyPI Supported
Claim Docling and update features and information
Claim Docling and update features and information
Claim pdf2docx and update features and information
Claim pdf2docx and update features and information