pdfRest
pdfRest is a developer PDF API with 40+ endpoints for conversion, extraction, and document manipulation. Convert to and from PDF (Word, Excel, PowerPoint, images, PDF/A, PDF/X). Merge, split, compress, and linearize. Extract text, images, and metadata. Add OCR to scanned documents. Watermark, encrypt, decrypt, restrict, and redact. Fill, flatten, import, and export form data. Requests are chainable, so multi-step document pipelines run in sequence without shuttling files back and forth between calls.
What sets pdfRest apart in the PDF API space is deployment flexibility. Most PDF APIs are cloud-only. pdfRest ships three ways: Cloud, Container and AWS Marketplace.
For teams running an evaluation: pdfRest is SOC 2 Type 2 audited and supports GDPR and HIPAA requirements. The Container deployment is the strongest fit for regulated workloads, specifically because processing happens entirely inside your own environment.
Learn more
Filestar
Do anything to any file. Tens of thousands of skills at your fingertips. Quickly convert files in a few clicks. Choose from over 30 000 file conversions. Both common and unusual file formats. Single files or in bulk. Easily merge one or many files at once. Combine files for many different file types. Merge documents, video, audio, Visio or other file formats. Split large files with many pages into several separate ones. For text file formats like .pdf, .doc and .txt. Divide files and documents into parts. Change or alter files. Rotate, add filters, replace file names, add watermarks, add text to images, and much more. One at a time or many at once. Simply compress or reduce the file size of your files. Wide selection of file compression formats and zip options to choose from. Smoothly extract selected pages or elements from a document. Collect images out of a file, or get all images or text from a document.
Learn more
PDF Constructor
Using an XML grammar incorporating features of XHTML, CSS, and SVG, PDF Constructor creates single or multiple-page PDF documents using existing or dynamically-created raster, vector, and text content. Build PDFs with content that is ready to go to print. Use CMYK and spot colors. Specify the bleed and trim. Use Type 1, TrueType, or OpenType fonts, always embedded and optionally subset. Produce web or screen-ready documents with bookmarks, hyperlinks, actions, and JavaScript. You can even build complete Acrobat Forms dynamically. Include JPEG and TIFF images in any colorspace and resolution. Apply your choice of transformations to ensure the image fits correctly into your layout. Include SVG drawings directly or by reference. Specify individual pages or entire PDF documents as new content or as a template on which to add new elements. Paragraph and character styles based on CSS2 can be specified for flowable content.
Learn more
PDFBox
The Apache PDFBox® library is an open-source Java tool for working with PDF documents. This project allows the creation of new PDF documents, manipulation of existing documents and the ability to extract content from documents. Apache PDFBox also includes several command-line utilities. Apache PDFBox is published under the Apache License v2.0. Extract Unicode text from PDF files. Split a single PDF into many files or merge multiple PDF files. Extract data from PDF forms or fill a PDF form. Validate PDF files against the PDF/A-1b standard. Print a PDF file using the standard Java printing API. Create a PDF from scratch, with embedded fonts and images. Save PDFs as image files, such as PNG or JPEG and digitally sign PDF files. See also the export control information related to the encryption features included in Apache PDFBox.
Learn more