Search Results for "pdf data mining" - Page 2

Showing 148 open source projects for "pdf data mining"

View related business solutions
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    Pysheeet

    Pysheeet

    Python Cheat Sheet

    Pysheeet is a community-driven collection of Python code snippets covering common patterns and tasks like sockets, file I/O, data structures, and more. Each snippet is concise and battle-tested, designed to save coding time and reduce boilerplate. With documentation hosted on Read the Docs and an active GitHub repo, it’s a go-to resource for Python developers.
    Downloads: 4 This Week
    Last Update:
    See Project
  • 2
    Tarjamento de Dados Pessoais e Sigilosos

    Tarjamento de Dados Pessoais e Sigilosos

    Ferramenta de Tarjamento de Dados Pessoais e Sigilosos

    TarjaPDF v2.3 — Ferramenta de Tarjamento de Dados Pessoais e Sigilosos Proteja dados sensíveis em PDFs com segurança irreversível. Interface moderna (tema claro/escuro), marcação manual (texto, linha, área) e detecção automática de CPF, RG, e-mail, telefone, nomes e endereços, inclusive em PDFs digitalizados via OCR integrado, em todas as páginas. Scan inteligente com revisão antes de tarjar e relatório LGPD a cada operação. Security by Design: salva só como PDF-imagem, impedindo a...
    Leader badge
    Downloads: 83 This Week
    Last Update:
    See Project
  • 3
    Hiring Agent

    Hiring Agent

    AI agent to evaluate and score resumes

    Hiring Agent is an AI-powered resume evaluation pipeline for screening technical candidates. It reads a resume PDF and converts the content into Markdown-like text. It then uses a local or hosted language model to extract structured candidate information into sectioned JSON. The system can enrich that resume data with GitHub profile and repository signals when a profile is available. After the data is collected, it produces an explainable evaluation with category scores, supporting evidence, bonus points, and deductions. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    ArchiveBox

    ArchiveBox

    Open source self-hosted web archiving

    ArchiveBox is a powerful, self-hosted internet archiving solution to collect, save, and view websites offline. Without active preservation effort, everything on the internet eventually disappears or degrades. Archive.org does a great job as a centralized service, but saved URLs have to be public, and they can't save every type of content. ArchiveBox is an open source tool that lets organizations & individuals archive both public & private web content while retaining control over their data....
    Downloads: 9 This Week
    Last Update:
    See Project
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • 5
    Mercury

    Mercury

    Convert Python notebook to web app and share with non-technical users

    Turn Python notebooks to web applications with open-source Mercury framework. Hide code and add interactive widgets. Non-technical users can tweak widgets and execute notebook with new parameters. The core of Mercury is Open Source under AGPLv3. We provide Mercury Pro with additional features, dedicated support and friendly commercial license. Mercury is a perfect tool to convert Python notebook to interactive web application and share with non-programmers. You define interactive widgets for...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 6
    LLMStack

    LLMStack

    No-code multi-agent framework to build LLM Agents, workflows

    LLMStack is a no-code platform for building generative AI agents, workflows and chatbots, connecting them to your data and business processes. Build tailor-made generative AI agents, applications and chatbots that cater to your unique needs by chaining multiple LLMs. Seamlessly integrate your own data, internal tools and GPT-powered models without any coding experience using LLMStack's no-code builder. Trigger your AI chains from Slack or Discord. Deploy to the cloud or on-premise.
    Downloads: 10 This Week
    Last Update:
    See Project
  • 7
    gdown

    gdown

    Google Drive public file downloader when curl/wget fails

    ...Users can download individual files, recursively copy folders, choose output paths, or inspect resolved names as JSON. Google Docs, Sheets, and Slides can be exported to formats such as DOCX, XLSX, PPTX, PDF, and CSV. Interrupted transfers can resume, while optional controls limit speed, route traffic through a proxy, adjust certificates, or change cookies and user agents. The tool also downloads regular HTTP and HTTPS URLs and can stream data to standard output. It is available as both a command-line program and a Python package.
    Downloads: 16 This Week
    Last Update:
    See Project
  • 8
    PaperQA2

    PaperQA2

    High accuracy RAG for answering questions from scientific documents

    PaperQA2 is a package for doing high-accuracy retrieval augmented generation (RAG) on PDFs or text files, with a focus on the scientific literature. See our recent 2024 paper to see examples of PaperQA2's superhuman performance in scientific tasks like question answering, summarization, and contradiction detection. In this example we take a folder of research paper PDFs, magically get their metadata - including citation counts and a retraction check, then parse and cache PDFs into a...
    Downloads: 11 This Week
    Last Update:
    See Project
  • 9
    PdfBooklet
    PdfBooklet is a Python Gtk application which allows to make books or booklets from existing pdf files. It can also adjust margins, rotate, scale, merge files or extract pages.
    Leader badge
    Downloads: 187 This Week
    Last Update:
    See Project
  • Train ML Models With SQL You Already Know Icon
    Train ML Models With SQL You Already Know

    BigQuery automates data prep, analysis, and predictions with built-in AI assistance.

    Build and deploy ML models using familiar SQL. Automate data prep with built-in Gemini. Query 1 TB and store 10 GB free monthly.
    Start Free
  • 10
    NeMo Retriever Library

    NeMo Retriever Library

    Document content and metadata extraction microservice

    NeMo Retriever Library is a scalable microservice framework designed for extracting, structuring, and enriching content from documents to support downstream generative AI applications. It processes various document types by splitting them into components such as text, tables, charts, and images, and then applies OCR and contextual analysis to convert them into structured data formats. The system is built on NVIDIA NIM microservices, enabling high-performance parallel processing and efficient...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    DocsGPT

    DocsGPT

    Private AI platform for agents, enterprise search and RAG pipelines

    DocsGPT is an open-source AI platform for deploying private RAG pipelines, AI agents, and enterprise search on your own infrastructure. Connect any data source (PDFs, DOCX, CSV, Excel, HTML, audio, GitHub, databases, URLs) and get accurate, hallucination-free answers with source citations. Choose your LLM: OpenAI, Anthropic, Google Gemini, or local models. Works with Qdrant, MongoDB, and Elasticsearch and more. Deploy via Docker or Kubernetes with full data sovereignty. Build...
    Downloads: 7 This Week
    Last Update:
    See Project
  • 12
    shuyuan

    shuyuan

    Reading book source

    shuyuan is a project oriented around reading and knowledge consumption, especially targeting large-scale text content such as books, articles, or educational material. The name suggests “academy” or “study hall,” and the tool aims to help users ingest, organize, and manage reading content — possibly offering features like text parsing, annotation, metadata generation, translation, or storage for later reference. The repository is set up to support document ingestion, indexing, and maybe some...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    text-extract-api

    text-extract-api

    Document (PDF, Word, PPTX ...) extraction and parse API

    text-extract-api is an open-source service designed to extract readable text from a wide variety of document formats through a simple API interface. The project focuses on converting complex files such as PDFs, images, scanned documents, and office files into structured plain text that can be processed by downstream applications or language models. Instead of requiring developers to integrate multiple document parsing libraries individually, the system centralizes text extraction...
    Downloads: 6 This Week
    Last Update:
    See Project
  • 14
    Khoj

    Khoj

    An AI personal assistant for your digital brain

    Get more done with your open-source AI personal assistant. Khoj is a desktop application to search and chat with your notes, documents, and images. It is an offline-first, open-source AI personal assistant that is accessible from Emacs, Obsidian or your Web browser. Khoj is a thinking tool that is transparent, fun, and easy to engage with. You can build faster and better by using Khoj to search and reason across all your data sources. Khoj learns from your notes and documents to function as...
    Downloads: 8 This Week
    Last Update:
    See Project
  • 15
    Jina

    Jina

    Build cross-modal and multimodal applications on the cloud

    ...Jina handles the infrastructure complexity, making advanced solution engineering and cloud-native technologies accessible to every developer. Build applications that deliver fresh insights from multiple data types such as text, image, audio, video, 3D mesh, PDF with Jina AI’s DocArray. Polyglot gateway that supports gRPC, Websockets, HTTP, GraphQL protocols with TLS. Intuitive design pattern for high-performance microservices. Seamless Docker container integration: sharing, exploring, sandboxing, versioning and dependency control via Jina Hub. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    CyberPPT

    CyberPPT

    A Codex Skill for generating high-density, editable PowerPoints

    CyberPPT is a Codex Skill for converting documents, research, and business data into dense, editable, consulting-style PowerPoint presentations. It extracts facts, numbers, claims, recommendations, conflicts, and caveats from formats such as PDF, DOCX, TXT, and XLSX. The workflow builds an MBB-style evidence table before developing storylines and converging on an SCR narrative. Users select from eight fixed visual systems, after which the skill plans each slide’s hierarchy, grid, charts, palette, and information density. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    deepdoctection

    deepdoctection

    A Repo For Document AI

    DeepDoctection is a document AI framework that applies deep learning techniques to analyze and extract structured data from scanned documents, PDFs, and images. deepdoctection is a Python library that orchestrates document extraction and document layout analysis tasks using deep learning models. It does not implement models but enables you to build pipelines using highly acknowledged libraries for object detection, OCR and selected NLP tasks and provides an integrated frameworks for...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Former.skill

    Former.skill

    AI skill generator designed to reconstruct conversational personality

    Ex Skill is an AI skill generator designed to reconstruct the conversational personality and shared memories of a former partner from personal data. It can analyze exported chats, photos, social media records, PDFs, images, and manually entered descriptions. The system builds separate memory and persona layers covering relationship history, habits, emotional patterns, communication style, and behavior. Generated personas can respond in a style modeled after the source material. New...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    QAnything

    QAnything

    Question and Answer based on Anything

    QAnything is a local knowledge-base question-answering system designed to let users ask questions over many kinds of files and databases. It supports offline installation, making it useful for organizations that need private document analysis without sending data to external services. Users can upload local files and receive fast, reliable answers based on the indexed content. The system supports formats such as PDF, Word, PowerPoint, Excel, Markdown, email, text, images, CSV, and web links. Its retrieval process uses a two-stage vector and reranking approach to maintain answer quality as the knowledge base grows. ...
    Downloads: 9 This Week
    Last Update:
    See Project
  • 20
    Perf Book

    Perf Book

    The book "Performance Analysis and Tuning on Modern CPU"

    This project is a practical guide to performance analysis and tuning on modern CPUs, bridging microarchitecture details with hands-on profiling. It explains how caches, TLBs, prefetchers, branch predictors, and out-of-order execution influence real program speed, then connects those concepts to concrete optimization strategies. Readers learn how to design trustworthy benchmarks, avoid measurement traps (warmup, turbo, frequency scaling), and interpret hardware performance counters. The book...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Mova PDF voice reader

    Mova PDF voice reader

    PDF Reader with Text-to-Speech (TTS) Mowa

    Mova is an accessible PDF reader for Linux that combines document viewing with text-to-speech functionality. It displays documents in a convenient two-page layout and reads the extracted text aloud while highlighting the currently spoken line. Users can start reading from the beginning of the displayed pages or click a specific line to continue from that position. Speech language and reading speed can be adjusted directly from the toolbar. Mova automatically saves reading progress for...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 22

    Create Index from PDF

    PDF Indexing Script: Searches PDF for words, records page numbers

    This Python script helps automate the process of creating an index for a PDF document. It reads a list of words from a text file, searches through each page of the PDF, and records the page numbers where each word appears. The script accounts for the first 24 pages of the PDF that use Roman numerals (i-xxiv) and adjusts the page numbers accordingly. It is designed to be case-insensitive, ensuring that variations in capitalization do not affect the search results. As it processes the PDF, the...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23

    FreeSEM

    Free and open-source desktop application designed for SEM

    ...The software supports methods such as exploratory factor analysis, covariance-based SEM, partial least squares SEM, and meta-SEM, and it provides model fit statistics like CFI, TLI, RMSEA, SRMR, and chi-square to evaluate models. It also enables exporting analysis results and reports to formats like Word, Excel, CSV, and PDF, making it useful for academic research and data analysis workflows.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 24

    realwatermark

    A Python application to add watermarks (text or image) to PDF files

    A Python application to add watermarks (text or image) to PDF files, converts them into image and back to PDF with options for OCR and compression.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 25
    PII-Blackout

    PII-Blackout

    100% offline, AI-powered PDF redaction

    ...PII Blackout automatically scans, detects, and blackouts sensitive data points across your documents in one click. Absolute, Irreversible Security (Image-Level Blackout) Unlike standard PDF editors that merely place a black shape over editable text (which can easily be copied or uncovered), PII Blackout flattens and bakes the redaction directly into the image surface of the document. The covered data is permanently destroyed and mathematically impossible to recover. 100% Offline & Local Processing
    Downloads: 3 This Week
    Last Update:
    See Project