Showing 74 open source projects for "scraper"

View related business solutions
  • Veeam Data Platform v13.1 - Get Your Free Trial Icon
    Veeam Data Platform v13.1 - Get Your Free Trial

    Secure by design, portable by default. Recover clean, fast, anywhere. Start a free trial.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try it Free
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 1
    scanless

    scanless

    Online port scan scraper

    scanless is a Python command-line utility and library for running port scans through third-party online scanning services. It is built for users who want quick visibility into exposed ports without running a local scanner directly from their own machine. The tool can scan an IP address or domain, choose a specific supported scanner, select a random scanner, or run all available scanners. It supports services such as ipfingerprints, spiderip, standingtech, viewdns, and yougetsignal. scanless...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Tholian Stealth

    Tholian Stealth

    Secure, Peer-to-Peer, Private and Automateable Web Browser

    ...It aims to prioritize user privacy and autonomy by minimizing tracking, blocking unnecessary requests, and restricting potentially harmful web technologies such as JavaScript execution. The platform operates as both a browser and a network service, capable of acting as a proxy, scraper, and content filtering system for other applications. Stealth introduces peer-to-peer networking features that allow users to share cached content and reduce bandwidth usage, which can be especially beneficial in constrained environments. It also includes advanced content filtering and optimization mechanisms that strip unwanted or malicious elements from web pages before rendering.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    Mangal 4

    Mangal 4

    The most advanced (yet simple) cli manga downloader

    The most advanced CLI manga downloader in the entire universe.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    AutoScraper

    AutoScraper

    A Smart, Automatic, Fast and Lightweight Web Scraper for Python

    This project is made for automatic web scraping to make scraping easy. It gets a URL or the HTML content of a web page and a list of sample data that we want to scrape from that page. This data can be text, URL or any HTML tag value of that page. It learns the scraping rules and returns similar elements. Then you can use this learned object with new URLs to get similar content or the exact same element of those new pages.
    Downloads: 0 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 5
    mlscraper

    mlscraper

    ML-based HTML scraper that learns extraction rules from examples

    ...It analyzes those examples within the HTML document and determines patterns or rules that can be used to extract the same type of information from similar pages. Once trained, the generated scraper can process new pages and return the extracted data in structured formats such as dictionaries or lists. This approach simplifies web scraping tasks by shifting the focus from rule-writing to example-based training. Internally, the project processes HTML documents, identifies relevant elements in the DOM, and builds extraction logic based on statistical or heuristic analysis of the training samples. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    NSFW Data Scraper

    NSFW Data Scraper

    Collection of scripts to aggregate image data

    NSFW Data Scraper is an open-source project that provides scripts for automatically collecting large datasets of images intended for training NSFW image classification systems. The repository focuses on aggregating image data from various online sources so that developers can build datasets suitable for training content moderation models. These datasets typically contain images categorized into different classes associated with adult or explicit content, which can then be used to train neural networks that detect unsafe or inappropriate material. ...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 7
    SecretAgent

    SecretAgent

    The web scraper that's nearly impossible to block

    SecretAgent is a headless browser that’s nearly impossible to detect. It achieves this by emulating real users. And it has powerful auto-replay functionality that lets you create and debug scripts in record setting time.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8

    scraper-helper

    A HTTP proxy that logs everything flowing through it

    ...It works with HTTPS, which means it performs a man in the middle attack SSL do it can decode all encrypted connections as well. It can create the X509 CA certificate needed to perform the MITM attack. All available documentation can be read online at http://scraper-helper.sourceforge.net/
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    NYT Vote Scraper

    NYT Vote Scraper

    Scrapes the NYT Votes Remaining Page JSON

    NYT Vote Scraper is a small but clever project that periodically fetches JSON data from the “Votes Remaining” page of The New York Times during the 2020 U.S. presidential election and commits the results into the repository, effectively using Git as a time-series database. The idea is to create a historical record — including diffs — of how vote counts and “votes remaining” estimates changed over time.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Start Free
  • 10
    X-RAY

    X-RAY

    The next web scraper, see through the <html> noise

    Supports strings, arrays, arrays of objects, and nested object structures. The schema is not tied to the structure of the page you're scraping, allowing you to pull the data in the structure of your choosing. The API is entirely composable, giving you great flexibility in how you scrape each page. Paginate through websites, scraping each page. X-ray also supports a request delay and a pagination limit. Scraped pages can be streamed to a file, so if there's an error on one page, you won't...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    JonDoFox Advanced Privacy Browser

    JonDoFox Advanced Privacy Browser

    Browser with fingerprinting- and psychological profiling protection

    ...This can't be reached with common addons, but our Browser provides it. In the last line of defense, we do provide fluctuating IPs with proxies Update v2.0: -Based on Firefox 70.0 Beta (15.09.2019) -Added proxy scraper/checker/configurator Note: currently, facebook chat is broken. To increase security, use a hosts file black list in adition to your adblocker, like this one: https://github.com/StevenBlack/hosts
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    django-dynamic-scraper

    django-dynamic-scraper

    Creating Scrapy scrapers via the Django admin interface

    Django Dynamic Scraper (DDS) is an app for Django build on top of the scraping framework Scrapy. While preserving many of the features of Scrapy it lets you dynamically create and manage spiders via the Django admin interface. With Django Dynamic Scraper (DDS) you can define your Scrapy scrapers dynamically via the Django admin interface and save your scraped items in the database you defined for your Django project.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    google-play-scraper

    google-play-scraper

    Node.js scraper to get data from Google Play

    Node.js module to scrape application data from the Google Play store. Retrieves the full detail of an application. Retrieves a list of applications from one of the collections at Google Play. Retrieves a list of apps that results of searching by the given term. Returns the list of applications by the given developer name. Given a string returns up to five suggestions to complete a search query term. Retrieves a page of reviews for a specific application. Returns a list of similar apps to the...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14

    WebExtractServer

    WebExtractServer use with WebExtractLte for use with web browsers

    Browse data, fetched by WebExtractLte directly in your browser. Designed to be used with Webscraper (webscraper.io) - third party web scraper tool, available as plugin for Chrome and Firefox.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    mzitu

    mzitu

    Python crawler that downloads image galleries and analyzes titles

    mzitu is a Python-based web crawling project designed to automatically download and organize image galleries from a specific photography site. It demonstrates how to build a scraper that navigates gallery pages, retrieves image links, and saves the images locally in a structured directory layout. It focuses on automating the collection of large sets of images by programmatically parsing page content and iterating through gallery entries. mzitu also includes a simple analysis script that processes downloaded folder names to generate statistics and visualizations. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 16
    ProxyGenerator

    ProxyGenerator

    Proxy Generator is a multi-functional Proxy Grabber and Checker

    Proxy Generator is a multi-functional Programm for Proxys Features: Proxy Grabber Proxy Scraper Proxy Checker
    Downloads: 1 This Week
    Last Update:
    See Project
  • 17
    Squirrel Album Post Scraper

    Squirrel Album Post Scraper

    download photos albums from Facebook

    Simple tool to download photos of a specific album on Facebook. with just one click!.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18

    HttpScraper

    A simple http scraper for my own amusment

    A simple http scraper for my own amusment
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Products of the project: Java HTMLParser - VietSpider Web Data Extractor - Extractor VietSpider News. Click on "Show project details" to see more feature about each product.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20

    My Twitter Scraper

    Filters Twitter in real time against your keywords and outputs to .csv

    Program allows users to filter main Twitter stream against specified keywords. Output is shown on screen, and when finished, a csv file is created containing all of the captured tweets, usernames, times, and location. Is built to run for extended periods with minimal growth in RAM usage as number of captured tweets increases. Uses both twitter4j and Open-CSV. YOU MUST HAVE JAVA 1.8 INSTALLED, and you will need to get your own OAuth tokens from Twitter. You can get them here:...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21

    ScraperEdit for XBMC

    XML bindings and a GUI for creating and editing XBMC Scrapers

    This program is an editor for creating XBMC Scrapers. It is similar to ScraperEditor, an other editor using ScraperXML, that runs under .Net environment. This program runs under Sun/Oracle's Java Runtime. HELP WANTED! I am looking for someone, who would help me writing documentation, like user's manual and on-line help. Also if someone want to help, translated language files are always welcome...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22

    Scra.php

    Scrape anything!

    The ultimate customiseable YAML-ised Web Scraper for PHP
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23

    GatherProxy

    Free Proxy & Socks Scraper

    Gather Proxy is a lightweight Windows utility designed to help users gather information about proxy servers and socks. Since this is a portable program, it is important to mention that it doesn’t leave any traces in the Windows Registry. You can copy it on any USB flash drive or other devices, and take it with you whenever you to need to generate proxy and socks lists on the breeze. Although it comes bundled with many useful functions, it boasts a clean and straightforward...
    Downloads: 6 This Week
    Last Update:
    See Project
  • 24
    MuhVieh - Filmverwaltung

    MuhVieh - Filmverwaltung

    Ein Skript zur Verwaltung der persönlichen Filmsammlung.

    Das Skript stellt eine Filmdatenbank zur Verfügung. Des Weiteren beinhaltet es die Aufschlüsselung nach Genres, eine Benutzerverwaltung und eine ansprechende Präsentation der Inhalte.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25

    PlusOne Scraper

    Generate an RSS feed of PlusOnes published to your Google Profile

    PlusOne Scraper generates an RSS feed of a user's PlusOnes (+1s) from a Google profile page. Designed to be self-hosted. Built with PHP and PHP Simple HTML DOM Paser. Code also taken from sgthayes.
    Downloads: 0 This Week
    Last Update:
    See Project