Showing 726 open source projects for "linux proxy scraper"

View related business solutions
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    camofox-browser

    camofox-browser

    Headless browser automation server for AI agents to visit sites

    camofox-browser is a headless browser automation server built specifically for AI agents that need to interact with websites that often block standard automation stacks. It wraps Camoufox, a Firefox fork that performs fingerprint spoofing at the C++ level, which means many browser characteristics are altered before page scripts can inspect them, rather than relying on JavaScript-layer stealth patches. The project is designed around a REST API, making it easier for agents and external tools...
    Downloads: 23 This Week
    Last Update:
    See Project
  • 2
    gliderlabs/ssh

    gliderlabs/ssh

    Easy SSH servers in Golang

    Gliderlabs SSH is a Go package that simplifies the creation of SSH servers. It provides a high-level API for building custom SSH server applications, abstracting the complexities of the underlying SSH protocol.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    POI POI POI

    POI POI POI

    Scalable KanColle browser and tool

    poi is a scalable browser and tool set for Kantai Collection(KanColle). Key features include proxy,HTTP, Socks5 and PAC (Experimental). Cache, including custom cache. Data synthesis and analysis. Notification and plugin support for extensive functionalities.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    Spider

    Spider

    High-performance Rust web crawler and scraper for large-scale data

    Spider is a high-performance web crawler and web scraping library written in Rust that enables developers to crawl and index websites efficiently. It focuses on speed, concurrency, and reliability by using asynchronous and multi-threaded processing to handle large volumes of web pages. It can rapidly crawl websites to collect links, retrieve page content, and extract structured information from HTML documents. Spider can operate concurrently across many pages, allowing it to gather large...
    Downloads: 20 This Week
    Last Update:
    See Project
  • Build Data Resilience - Take the Assessment Today Icon
    Build Data Resilience - Take the Assessment Today

    Can you recover when it matters most? Take this quick assessment to identify gaps and build greater recovery confidence.

    Is your recovery strategy as strong as you think? Take this quick self-assessment to check your recovery readiness and gain tailored insights. In only 2 minutes, you'll learn where you fall on the recovery readiness scale.
    Take the Assessment
  • 5
    ScrapeGraphAI

    ScrapeGraphAI

    Python scraper based on AI

    Extracting content from websites and local documents using LLM. ScrapeGraphAI is a web scraping python library that uses LLM and direct graph logic to create scraping pipelines for websites and local documents (XML, HTML, JSON, Markdown, etc.). Just say which information you want to extract and the library will do it for you.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    ArgoX

    ArgoX

    Argo Xray for VPS one-click script

    ArgoX is a shell-based deployment project for setting up Xray proxy services with Cloudflare Argo Tunnel support on VPS environments. It focuses on one-click installation and interactive configuration, making it easier to deploy a multi-protocol proxy stack without manually assembling every component. The project supports several protocol options, including VLESS, VMess, Trojan, Shadowsocks, Hysteria2, Reality, WebSocket, gRPC, and XHTTP-based modes. ArgoX uses Cloudflare tunnel behavior to...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    rnet

    rnet

    Python HTTP client with TLS and HTTP/2 fingerprint emulation support

    rnet is an ergonomic and modular Python HTTP client designed for developers who need advanced control over network requests and protocol behavior. It provides a flexible API for making HTTP requests while supporting both asynchronous and blocking workflows, allowing it to integrate easily into different Python applications and runtimes. rnet focuses on low-level protocol customization, giving users fine-grained control over TLS and HTTP/2 configuration in order to emulate specific browser...
    Downloads: 15 This Week
    Last Update:
    See Project
  • 8
    Ulixee Hero

    Ulixee Hero

    The web browser built for scraping

    It's the first modern headless browsers designed specifically for scraping instead of just automated testing. Hero provides access to the W3C DOM specification without the need for Puppeteer's complicated evaluate callbacks and multi-context switching. We've recreated a fully compliant DOM directly in NodeJS allowing you bypass the headaches of previous scraper tools. The powerful Chrome engine sits under the hood, allowing for lightning fast rendering. Emulators make it easy to disguise...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 9
    MDCx

    MDCx

    Movie metadata scraper and organizer for media libraries and NFO

    MDCx is an open source media metadata scraping and organization tool designed to automate the process of collecting detailed information for movie files. It retrieves metadata from multiple online sources and applies it to local media collections, helping users maintain structured and well-organized libraries. MDCx can download information such as titles, cast data, artwork, and other metadata, then generate standardized NFO files compatible with media management systems. It also supports...
    Downloads: 9 This Week
    Last Update:
    See Project
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Start Free
  • 10
    NextDNS

    NextDNS

    NextDNS CLI client (DoH Proxy)

    NextDNS protects you from all kinds of security threats, blocks ads and trackers on websites and in apps and provides a safe and supervised Internet for kids, on all devices and on all networks. Determine your threat model and fine-tune your security strategy by enabling 10+ different types of protections. Use the most trusted threat intelligence feeds containing millions of malicious domains, all updated in real-time. Go beyond the domain, we analyze DNS questions and answers on-the-fly (in...
    Downloads: 14 This Week
    Last Update:
    See Project
  • 11
    KubeVPN

    KubeVPN

    KubeVPN offers a Cloud Native Dev Environment

    KubeVPN is a cloud-native development tool that connects a local machine to a Kubernetes cluster network. It lets developers access cluster resources directly through service names, Pod IPs, or Service IPs, which can make local debugging feel closer to working inside the cluster. The project supports traffic interception so inbound requests for a remote Kubernetes workload can be redirected to a local development process. It can also run a Kubernetes pod locally in Docker while preserving...
    Downloads: 5 This Week
    Last Update:
    See Project
  • 12
    crwlr

    crwlr

    Library for Rapid (Web) Crawler and Scraper Development

    This library provides kind of a framework and a lot of ready-to-use, so-called steps, that you can use as building blocks, to build your own crawlers and scrapers with. Before diving into the library, let's have a look at the terms crawling and scraping. For most real-world use cases, those two things go hand in hand, which is why this library helps with and combines both. A (web) crawler is a program that (down)loads documents and follows the links in it to load them as well. A crawler...
    Downloads: 6 This Week
    Last Update:
    See Project
  • 13
    scrape-it

    scrape-it

    A Node.js scraper for humans

    scrape-it is a Node.js library for converting web pages and HTML documents into structured JavaScript objects. Developers describe the desired output through a declarative schema based on CSS selectors. Fields can extract text, raw HTML, element attributes, repeated items, and nested lists. Conversion functions can transform captured values into dates, numbers, or other application-specific types. The library supports promises, async and await workflows, and callback-based usage. It can...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    tumblr-crawler

    tumblr-crawler

    Python crawler to download photos and videos from Tumblr blogs

    tumblr-crawler is an open source Python-based utility designed to download media content from Tumblr blogs. It provides a script that automatically retrieves photos and videos from specified Tumblr sites and saves them locally for offline access. Users can specify one or multiple blogs to crawl by editing a configuration file or by passing parameters through the command line. Once executed, the script fetches media from the Tumblr API and stores the downloaded files in folders named after...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Pomerium

    Pomerium

    Pomerium is an identity and context-aware access proxy

    Secure, context-aware access that just works. Access internal resources securely. Implement zero trust. Achieve compliance. All without the headache of a VPN. For teams that prefer a hosted solution while keeping data governance. For organizations that need advanced scaling, access control, and governance capabilities. IT and developers need a scalable access control solution to keep users productive, happy, and secure. Pomerium uses identity and context to ensure secure access to internal...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    wireguard-install

    wireguard-install

    WireGuard road warrior installer for Ubuntu, Debian, AlmaLinux

    WireGuard Install is an interactive Bash installer for quickly deploying a road-warrior WireGuard VPN server on Linux. It validates the operating system, privileges, networking environment, and kernel requirements before beginning configuration. The assistant selects server addresses, ports, DNS settings, and an initial client through a guided terminal workflow. It installs WireGuard packages, enables packet forwarding, and creates the firewall and routing rules needed for tunneled traffic....
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    autocrawler

    autocrawler

    Multiprocess Selenium crawler for downloading images by keywords

    AutoCrawler is a Python-based image crawling tool designed to automatically download large numbers of images from search engines using automated browser interaction. It uses Selenium and a Chrome browser driver to navigate image search pages and collect image sources based on keywords provided by the user. AutoCrawler supports multiprocess and multithreaded downloading, which allows it to retrieve images faster by running several tasks simultaneously. Users provide search terms through a...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 18
    paqet

    paqet

    Ferries Packets Across Forbidden Boundaries

    paqet is a bidirectional packet-level proxy that transports traffic over raw packets instead of relying on the host operating system’s normal TCP/IP stack. It forwards traffic between a local client and a remote server using KCP as the reliable transport layer. The project is written in Go and is aimed at users who need a low-level tunnel that can operate in restricted or heavily filtered network environments. It is described as an alpha-stage tool, so compatibility and protocol behavior may...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Blocky

    Blocky

    Fast and lightweight DNS proxy as ad-blocker for local network

    Blocky is a fast, lightweight DNS proxy and network-wide ad blocker designed for home labs and small networks that want Pi-hole-like filtering with more flexible DNS routing and modern protocol support. It blocks DNS queries using external deny lists while also supporting allow lists, and it can scope policies by client groups so different devices or households can have different rulesets. Unlike a single-purpose blocker, it supports advanced DNS behavior such as conditional forwarding and...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 20
    HTTPLeaks

    HTTPLeaks

    All possible ways, a website can leak HTTP requests

    HTTPLeaks is a security test project that catalogs ways HTML can cause a browser or renderer to send external HTTP requests. Its primary artifact is a single HTML file containing many combinations of elements and attributes that may trigger network activity. Researchers can use it to evaluate Content Security Policy enforcement, webmail sanitization, proxy rewriting, and privacy controls. The project highlights cases where opening content may unexpectedly disclose that it was viewed or...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    WIREDOOR

    WIREDOOR

    Self hosted ingress-as-a-service platform

    WIREDOOR is a self-hosted ingress-as-a-service platform for exposing private, local, or firewalled services to the internet. It uses secure WireGuard tunnels and reverse proxy behavior so internal applications can be reached without directly opening the private network. The platform is designed for users who want cloud-like ingress management while keeping the infrastructure under their own control. Wiredoor can expose HTTP and TCP services and is positioned as a secure way to publish...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    piko

    piko

    An open-source alternative to Ngrok, designed to serve production

    Piko is an open-source reverse proxy and tunneling system designed as a self-hostable alternative to ngrok. It lets upstream services open outbound-only connections to Piko, then routes incoming traffic back through those established tunnels. This design is useful for services that are private, behind NAT, firewalled, or otherwise not publicly routable. Piko is built for production-oriented hosting rather than only short-lived development demos. It supports a server and agent model, with the...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Cntlm is an NTLM / NTLMv2 authenticating HTTP/1.1 proxy. It caches auth'd connections for reuse, offers TCP/IP tunneling (port forwarding) thru parent proxy and much much more. It's in C, very fast and resource-efficient. Go to http://cntlm.sf.net/
    Leader badge
    Downloads: 303 This Week
    Last Update:
    See Project
  • 24
    Microsoft Azure CLI

    Microsoft Azure CLI

    Azure command-line interface

    A great cloud needs great tools; we're excited to introduce Azure CLI, our next-generation multi-platform command-line experience for Azure. Take a test run now from Azure Cloud Shell! We support tab completion for groups, commands, and some parameters. You can use the --query parameter and the JMESPath query syntax to customize your output. With the Azure CLI Tools Visual Studio Code extension, you can create .azcli files and use these features. IntelliSense for commands and their...
    Downloads: 50 This Week
    Last Update:
    See Project
  • 25
    gain

    gain

    Asyncio-based Python framework for building fast web crawling spiders

    Gain is a Python web crawling framework designed to simplify the process of building efficient and scalable web scrapers. It is built on top of asynchronous technologies such as asyncio, aiohttp, and uvloop to support high-performance crawling with concurrent network requests. It provides a structured framework for creating spiders that can navigate websites, extract structured data, and process the collected results. Developers define crawlers using components such as spiders, parsers, and...
    Downloads: 1 This Week
    Last Update:
    See Project