Showing 126 open source projects for "data collection algorithm"

View related business solutions
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Try It Free
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Try Free
  • 1
    skycaiji

    skycaiji

    Open source web scraping system for automated data collection tasks

    SkyCaiji is an open source web scraping and data collection system designed to gather information from websites through configurable extraction rules. It focuses on simplifying the process of building crawlers by allowing users to visually define scraping rules rather than writing complex code. It can collect structured or unstructured data from many types of webpages and automate the extraction process for large datasets.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Fluentd

    Fluentd

    Fluentd: Unified Logging Layer (project under CNCF)

    Fluentd is a CNCF‑graduated open-source data collector that unifies log data collection and consumption across diverse systems. It supports robust reliability, buffering, extensible plugin architecture, and real-time log routing. Fluentd serves as a unified logging layer for structured/unstructured data processing.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    douyin

    douyin

    Open source Douyin crawler for collecting and downloading public data

    DouyinCrawler is an open source data collection tool designed to gather publicly available information from the Douyin platform. It demonstrates how to build a Python-based web crawler combined with a graphical interface and command line functionality. It allows users to collect data from various types of Douyin content, including user profiles, videos, hashtags, and music pages.
    Downloads: 9 This Week
    Last Update:
    See Project
  • 4
    Immutable.js

    Immutable.js

    Immutable collections for JavaScript

    Immutable.js offers a collection of Persistent Immutable data structures for JavaScript. Immutable data is unchangeable once created, which makes application development so much simpler. There’s no defensive copying, and you get advanced memoization and change detection techniques with simple logic. Persistent data gives you a mutative API, one that doesn’t update data in-place but always produces new and updated data.
    Downloads: 0 This Week
    Last Update:
    See Project
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New to Google Cloud? Get $300 in credits to explore Compute Engine, BigQuery, Cloud Run, Gemini Enterprise Agent Platform, and more.

    Start your next project with $300 in free Google Cloud credit. Spin up VMs, run containers, query petabytes in BigQuery, or build agents with Gemini Enterprise Agent Platform. Once your credits are used, keep building with 20+ always-free tier products including Compute Engine, Cloud Storage, GKE, and Cloud Run functions. No commitment required—just sign up and start building.
    Claim $300 Free
  • 5
    POI POI POI

    POI POI POI

    Scalable KanColle browser and tool

    poi is a scalable browser and tool set for Kantai Collection(KanColle). Key features include proxy,HTTP, Socks5 and PAC (Experimental). Cache, including custom cache. Data synthesis and analysis. Notification and plugin support for extensive functionalities.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    Anna’s Archive

    Anna’s Archive

    Comprehensive search engine for books, papers, comics, magazines

    Anna’s Archive is a large-scale open-source search engine and data aggregation platform designed to index and provide access to a vast collection of books, academic papers, comics, magazines, and other digital texts through a unified interface. The project includes all the infrastructure required to run a full instance locally or in production, combining web servers, databases, and search indexing systems into a scalable architecture.
    Downloads: 20 This Week
    Last Update:
    See Project
  • 7
    Helium Browser

    Helium Browser

    Private, fast, and honest web browser

    ...Helium blocks ads and trackers by default through an integrated, unbiased uBlock Origin extension prepackaged as a native browser component. Its UI and feature set emphasize minimalism, no “smart” recommendations, account sync, or background data collection, resulting in a distraction-free browsing experience that respects user autonomy. The browser is available across macOS, Linux, and Windows, each version built from a fully open source pipeline for reproducibility and trust. Development focuses on maintaining compatibility with modern web standards while decoupling Chromium from its Google dependencies and services.
    Downloads: 123 This Week
    Last Update:
    See Project
  • 8
    FEAPDER

    FEAPDER

    Powerful Python crawler framework for scalable web scraping tasks

    feapder is a Python-based web crawling framework designed to simplify the process of building scalable and efficient web scrapers. It focuses on providing a developer-friendly environment that makes it easier to create, run, and manage crawlers for a variety of data collection tasks. It includes several built-in spider types, such as AirSpider, Spider, TaskSpider, and BatchSpider, which address different crawling scenarios ranging from lightweight scraping to distributed and batch-based jobs. feapder supports features such as breakpoint resume, allowing crawlers to continue from where they stopped without losing progress. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Weibo Crawler

    Weibo Crawler

    Python crawler for collecting and downloading Sina Weibo user data

    weibo-crawler is a Python-based data collection tool designed to retrieve information from Sina Weibo user accounts. It automates the process of gathering posts, user profile details, and engagement metrics from one or more target accounts. weibo-crawler can extract comprehensive information about users, including profile attributes such as nickname, follower count, following count, and account metadata.
    Downloads: 5 This Week
    Last Update:
    See Project
  • Earn up to 16% annual interest with Nexo. Icon
    Earn up to 16% annual interest with Nexo.

    Let your crypto work for you

    Put idle assets to work with competitive interest rates, borrow without selling, and trade with precision. All in one platform. Geographic restrictions, eligibility, and terms apply.
    Get started with Nexo.
  • 10
    crawler

    crawler

    Collection of JS reverse engineering examples for web scraping study

    crawler is a collection of web scraping and JavaScript reverse engineering examples designed for learning how modern websites protect their data and how those protections can be analyzed. It contains many case studies that demonstrate how to analyze and replicate request parameters, cookies, and encryption logic used by real websites. Each directory in the project focuses on a specific target service or scenario, showing how browser network requests and JavaScript code can be studied to reproduce API calls programmatically. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 11
    Dataproc Templates

    Dataproc Templates

    Dataproc templates and pipelines for solving simple in-cloud data task

    Dataproc templates are designed to address various in-cloud data tasks, including data import/export/backup/restore and bulk API operations. These templates leverage the power of Google Cloud's Dataproc, supporting both Dataproc Serverless and Dataproc clusters. Google provides this collection of pre-implemented Dataproc templates as a reference and for easy customization.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 12
    V2Rayfree

    V2Rayfree

    Latest public welfare free v2ray node subscription address

    V2Rayfree is a repository that aggregates and distributes free V2Ray node configurations, typically used for bypassing network restrictions and enabling private or unrestricted internet access. It functions as a continuously updated collection of subscription links and node information that users can import into compatible V2Ray clients. The project focuses on accessibility by offering free-tier options with defined bandwidth, device limits, and streaming capabilities, making it appealing to users seeking low-cost or no-cost solutions. It often includes curated recommendations for external providers, along with configuration data that supports multiple protocols such as V2Ray, Shadowsocks, and Trojan. ...
    Downloads: 60 This Week
    Last Update:
    See Project
  • 13
    Scweet

    Scweet

    Scrape tweets, profiles, followers and following from Twitter/X

    Scweet is a Python-based Twitter/X scraping library and CLI designed to collect tweets, profile timelines, followers, following lists, and user profile data without requiring the official Twitter/X API or a developer account. Instead of depending on deprecated unauthenticated scraping methods, it works by using X’s web GraphQL API together with authenticated browser cookies, which gives it a more current and practical approach for data extraction. The project supports a broad set of collection patterns, including searches by keyword, hashtag, user, date range, engagement thresholds, language, and location, making it useful for research, monitoring, and data gathering workflows. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    xhs-spider

    xhs-spider

    Desktop tool for collecting and exporting Xiaohongshu post data

    XHS-Spider is a desktop data collection tool designed to gather content and metadata from the Xiaohongshu platform. It provides a graphical interface that allows users to explore posts, collect information, and download media such as images and videos from individual notes or search results. It was developed primarily as a learning project to demonstrate approaches to building web crawlers and experimenting with technologies such as WebView2 and WPF UI.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 15
    OpenSEO

    OpenSEO

    Open source alternative to Semrush and Ahrefs

    ...It provides focused workflows for keyword research, rank tracking, competitor analysis, backlink research, site auditing, and AI visibility monitoring. A modern web interface keeps these tasks accessible without hiding them inside a large collection of unrelated reports. The platform exposes an MCP server so AI agents can work directly with SEO data and reusable agent skills. Users can run a hosted edition or self-host the application with Docker or Cloudflare. Self-hosted installations connect to DataForSEO with the user's own API key and use pay-as-you-go data rather than a mandatory software subscription. ...
    Downloads: 8 This Week
    Last Update:
    See Project
  • 16
    syslog-ng

    syslog-ng

    Log management solution that improves the performance of SIEM

    ...Instead of deploying multiple agents on hosts, organizations can unify their log data collection and management. syslog-ng Store Box provides automated archiving, tamper-proof encrypted storage, granular access controls to protect log data. The largest appliance can store up to 10TB of raw logs.
    Downloads: 5 This Week
    Last Update:
    See Project
  • 17
    Geziyor

    Geziyor

    Blazing fast Go framework for web crawling and data scraping tasks

    ...With built-in tools for caching, metrics collection, and proxy management, it enables developers to build robust and customizable scraping systems using Go.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 18
    OpenSearch

    OpenSearch

    Open source distributed and RESTful search engine

    ...It offers excellent performance and can scale up and down as the needs of the application grow or shrink. Its distributed design means that you interact with OpenSearch clusters. Each cluster is a collection of one or more nodes, servers that store your data and process search requests. You can run OpenSearch locally on a laptop, its system requirements are minimal, but you can also scale a single cluster to hundreds of powerful machines in a data center.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 19
    SwifterSwift

    SwifterSwift

    A collection of more than 500 native Swift extensions

    SwifterSwift is a collection of over 500 native Swift extensions, with handy methods, syntactic sugar, and performance improvements for a wide range of primitive data types, UIKit and Cocoa classes, over 500 in 1– for iOS, macOS, tvOS, watchOS, and Linux. CocoaPods is a dependency manager for Cocoa projects. To integrate SwifterSwift into your Xcode project using CocoaPods, specify it in your Podfile.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 20
    Aeron

    Aeron

    Efficient UDP unicast, UDP multicast, and IPC message transport

    ...Message streams can be recorded by the Archive module to persistent storage for later, or real-time, replay. Aeron Cluster provides support for fault-tolerant services as replicated state machines based on the Raft consensus algorithm. Performance is the key focus. A design goal for Aeron is to be the highest throughput with the lowest and most predictable latency of any messaging system. Aeron integrates with Simple Binary Encoding (SBE) for the best possible message encoding and decoding performance. Many of the data structures used in the creation of Aeron have been factored out to the Agrona project.
    Downloads: 17 This Week
    Last Update:
    See Project
  • 21
    Docker-ELK

    Docker-ELK

    The Elastic stack (ELK) powered by Docker and Compose

    A turnkey Docker Compose stack to spin up the ELK stack (Elasticsearch, Logstash, Kibana) for log collection, analysis, and visualization. Based on official Elastic images and enhanced with configuration defaults optimized for local development and testing.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 22
    Rybbit

    Rybbit

    Open-source and privacy-friendly alternative to Google Analytics

    Rybbit is an open-source, privacy-friendly web and product analytics platform positioned as a modern alternative to heavyweight tracking suites. It focuses on a clean UI and intuitive metrics so non-analysts can answer everyday questions without wading through dozens of reports. The data model is event-driven, enabling funnels, retention views, and simple user journeys without complex configuration. Because privacy is a first-class goal, it aims to minimize personal data collection and offers a self-host path so you can keep analytics in your own infrastructure. The project ships with a lightweight snippet and a Docker-based deployment, making it quick to trial or roll out. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 23
    KCP

    KCP

    A fast and reliable ARQ protocol

    KCP is a fast and reliable protocol that can achieve the transmission effect of a reduction of the average latency by 30% to 40% and reduction of the maximum delay by a factor of three, at the cost of 10% to 20% more bandwidth wasted than TCP. It is implemented by using the pure algorithm, and is not responsible for the sending and receiving of the underlying protocol (such as UDP), requiring the users to define their own transmission mode for the underlying data packet, and provide it to KCP in the way of callback. Even the clock needs to be passed in from the outside, without any internal system calls. The entire protocol has only two source files of ikcp.h, ikcp.c, which can be easily integrated into the user's own protocol stack. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    Waterfox

    Waterfox

    Privacy browser built for users who value customization & control

    ...It began as a 64-bit-focused fork and has grown into a full-featured browser available on Windows, macOS, Linux, and Android, making it a versatile alternative to mainstream browsers. Waterfox removes or disables much of Firefox’s telemetry and unnecessary data collection by default, giving users greater control over what is shared with third parties while still supporting classic and modern add-ons. The project emphasizes speed and efficiency through optimizations and choice, letting users tweak settings or install favored extensions without being locked into ecosystem constraints. Because it’s open source and actively developed by the community, new features, security updates, and performance improvements arrive regularly to align with web standards.
    Downloads: 26 This Week
    Last Update:
    See Project
  • 25
    watercrawl

    watercrawl

    AI-ready web crawler that extracts and structures website content

    ...WaterCrawl supports customizable extraction rules so users can focus only on relevant elements while ignoring unnecessary page components. WaterCrawl also offers real-time monitoring capabilities, allowing users to track crawling progress, performance metrics, and errors during large data collection jobs. Developers can integrate the tool into applications through a REST API and multiple client SDKs, enabling automated data pipelines and AI data preparation workflows.
    Downloads: 0 This Week
    Last Update:
    See Project