Showing 134 open source projects for "dataset"

View related business solutions
  • PRTG Catches Network Issues Before They Cause Downtime Icon
    PRTG Catches Network Issues Before They Cause Downtime

    Threshold-based alerts flag problems early, so your team can act before users notice, not after.

    Reactive troubleshooting usually means hearing about a problem from frustrated users, not your monitoring tool. PRTG sets threshold-based alerts across devices, servers and applications, notifying your team by email, SMS or push the moment a metric crosses a set limit. That means catching a failing disk or overloaded server before it becomes an outage and getting time back from firefighting. Start a free trial and set your first alerts today.
    Download 30-Day Trial
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Start Free
  • 1
    Passport Index Dataset

    Passport Index Dataset

    Passport Index 2023: visa requirements for 199 countries, in .csv

    There are 6 datasets with identical visa requirements data. Three datasets are matrix and three are long (tidy) formats. Each comes in 3 versions: with country codes as specified in ISO-2 (two-letter codes), ISO-3 (three-letter codes), and full country names from no particular standard. In distance matrices (files with matrix in the filename), the first column represents a passport (=from), each remaining column represents a destination (=to). Files in tidy format (with tidy in filename)...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 2
    Exclusively Dark Image Dataset

    Exclusively Dark Image Dataset

    ExDARK dataset is the largest collection of low-light images

    The Exclusively Dark (ExDARK) dataset is one of the largest curated collections of real-world low-light images designed to support research in computer vision tasks under challenging lighting conditions. It contains 7,363 images captured across ten different low-light scenarios, ranging from extremely dark environments to twilight. Each image is annotated with both image-level labels and object-level bounding boxes for 12 object categories, making it suitable for detection and classification tasks. ...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 3
    Image Harmonization Dataset iHarmony4

    Image Harmonization Dataset iHarmony4

    The first large-scale public benchmark dataset for image harmonization

    This repository provides the iHarmony4 dataset, which is a large-scale dataset designed for image harmonization tasks. Image harmonization involves adjusting the appearance of a foreground in a composite image so that it is consistent with the background (in color, tone, illumination, etc.). The iHarmony4 dataset comprises four sub-datasets (HCOCO, HAdobe5k, HFlickr, Hday2night), each making composite images by combining a foreground from one image with a background from another, along with associated ground truth harmonized images and foreground masks. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    GH Archive

    GH Archive

    GH Archive is a project to record the public GitHub timeline

    ...The dataset is also published through Google BigQuery for large-scale SQL-style exploration without downloading every archive. Its structure supports trend analysis, visualizations, machine learning, and open-source ecosystem research. The repository contains the crawler, supporting scripts, and website code, while the actual event files are hosted separately.
    Downloads: 2 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 5
    all AI news

    all AI news

    A list of online news & info sources in the AI/ML/Data Science space

    ...Overall, it provides a foundational dataset for tracking AI industry trends and updates.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    NYC Taxi Data

    NYC Taxi Data

    Import public NYC taxi and for-hire vehicle (Uber, Lyft)

    ...It also contains example analyses—spatial and temporal visualizations like maps, time-series plots, and hotspot detection—highlighting insights such as patterns of demand, peak times, and geospatial distributions. The repository is often used as a benchmark dataset and example for teaching, benchmarking, and demonstration purposes in the data science and urban analytics communities.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 7
    Marble Skill Taxonomy

    Marble Skill Taxonomy

    Structured taxonomy of what children learn across the primary years

    Marble Skill Taxonomy is an open structured dataset that maps what children learn across primary and elementary education. It breaks curriculum content into fine-grained micro-topics instead of leaving learning goals as broad standards. Each topic includes subject, domain, age range, description, evidence criteria, assessment prompt, and stable identifiers. The dataset also defines prerequisite relationships, so developers can trace what a learner should master before a concept. ...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 8
    React Chart.js

    React Chart.js

    React components for Chart.js, the most popular charting library

    ...In order to improve performance, offer new features, and improve maintainability, it was necessary to break backwards compatibility, but we aimed to do so only when worth the benefit. You will find that any event which causes the chart to re-render, such as hover tooltips, etc., will cause the first dataset to be copied over to other datasets, causing your lines and bars to merge together. This is because to track changes in the dataset series, the library needs a key to be specified. If none is found, it can't tell the difference between the datasets while updating. Specify a different property to be used as a key by passing a datasetIdKey prop to your chart component.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 9
    Cloud Native Landscape

    Cloud Native Landscape

    The Cloud Native Interactive Landscape filters

    Cloud Native Landscape is an open-source project that serves as a comprehensive visual map and database of the cloud-native ecosystem, cataloging hundreds of tools, platforms, and services across categories such as orchestration, observability, networking, and storage. Rather than being a traditional software application, it is a curated dataset and visualization system that helps developers and organizations navigate the rapidly evolving cloud-native space. The project organizes technologies into structured categories, making it easier to understand relationships and identify tools relevant to specific use cases. It includes strict contribution guidelines to maintain consistency and ensure that listed projects meet certain criteria, such as community adoption and relevance. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Start Free
  • 10
    BIMserver

    BIMserver

    The open source BIMserver platform

    ...The main advantage of this approach is the ability to query, merge and filter the BIM model and generate IFC output (i.e. files) on the fly. Thanks to its multi-user support, multiple people can work on their own part of the dataset, while the complete dataset is updated on the fly. Other users can get notifications when the model (or a part of it) is updated.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Charts.css

    Charts.css

    Open source CSS framework for data visualization

    Charts.css is a modern CSS framework. It uses CSS utility classes to style HTML elements as charts. No dependencies. 72kb file size. Less than 6kb gzipped file size! Visualization helps end-users understand data. Charts.css help frontend developers turn data into beautiful charts and graphs using simple CSS classes. The data is structured using semantic HTML tags and styled using CSS classes which change the visual representation displayed to the end-user. The framework offers developers...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    SponsorBlock

    SponsorBlock

    Skip YouTube video sponsors (browser extension)

    SponsorBlock is an open-source crowdsourced browser extension and open API for skipping sponsor segments in YouTube videos. Users submit when a sponsor happens from the extension, and the extension automatically skips sponsors it knows about using a privacy-preserving query system. It also supports skipping other categories, such as intros, outros, and reminders to subscribe, and skipping to the point with highlights. The extension also features an upvote/downvote system with a weighted...
    Downloads: 61 This Week
    Last Update:
    See Project
  • 13
    UCO3D

    UCO3D

    Uncommon Objects in 3D dataset

    uCO3D is a large-scale 3D vision dataset and toolkit centered on turn-table videos of everyday objects drawn from the LVIS taxonomy. It provides about 170,000 full videos per object instance rather than still frames, along with per-video annotations including object masks, calibrated camera poses, and multiple flavors of point clouds. Each sequence also ships with a precomputed 3D Gaussian Splat reconstruction, enabling fast, differentiable rendering workflows and modern implicit/point-based modeling experiments. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    Petastorm

    Petastorm

    Petastorm library enables single machine or distributed training

    ...This library enables single machine or distributed training and evaluation of deep learning models directly from datasets in Apache Parquet format. Petastorm supports popular Python-based machine learning (ML) frameworks such as Tensorflow, PyTorch, and PySpark. It can also be used from pure Python code. A dataset created using Petastorm is stored in Apache Parquet format. On top of a Parquet schema, petastorm also stores higher-level schema information that makes multidimensional arrays into a native part of a petastorm dataset. Petastorm supports extensible data codecs. These enable a user to use one of the standard data compressions (jpeg, png) or implement her own.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    emojilib

    emojilib

    Emoji keyword library

    Emoji keyword library. Make emoji searchable with this keyword library. If you are looking for the unicode emoji dataset, including version, grouping, ordering, and skin tone support flag, check out unicode-emoji-json.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Flipper-IRDB

    Flipper-IRDB

    A collective of different IRs for the Flipper (maintained)

    Flipper-IRDB is a community-maintained collection of infrared remote-control files for Flipper Zero devices. It organizes captured signals by device type, brand, and model so users can find compatible controls quickly. The collection covers televisions, audio equipment, air conditioners, projectors, lighting, toys, appliances, and many other infrared-controlled products. Files use Flipper’s readable IR format and follow shared naming conventions for devices and buttons. Community...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 17
    v2ray-rules-dat

    v2ray-rules-dat

    V2Ray routing rules file enhanced version, which can replace V2Ray

    v2ray-rules-dat is a repository that compiles and distributes enhanced rule data (domain lists, geo-IP/geo-domain data, block/proxy/detect lists) intended for use with tools like V2Ray, Xray-core, and similar network/proxy frameworks. The dataset serves as an alternative or supplement to official geoip/ geosite data files, often providing more up-to-date, community-curated entries — enabling better routing, blocking, or traffic management when using those proxy tools. The repository is regularly updated (weekly sync upstream) and provides releases containing comprehensive data files (e.g. geoip.dat, geosite.dat, plus multiple .txt rule lists) plus checksums for integrity verification. ...
    Downloads: 16 This Week
    Last Update:
    See Project
  • 18
    Apache Spark

    Apache Spark

    A unified analytics engine for large-scale data processing

    Apache Spark is a unified engine for large-scale data processing, offering APIs for batch jobs, streaming, machine learning, and graph computation. It builds on resilient distributed datasets (RDDs) and the newer DataFrame/Dataset abstractions to provide fault-tolerant, in-memory computation across clusters. Spark’s execution engine handles scheduling, shuffles, caching, and data locality so users can focus on transformations rather than infrastructure plumbing. With Spark Streaming (microbatches) and Structured Streaming, it delivers low-latency event processing suitable for real-time analytics. ...
    Downloads: 16 This Week
    Last Update:
    See Project
  • 19
    Kaggle CLI

    Kaggle CLI

    The official CLI to interact with Kaggle

    Kaggle CLI is Kaggle’s official command-line interface for interacting with the Kaggle platform from a terminal. It lets users authenticate, search resources, download files, submit competition entries, manage datasets, work with models, run notebooks, and read discussion content without relying only on the web interface. The tool is useful for data scientists who want to automate Kaggle workflows inside scripts, CI jobs, notebooks, or reproducible local environments. It supports both...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 20
    dejavu

    dejavu

    The missing web UI for Elasticsearch

    dejavu is the missing web UI for Elasticsearch. Existing web UIs leave much to be desired or are built with server-side page rendering techniques that make it less responsive and bulkier to run (I am looking at you, Kibana). We started building dejavu with the goal of creating a modern Web UI (no page reloads, infinite scroll, filtered views, realtime updates, search UI builder) for Elasticsearch with 100% client-side rendering so one can easily run it as a hosted app on github pages, as a...
    Downloads: 5 This Week
    Last Update:
    See Project
  • 21
    User Agents

    User Agents

    A JavaScript library for generating random user agents with data

    ...Unlike simpler random user agent generators, it uses frequency-weighted datasets to ensure that generated values reflect how browsers are actually used in the wild. The dataset is updated automatically on a daily basis, ensuring that generated user agents remain current and relevant over time. In addition to user agent strings, the library can produce detailed browser fingerprint data such as screen size, platform, connection type, and device category. It also includes flexible filtering capabilities that allow developers to generate user agents matching specific criteria such as device type, operating system, or browser version.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 22
    Fast Paginate for Laravel

    Fast Paginate for Laravel

    A fast implementation of offset/limit pagination for Laravel.

    This is a fast limit/offset pagination macro for Laravel. It can be used in place of the standard paginate methods. This package uses a SQL method similar to a "deferred join" to achieve this speedup. A deferred join is a technique that defers access to requested columns until after the offset and limit have been applied. In our case we don't actually do a join, but rather a where in with a subquery. Using this technique we create a subquery that can be optimized with specific indexes for...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Angular DataTables

    Angular DataTables

    DataTables with Angular

    An Angular2+ library for building complex HTML tables using DataTables JQuery plug-in. Implementation of the example on custom filtering with range search. The HTML element provides a Promise that returns the instance of the DataTable. Implementation of the example on individual column searching (text inputs). Sometimes, your DataTable options are stored or computed server-side. All you need to do is to return the expected result as a promise. You can use Angular Pipe to transform data on...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    Synthetic Data Vault (SDV)

    Synthetic Data Vault (SDV)

    Synthetic Data Generation for tabular, relational and time series data

    The Synthetic Data Vault (SDV) is a Synthetic Data Generation ecosystem of libraries that allows users to easily learn single-table, multi-table and timeseries datasets to later on generate new Synthetic Data that has the same format and statistical properties as the original dataset. Synthetic data can then be used to supplement, augment and in some cases replace real data when training Machine Learning models. Additionally, it enables the testing of Machine Learning or other data dependent software systems without the risk of exposure that comes with data disclosure. Underneath the hood it uses several probabilistic graphical modeling and deep learning based techniques. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 25
    Synthetic Data Kit

    Synthetic Data Kit

    Tool for generating high quality Synthetic datasets

    ...It ships an opinionated, modular workflow that covers ingesting heterogeneous sources (documents, transcripts), prompting models to create labeled examples, and exporting to fine-tuning schemas with minimal glue code. The kit’s design goal is to shorten the “data prep” bottleneck by turning dataset creation into a repeatable pipeline rather than ad-hoc notebooks. It supports generation of rationales/chain-of-thought variants, configurable sampling, and guardrails so outputs meet format constraints and quality checks. Examples and guides show how to target task-specific behaviors like tool use or step-by-step reasoning, then save directly into training-ready files.
    Downloads: 1 This Week
    Last Update:
    See Project