46 projects for "dataset" with 2 filters applied:

  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    Exclusively Dark Image Dataset

    Exclusively Dark Image Dataset

    ExDARK dataset is the largest collection of low-light images

    The Exclusively Dark (ExDARK) dataset is one of the largest curated collections of real-world low-light images designed to support research in computer vision tasks under challenging lighting conditions. It contains 7,363 images captured across ten different low-light scenarios, ranging from extremely dark environments to twilight. Each image is annotated with both image-level labels and object-level bounding boxes for 12 object categories, making it suitable for detection and classification tasks. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 2
    GH Archive

    GH Archive

    GH Archive is a project to record the public GitHub timeline

    ...The dataset is also published through Google BigQuery for large-scale SQL-style exploration without downloading every archive. Its structure supports trend analysis, visualizations, machine learning, and open-source ecosystem research. The repository contains the crawler, supporting scripts, and website code, while the actual event files are hosted separately.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 3
    NYC Taxi Data

    NYC Taxi Data

    Import public NYC taxi and for-hire vehicle (Uber, Lyft)

    ...It also contains example analyses—spatial and temporal visualizations like maps, time-series plots, and hotspot detection—highlighting insights such as patterns of demand, peak times, and geospatial distributions. The repository is often used as a benchmark dataset and example for teaching, benchmarking, and demonstration purposes in the data science and urban analytics communities.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 4
    all AI news

    all AI news

    A list of online news & info sources in the AI/ML/Data Science space

    ...Overall, it provides a foundational dataset for tracking AI industry trends and updates.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Build Securely on AWS with Proven Frameworks Icon
    Build Securely on AWS with Proven Frameworks

    Lay a foundation for success with Tested Reference Architectures developed by Fortinet’s experts. Learn more in this white paper.

    Moving to the cloud brings new challenges. How can you manage a larger attack surface while ensuring great network performance? Turn to Fortinet’s Tested Reference Architectures, blueprints for designing and securing cloud environments built by cybersecurity experts. Learn more and explore use cases in this white paper.
    Download Now
  • 5
    Marble Skill Taxonomy

    Marble Skill Taxonomy

    Structured taxonomy of what children learn across the primary years

    Marble Skill Taxonomy is an open structured dataset that maps what children learn across primary and elementary education. It breaks curriculum content into fine-grained micro-topics instead of leaving learning goals as broad standards. Each topic includes subject, domain, age range, description, evidence criteria, assessment prompt, and stable identifiers. The dataset also defines prerequisite relationships, so developers can trace what a learner should master before a concept. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    v2ray-rules-dat

    v2ray-rules-dat

    V2Ray routing rules file enhanced version, which can replace V2Ray

    v2ray-rules-dat is a repository that compiles and distributes enhanced rule data (domain lists, geo-IP/geo-domain data, block/proxy/detect lists) intended for use with tools like V2Ray, Xray-core, and similar network/proxy frameworks. The dataset serves as an alternative or supplement to official geoip/ geosite data files, often providing more up-to-date, community-curated entries — enabling better routing, blocking, or traffic management when using those proxy tools. The repository is regularly updated (weekly sync upstream) and provides releases containing comprehensive data files (e.g. geoip.dat, geosite.dat, plus multiple .txt rule lists) plus checksums for integrity verification. ...
    Downloads: 21 This Week
    Last Update:
    See Project
  • 7
    Karpathy-Inspired Claude Code Guidelines

    Karpathy-Inspired Claude Code Guidelines

    A single CLAUDE.md file to improve Claude Code behavior

    Karpathy-Inspired Claude Code Guidelines is a curated learning and experimentation repository inspired by the work and teaching philosophy of Andrej Karpathy, designed to help learners build practical competence in deep learning, neural networks, and AI infrastructure. The project organizes a progressive path through exercises, notebooks, code examples, and practical mini-projects that echo Karpathy’s approach to “learning by doing,” where students build core concepts from first principles...
    Downloads: 4 This Week
    Last Update:
    See Project
  • 8
    Fast Paginate for Laravel

    Fast Paginate for Laravel

    A fast implementation of offset/limit pagination for Laravel.

    This is a fast limit/offset pagination macro for Laravel. It can be used in place of the standard paginate methods. This package uses a SQL method similar to a "deferred join" to achieve this speedup. A deferred join is a technique that defers access to requested columns until after the offset and limit have been applied. In our case we don't actually do a join, but rather a where in with a subquery. Using this technique we create a subquery that can be optimized with specific indexes for...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 9
    Kaggle CLI

    Kaggle CLI

    The official CLI to interact with Kaggle

    Kaggle CLI is Kaggle’s official command-line interface for interacting with the Kaggle platform from a terminal. It lets users authenticate, search resources, download files, submit competition entries, manage datasets, work with models, run notebooks, and read discussion content without relying only on the web interface. The tool is useful for data scientists who want to automate Kaggle workflows inside scripts, CI jobs, notebooks, or reproducible local environments. It supports both...
    Downloads: 2 This Week
    Last Update:
    See Project
  • Build Securely on Azure with Proven Frameworks Icon
    Build Securely on Azure with Proven Frameworks

    Lay a foundation for success with Tested Reference Architectures developed by Fortinet’s experts. Learn more in this white paper.

    Moving to the cloud brings new challenges. How can you manage a larger attack surface while ensuring great network performance? Turn to Fortinet’s Tested Reference Architectures, blueprints for designing and securing cloud environments built by cybersecurity experts. Learn more and explore use cases in this white paper.
    Download Now
  • 10
    Recursive Language Models

    Recursive Language Models

    General plug-and-play inference library for Recursive Language Models

    ...It provides a consistent API that abstracts away many of the repetitive engineering patterns in RL research and application work, letting developers focus on modeling, experimentation, and fine-tuning rather than infrastructure plumbing. Within the framework, you can define custom agents, environments, policy networks, and reward structures while leveraging built-in dataset utilities, logging, and checkpointing for reproducible experiments. RLM also includes integration with popular simulation environments and benchmark suites, giving researchers a ready-made playground for algorithm comparison and performance tracking.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 11
    Apache Spark

    Apache Spark

    A unified analytics engine for large-scale data processing

    Apache Spark is a unified engine for large-scale data processing, offering APIs for batch jobs, streaming, machine learning, and graph computation. It builds on resilient distributed datasets (RDDs) and the newer DataFrame/Dataset abstractions to provide fault-tolerant, in-memory computation across clusters. Spark’s execution engine handles scheduling, shuffles, caching, and data locality so users can focus on transformations rather than infrastructure plumbing. With Spark Streaming (microbatches) and Structured Streaming, it delivers low-latency event processing suitable for real-time analytics. ...
    Downloads: 7 This Week
    Last Update:
    See Project
  • 12
    User Agents

    User Agents

    A JavaScript library for generating random user agents with data

    ...Unlike simpler random user agent generators, it uses frequency-weighted datasets to ensure that generated values reflect how browsers are actually used in the wild. The dataset is updated automatically on a daily basis, ensuring that generated user agents remain current and relevant over time. In addition to user agent strings, the library can produce detailed browser fingerprint data such as screen size, platform, connection type, and device category. It also includes flexible filtering capabilities that allow developers to generate user agents matching specific criteria such as device type, operating system, or browser version.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 13
    kagglehub

    kagglehub

    Python library to access Kaggle resources

    kagglehub is a Python library for accessing Kaggle resources directly from Python code. It provides a simple API for downloading datasets, models, competition files, and notebook outputs without requiring users to manually manage every URL or file path. The library is designed to work both inside and outside Kaggle Notebooks, with native behavior that can adapt when it runs in Kaggle’s hosted notebook environment. It is useful for machine learning workflows where data, models, and notebook...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 14
    Synthetic Data Kit

    Synthetic Data Kit

    Tool for generating high quality Synthetic datasets

    ...It ships an opinionated, modular workflow that covers ingesting heterogeneous sources (documents, transcripts), prompting models to create labeled examples, and exporting to fine-tuning schemas with minimal glue code. The kit’s design goal is to shorten the “data prep” bottleneck by turning dataset creation into a repeatable pipeline rather than ad-hoc notebooks. It supports generation of rationales/chain-of-thought variants, configurable sampling, and guardrails so outputs meet format constraints and quality checks. Examples and guides show how to target task-specific behaviors like tool use or step-by-step reasoning, then save directly into training-ready files.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 15
    LLM Datasets

    LLM Datasets

    Curated list of datasets and tools for post-training

    ...Quality is a recurring theme: examples and utilities help filter low-value samples, enforce length limits, and split train/validation consistently so results are comparable. Licensing and provenance are surfaced to encourage compliant usage and to guide dataset selection in commercial settings. For practitioners, the repo is a practical “starting pantry” that accelerates experimentation and helps keep data wrangling from dominating the project timeline.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 16
    trustfall

    trustfall

    A query engine for any combination of data sources

    trustfall is a flexible and extensible query engine designed to retrieve and analyze data from a wide variety of heterogeneous sources using a unified query language. It allows developers to query APIs, databases, files, and even structured outputs from AI systems as if they were part of a single, cohesive graph-based dataset. The system is inspired by GraphQL-like paradigms, enabling expressive and composable queries that can traverse relationships across different data domains seamlessly. Trustfall operates by parsing, compiling, and executing queries through adapters that connect to specific data sources, making it highly modular and adaptable to new environments. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    EasyR1

    EasyR1

    An Efficient, Scalable, Multi-Modality RL Training Framework

    ...The framework is also organized to help you compare training strategies (e.g., pure SFT vs. preference optimization) so you can see what actually moves metrics in math, code, and multi-step reasoning. For teams exploring open reasoning models, EasyR1 provides an opinionated yet flexible path from dataset to deployable checkpoints.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    ARC-AGI

    ARC-AGI

    The Abstraction and Reasoning Corpus

    ...The dataset is structured as grid-based puzzles, where each task requires understanding transformations such as symmetry, counting, or spatial manipulation. Unlike traditional machine learning benchmarks, ARC emphasizes generalization and reasoning over statistical pattern recognition, making it particularly challenging for current AI systems.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    XMLRAD

    XMLRAD

    Web Application Server Stack for Delphi/FreePascal/Lazarus

    Build HTML5 web apps based on DSL (Domain Specific Language) for Delphi/FreePascal/Lazarus: - XMLGram to create XMLServices using DBExtract to select from DB, DBBatch to update DB, macros using programming patterns such as Multicast and RetroFit, - XQL (eXtensible Query Language) to select records in the embedded database XQLite, - XML to load/store record/dataset, - XTL (eXtensible Template Language) to design HTML5 pages based on templates, includes templates for multiple layouts: desktop, mobile and tablet devices XMLRAD includes an Embedded Web Server running natively on Windows and Linux with Delphi/FreePascal/Lazarus.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    YOLOV4 Pytorch

    YOLOV4 Pytorch

    This is a source code for YoloV4-pytorch that can be used to train you

    YOLOV4 Pytorch is a PyTorch implementation of the YOLOv4 object detection model for training and running custom detection systems. The repository is structured around practical workflows, including training, prediction, evaluation, anchor generation, model configuration, and dataset annotation utilities. It supports VOC-style datasets and includes scripts for prediction, mAP evaluation, FPS testing, video prediction, batch prediction, and heatmap generation. The project added multi-GPU training, seed settings for reproducible results, adaptive learning rate behavior based on batch size, and both step and cosine learning rate schedules. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    YOLOV3 Pytorch

    YOLOV3 Pytorch

    This is a source code for yolo3-pytorch

    YOLOV3 Pytorch is a PyTorch implementation of the YOLOv3 object detection model built for training, prediction, and evaluation. The repository provides a complete workflow for users who want to train their own object detector with VOC-style data or use pretrained weights. It includes utilities for annotation conversion, anchor generation, image prediction, video prediction, batch prediction, FPS measurement, heatmap output, and mAP evaluation. The project added multi-GPU training, target...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Unet

    Unet

    Source code for unet-pytorch, which can train its own model

    ...The repository is built around training, prediction, and mIoU evaluation for VOC-style segmentation data and medical-style datasets. It includes scripts for general training, medical dataset training, prediction, annotation handling, model summaries, and evaluation. The project supports multiple backbones, data processing utilities, extensive comments, and adjustable training parameters. Its README notes that U-Net is better suited to datasets with fewer features and shallow visual structures, such as medical image segmentation, rather than complex VOC-style scenes. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 23
    DeepLabv3 Plus

    DeepLabv3 Plus

    Encoder-Decoder with Atrous Separable Convolution

    ...It implements the encoder-decoder architecture with atrous separable convolution and provides a practical workflow for training, prediction, and mIoU evaluation. The repository supports VOC-style segmentation datasets and includes utilities for annotation generation, JSON dataset conversion, model summary inspection, prediction, and metric calculation. It provides pretrained weight workflows for MobileNetV2 and Xception backbones and notes that the correct backbone should be selected during training and prediction. The project also supports multi-GPU training, multiple backbones, learning rate schedules with step and cosine options, optimizer selection, and adaptive learning rate behavior based on batch size. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    Faster-Rcnn

    Faster-Rcnn

    This is a pytorch implementation library of faster-rcnn

    Faster-Rcnn is a PyTorch implementation of the Faster R-CNN two-stage object detection model. It is designed for training and evaluating detectors on VOC-format datasets, including VOC07+12 and custom datasets arranged with VOC-style annotations and images. The repository includes scripts for training, prediction, evaluation, annotation generation, and model summary inspection. It supports backbone options through pretrained VGG and ResNet weights, making it useful for comparing feature...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    SVoice (Speech Voice Separation)

    SVoice (Speech Voice Separation)

    We provide a PyTorch implementation of the paper Voice Separation

    ...Separate models are trained for different speaker counts, and the largest-capacity model dynamically determines the actual number of speakers in a mixture. The repository includes all necessary scripts for training, dataset preparation, distributed training, evaluation, and audio separation.
    Downloads: 1 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • 2
  • Next