30 projects for "dataset" with 2 filters applied:

  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Try It Free
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 1
    GH Archive

    GH Archive

    GH Archive is a project to record the public GitHub timeline

    ...The dataset is also published through Google BigQuery for large-scale SQL-style exploration without downloading every archive. Its structure supports trend analysis, visualizations, machine learning, and open-source ecosystem research. The repository contains the crawler, supporting scripts, and website code, while the actual event files are hosted separately.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 2
    NYC Taxi Data

    NYC Taxi Data

    Import public NYC taxi and for-hire vehicle (Uber, Lyft)

    ...It also contains example analyses—spatial and temporal visualizations like maps, time-series plots, and hotspot detection—highlighting insights such as patterns of demand, peak times, and geospatial distributions. The repository is often used as a benchmark dataset and example for teaching, benchmarking, and demonstration purposes in the data science and urban analytics communities.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 3
    all AI news

    all AI news

    A list of online news & info sources in the AI/ML/Data Science space

    ...Overall, it provides a foundational dataset for tracking AI industry trends and updates.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    Marble Skill Taxonomy

    Marble Skill Taxonomy

    Structured taxonomy of what children learn across the primary years

    Marble Skill Taxonomy is an open structured dataset that maps what children learn across primary and elementary education. It breaks curriculum content into fine-grained micro-topics instead of leaving learning goals as broad standards. Each topic includes subject, domain, age range, description, evidence criteria, assessment prompt, and stable identifiers. The dataset also defines prerequisite relationships, so developers can trace what a learner should master before a concept. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 5
    v2ray-rules-dat

    v2ray-rules-dat

    V2Ray routing rules file enhanced version, which can replace V2Ray

    v2ray-rules-dat is a repository that compiles and distributes enhanced rule data (domain lists, geo-IP/geo-domain data, block/proxy/detect lists) intended for use with tools like V2Ray, Xray-core, and similar network/proxy frameworks. The dataset serves as an alternative or supplement to official geoip/ geosite data files, often providing more up-to-date, community-curated entries — enabling better routing, blocking, or traffic management when using those proxy tools. The repository is regularly updated (weekly sync upstream) and provides releases containing comprehensive data files (e.g. geoip.dat, geosite.dat, plus multiple .txt rule lists) plus checksums for integrity verification. ...
    Downloads: 21 This Week
    Last Update:
    See Project
  • 6
    Karpathy-Inspired Claude Code Guidelines

    Karpathy-Inspired Claude Code Guidelines

    A single CLAUDE.md file to improve Claude Code behavior

    Karpathy-Inspired Claude Code Guidelines is a curated learning and experimentation repository inspired by the work and teaching philosophy of Andrej Karpathy, designed to help learners build practical competence in deep learning, neural networks, and AI infrastructure. The project organizes a progressive path through exercises, notebooks, code examples, and practical mini-projects that echo Karpathy’s approach to “learning by doing,” where students build core concepts from first principles...
    Downloads: 4 This Week
    Last Update:
    See Project
  • 7
    Recursive Language Models

    Recursive Language Models

    General plug-and-play inference library for Recursive Language Models

    ...It provides a consistent API that abstracts away many of the repetitive engineering patterns in RL research and application work, letting developers focus on modeling, experimentation, and fine-tuning rather than infrastructure plumbing. Within the framework, you can define custom agents, environments, policy networks, and reward structures while leveraging built-in dataset utilities, logging, and checkpointing for reproducible experiments. RLM also includes integration with popular simulation environments and benchmark suites, giving researchers a ready-made playground for algorithm comparison and performance tracking.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 8
    User Agents

    User Agents

    A JavaScript library for generating random user agents with data

    ...Unlike simpler random user agent generators, it uses frequency-weighted datasets to ensure that generated values reflect how browsers are actually used in the wild. The dataset is updated automatically on a daily basis, ensuring that generated user agents remain current and relevant over time. In addition to user agent strings, the library can produce detailed browser fingerprint data such as screen size, platform, connection type, and device category. It also includes flexible filtering capabilities that allow developers to generate user agents matching specific criteria such as device type, operating system, or browser version.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 9
    kagglehub

    kagglehub

    Python library to access Kaggle resources

    kagglehub is a Python library for accessing Kaggle resources directly from Python code. It provides a simple API for downloading datasets, models, competition files, and notebook outputs without requiring users to manually manage every URL or file path. The library is designed to work both inside and outside Kaggle Notebooks, with native behavior that can adapt when it runs in Kaggle’s hosted notebook environment. It is useful for machine learning workflows where data, models, and notebook...
    Downloads: 1 This Week
    Last Update:
    See Project
  • Build Data Resilience - Take the Assessment Today Icon
    Build Data Resilience - Take the Assessment Today

    Can you recover when it matters most? Take this quick assessment to identify gaps and build greater recovery confidence.

    Is your recovery strategy as strong as you think? Take this quick self-assessment to check your recovery readiness and gain tailored insights. In only 2 minutes, you'll learn where you fall on the recovery readiness scale.
    Take the Assessment
  • 10
    LLM Datasets

    LLM Datasets

    Curated list of datasets and tools for post-training

    ...Quality is a recurring theme: examples and utilities help filter low-value samples, enforce length limits, and split train/validation consistently so results are comparable. Licensing and provenance are surfaced to encourage compliant usage and to guide dataset selection in commercial settings. For practitioners, the repo is a practical “starting pantry” that accelerates experimentation and helps keep data wrangling from dominating the project timeline.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 11
    ARC-AGI

    ARC-AGI

    The Abstraction and Reasoning Corpus

    ...The dataset is structured as grid-based puzzles, where each task requires understanding transformations such as symmetry, counting, or spatial manipulation. Unlike traditional machine learning benchmarks, ARC emphasizes generalization and reasoning over statistical pattern recognition, making it particularly challenging for current AI systems.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    YOLOV4 Pytorch

    YOLOV4 Pytorch

    This is a source code for YoloV4-pytorch that can be used to train you

    YOLOV4 Pytorch is a PyTorch implementation of the YOLOv4 object detection model for training and running custom detection systems. The repository is structured around practical workflows, including training, prediction, evaluation, anchor generation, model configuration, and dataset annotation utilities. It supports VOC-style datasets and includes scripts for prediction, mAP evaluation, FPS testing, video prediction, batch prediction, and heatmap generation. The project added multi-GPU training, seed settings for reproducible results, adaptive learning rate behavior based on batch size, and both step and cosine learning rate schedules. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    YOLOV3 Pytorch

    YOLOV3 Pytorch

    This is a source code for yolo3-pytorch

    YOLOV3 Pytorch is a PyTorch implementation of the YOLOv3 object detection model built for training, prediction, and evaluation. The repository provides a complete workflow for users who want to train their own object detector with VOC-style data or use pretrained weights. It includes utilities for annotation conversion, anchor generation, image prediction, video prediction, batch prediction, FPS measurement, heatmap output, and mAP evaluation. The project added multi-GPU training, target...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    Unet

    Unet

    Source code for unet-pytorch, which can train its own model

    ...The repository is built around training, prediction, and mIoU evaluation for VOC-style segmentation data and medical-style datasets. It includes scripts for general training, medical dataset training, prediction, annotation handling, model summaries, and evaluation. The project supports multiple backbones, data processing utilities, extensive comments, and adjustable training parameters. Its README notes that U-Net is better suited to datasets with fewer features and shallow visual structures, such as medical image segmentation, rather than complex VOC-style scenes. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 15
    DeepLabv3 Plus

    DeepLabv3 Plus

    Encoder-Decoder with Atrous Separable Convolution

    ...It implements the encoder-decoder architecture with atrous separable convolution and provides a practical workflow for training, prediction, and mIoU evaluation. The repository supports VOC-style segmentation datasets and includes utilities for annotation generation, JSON dataset conversion, model summary inspection, prediction, and metric calculation. It provides pretrained weight workflows for MobileNetV2 and Xception backbones and notes that the correct backbone should be selected during training and prediction. The project also supports multi-GPU training, multiple backbones, learning rate schedules with step and cosine options, optimizer selection, and adaptive learning rate behavior based on batch size. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Faster-Rcnn

    Faster-Rcnn

    This is a pytorch implementation library of faster-rcnn

    Faster-Rcnn is a PyTorch implementation of the Faster R-CNN two-stage object detection model. It is designed for training and evaluating detectors on VOC-format datasets, including VOC07+12 and custom datasets arranged with VOC-style annotations and images. The repository includes scripts for training, prediction, evaluation, annotation generation, and model summary inspection. It supports backbone options through pretrained VGG and ResNet weights, making it useful for comparing feature...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    SVoice (Speech Voice Separation)

    SVoice (Speech Voice Separation)

    We provide a PyTorch implementation of the paper Voice Separation

    ...Separate models are trained for different speaker counts, and the largest-capacity model dynamically determines the actual number of speakers in a mixture. The repository includes all necessary scripts for training, dataset preparation, distributed training, evaluation, and audio separation.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 18
    TensorFlow Examples

    TensorFlow Examples

    TensorFlow Tutorial and Examples for Beginners (support TF v1 & v2)

    ...For clarity and educational value, each example is accompanied by explanatory comments or markdown cells to illustrate what the code does and why — a design that makes it especially suitable for self-learners or students following along with real data. Besides raw implementations, the repo often shows best practices using higher-level constructs (e.g. dataset pipelines, estimators, layers) which reflect modern TensorFlow workflows rather than only textbook-style code.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    ReinventCommunity

    ReinventCommunity

    Jupyter Notebook tutorials for REINVENT 3.2

    This repository is a collection of useful jupyter notebooks, code snippets and example JSON files illustrating the use of Reinvent 3.2.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    Nerfies

    Nerfies

    This is the code for Deformable Neural Radiance Fields

    ...The training pipeline handles imperfect captures by modeling camera poses, exposure variations, and background segmentation, producing stable geometry and appearance. A set of utilities manages dataset preparation, pose estimation, and checkpoints so researchers can reproduce results on their own footage. The work sits at the intersection of graphics and vision, showing how learned volumetric rendering can handle human motion without dense markers or studio rigs.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    DeepDiff

    DeepDiff

    Amazingly incredible extraordinary lightning fast diffing in Swift

    ...The library relies on models conforming to a diff-aware protocol so items can be uniquely identified and compared. This makes it useful for deciding which elements were inserted, deleted, moved, or replaced between two versions of a dataset. DeepDiff helps reduce unnecessary UI reloads by making updates more targeted and animation-friendly. The repository is archived, so it is best treated as a stable reference or legacy dependency rather than an actively evolving Swift package.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Movies for Hackers

    Movies for Hackers

    A curated list of movies every hacker & cyberpunk must watch

    ...The selection highlights movies that explore themes of security, privacy, code, networks, and the social impact of technology, useful for entertainment and cultural context for technologists. Because it’s a plain-text, cross-platform resource, anyone can fork the list, propose additions, or reuse the dataset in their own tooling.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 23
    nlp_chinese_corpus

    nlp_chinese_corpus

    Large Scale Chinese Corpus for NLP

    ...The repository gathers several major datasets, including Chinese Wikipedia entries, news articles, encyclopedia-style question answering data, community question answering data, and Chinese-English translation sentence pairs. Each dataset includes descriptions, download links, structure notes, and examples to help users understand how the data is formatted. The corpora can support tasks such as language model pretraining, word vector training, question answering, title generation, keyword generation, translation, and sentence representation learning. Overall, it is a practical resource hub for building or testing Chinese NLP models with larger and more varied datasets.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    RobotsDisallowed

    RobotsDisallowed

    A curated list of the most common and most interesting robots.txt

    ...The project aggregates domains, notes the targeted bots or user agents, and surfaces patterns for researchers, policymakers, and tool builders. It serves both as a transparency effort and as a resource for people designing allow/deny strategies for automated access. The dataset invites community contributions to keep the picture current as new bots emerge and policies shift. It also highlights the intersection of web standards, ethics, and AI governance by showing how site owners operationalize consent and restriction at scale.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    benchm-ml

    benchm-ml

    A benchmark of commonly used open source implementations

    This repository is designed to provide a minimal benchmark framework comparing commonly used machine learning libraries in terms of scalability, speed, and classification accuracy. The focus is on binary classification tasks without missing data, where inputs can be numeric or categorical (after one-hot encoding). It targets large scale settings by varying the number of observations (n) up to millions and the number of features (after expansion) to about a thousand, to stress test different...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • 2
  • Next