Showing 643 open source projects for "dataset"

View related business solutions
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Try Free
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Get Started Free
  • 1
    Azure Machine Learning Python SDK

    Azure Machine Learning Python SDK

    Python notebooks with ML and deep learning examples

    ...The content spans a wide range of real-world tasks — from foundational quickstarts that teach users how to configure an Azure ML workspace and connect to compute resources, to advanced tutorials on using pipelines, automated machine learning, and dataset handling. Because it is designed to work with Azure Machine Learning compute instances, many notebooks can be executed directly in the cloud without additional setup, but they can also run locally with the appropriate SDK and packages installed. Each notebook includes code, narrative explanations, and example workflows that help users build reproducible machine learning solutions, which are key for operationalizing models in production.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Resemblyzer

    Resemblyzer

    A python package to analyze and compare voices with deep learning

    ...It works by turning speech audio into a compact voice embedding that represents the speaker’s vocal characteristics. These embeddings can then be used for speaker similarity, clustering, diarization experiments, voice comparison, and audio dataset exploration. The project is useful for researchers and developers who need a practical way to reason about speaker identity without building a voice encoder from scratch. It can help identify whether two recordings sound like the same speaker or visualize voice relationships across many samples. Its main value is making speaker representation accessible through a simple Python workflow.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3

    SPIKE dataset

    Images and models for detection of wheat spikes in the field

    The SPIKE dataset contains images with manually annotated ground truth labels for training and testing convolutional neural networks for spike detection of wheat in the field. The model files available here are to be used with the Faster R-CNN implementation of the freely available Microsoft Cognitive Toolkit, found at the following link https://www.microsoft.com/en-us/cognitive-toolkit/ Futhermore, an additional software tool will soon be available for using the models with an executable program requiring less setup and more visual steps of the process.
    Downloads: 39 This Week
    Last Update:
    See Project
  • 4
    DCGAN in TensorLayerX

    DCGAN in TensorLayerX

    The Simplest DCGAN Implementation

    This is an implementation of Deep Convolutional Generative Adversarial Networks. First, download the aligned face images from google or baidu to a data folder. Please place dataset 'img_align_celeba.zip' under 'data/celebA/' by default.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • 5
    benchm-ml

    benchm-ml

    A benchmark of commonly used open source implementations

    This repository is designed to provide a minimal benchmark framework comparing commonly used machine learning libraries in terms of scalability, speed, and classification accuracy. The focus is on binary classification tasks without missing data, where inputs can be numeric or categorical (after one-hot encoding). It targets large scale settings by varying the number of observations (n) up to millions and the number of features (after expansion) to about a thousand, to stress test different...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    RobotsDisallowed

    RobotsDisallowed

    A curated list of the most common and most interesting robots.txt

    ...The project aggregates domains, notes the targeted bots or user agents, and surfaces patterns for researchers, policymakers, and tool builders. It serves both as a transparency effort and as a resource for people designing allow/deny strategies for automated access. The dataset invites community contributions to keep the picture current as new bots emerge and policies shift. It also highlights the intersection of web standards, ethics, and AI governance by showing how site owners operationalize consent and restriction at scale.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    Facets

    Facets

    Visualizations for machine learning datasets

    ...Understanding your data is critical to building a powerful machine learning system. Facets contains two robust visualizations to aid in understanding and analyzing machine learning datasets. Get a sense of the shape of each feature of your dataset using Facets Overview, or explore individual observations using Facets Dive. Explore Facets Overview and Facets Dive on the UCI Census Income dataset, used for predicting whether an individual’s income exceeds $50K/yr based on their census data. The census data contains features such as age, education level, and occupation for each individual. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    RoboSat

    RoboSat

    Semantic segmentation on aerial and satellite imagery

    RoboSat is an end-to-end pipeline written in Python 3 for feature extraction from aerial and satellite imagery. Features can be anything visually distinguishable in the imagery for example: buildings, parking lots, roads, or cars.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    ChainerCV

    ChainerCV

    ChainerCV: a Library for Deep Learning in Computer Vision

    ...Bounding boxes in an image are represented as a two-dimensional array of shape (R,4), where R is the number of bounding boxes and the second axis corresponds to the coordinates of bounding boxes. ChainerCV supports dataset loaders, which can be used to easily index examples with list-like interfaces. Dataset classes whose names end with BboxDataset contain annotations of where objects locate in an image and which categories they are assigned to. These datasets can be indexed to return a tuple of an image, bounding boxes and labels. ChainerCV provides several network implementations that carry out object detection.
    Downloads: 0 This Week
    Last Update:
    See Project
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 10
    maskrcnn-benchmark

    maskrcnn-benchmark

    Fast, modular reference implementation of Instance Segmentation

    Mask R-CNN Benchmark is a PyTorch-based framework that provides high-performance implementations of object detection, instance segmentation, and keypoint detection models. Originally built to benchmark Mask R-CNN and related models, it offers a clean, modular design to train and evaluate detection systems efficiently on standard datasets like COCO. The framework integrates critical components—region proposal networks (RPNs), RoIAlign layers, mask heads, and backbone architectures such as...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11

    pyrus-seq

    Pyrus analyzes paired-end whole genome sequencing data

    Pyrus analyzes paired-end and non-paired-end whole genome sequencing data. The program is designed to identify and statistically weight potential chromosomal rearrangements identified in the data using parameters derived from the dataset itself.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    CakeChat

    CakeChat

    CakeChat: Emotional Generative Dialog System

    CakeChat is a backend for chatbots that are able to express emotions via conversations. The code is flexible and allows to condition model's responses by an arbitrary categorical variable. For example, you can train your own persona-based neural conversational model or create an emotional chatting machine. Hierarchical Recurrent Encoder-Decoder (HRED) architecture for handling deep dialog context. Multilayer RNN with GRU cells. The first layer of the utterance-level encoder is always...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    MAML-Pytorch

    MAML-Pytorch

    Elegant PyTorch implementation of paper Model-Agnostic Meta-Learning

    ...It focuses on reproducing and exploring the MAML approach for few-shot learning research. The repository supports MiniImagenet and Omniglot, two common benchmark datasets for meta-learning experiments. It includes separate training scripts, dataset loaders, learner components, and meta-learning logic. The project also notes that MAML can be difficult to train and presents the implementation as a practical starting point for research. Overall, it is useful for students and researchers who want to study fast adaptation, few-shot classification, and gradient-based meta-learning in PyTorch.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 14
    Active Learning

    Active Learning

    Framework and examples for active learning with machine learning model

    ...The system allows researchers to study how models can improve labeling efficiency by selectively querying the most informative data points rather than relying on uniformly sampled training sets. The main experiment runner (run_experiment.py) supports a wide range of configurations, including batch sizes, dataset subsets, model selection, and data preprocessing options. It includes several established active learning strategies such as uncertainty sampling, k-center greedy selection, and bandit-based methods, while also allowing for custom algorithm implementations. The framework integrates with both classical machine learning models (SVM, logistic regression) and neural networks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Papers with Code

    Papers with Code

    List of different papers for coding

    ...It was originally created to support the discovery and reproducibility of AI research by connecting scholarly work with working software projects. Although the repository itself is no longer actively maintained, it still provides a historical dataset that reflects many influential research publications and their associated implementations.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 16
    Yolo_mark

    Yolo_mark

    GUI for marking bounded boxes of objects in images

    ...It can also extract every specified frame interval from a video before manual labeling begins. OpenCV support, sample configuration files, object tracking, coordinate display, and training instructions make it a compact workflow for legacy YOLO dataset creation.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17

    Real Vs Random Double Star Counts

    C code to parse the Gaia DR2 dataset to find real and random binaries.

    d2rSequenceTst.c : Checks that all stars are in increasing right ascension order. g2dataStructures.h: Data structure for Gaia DR2 data gaiaGR2toNa.c: Takes the Gaia DR2 ascii, csv data and formats them into g2data structures. sortNaGR2dataByRa.c: Sorts g2data structures into a list ordered by increasing right ascension. tstGaiaDR2data.c: A generic program to test a g2data file. utilities.tgz: Files called out by the above C programs in their include statements.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Omniglot

    Omniglot

    Omniglot data set for one-shot learning

    This repository hosts the Omniglot dataset for one-shot learning, containing handwritten characters across multiple alphabets along with stroke data. It includes both MATLAB and Python starter scripts (e.g. demo.m, demo.py) to illustrate how to load the images and stroke sequences and run baseline experiments (such as classification by modified Hausdorff distance). The dataset provides both an image representation of each character and the time-ordered stroke coordinates ([x, y, t]) for each instance. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 19
    The SPInDel workbench is a computational platform to facilitate the planning and management of SPInDel projects, alignment of nucleotide sequences, visualization and selection of conserved regions, calculation of PCR primers properties, prediction of SPInDel profiles and diverse statistical and phylogenetic analyses. It includes a large dataset comprising nearly 1,800 numeric profiles for the identification of eukaryotic, prokaryotic and viral species.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    lazynlp

    lazynlp

    Library to scrape and clean web pages to create massive datasets

    LazyNLP is a lightweight tool for collecting and curating large-scale text datasets for machine learning and NLP applications with minimal manual effort.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    SSD

    SSD

    A PyTorch Implementation of Single Shot MultiBox Detector

    SSD is a PyTorch implementation of the Single Shot MultiBox Detector, a well-known object detection architecture introduced in the original SSD paper. It is built to help users train, evaluate, and experiment with object detection models using PyTorch rather than the original Caffe implementation. The repository includes the major components needed for an object detection workflow, including training scripts, evaluation scripts, demos, and utility modules. It supports commonly used benchmark...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Tacotron-2

    Tacotron-2

    DeepMind's Tacotron-2 Tensorflow implementation

    ...It reproduces the original paper’s hyperparameters exactly via paper_hparams.py, while also offering a tuned hparams.py with extra improvements that often yield better audio quality in practice. The repository is structured as a full training pipeline: dataset preparation, preprocessing into spectrograms, Tacotron training, WaveNet (or Griffin-Lim) vocoder training, and final waveform synthesis. It includes directory layouts and logging directories for multiple datasets such as LJSpeech and M-AILABS en_US/en_UK, making it easier to adapt to new English corpora. Separate log trees track mel-spectrograms, attention plots, evaluation audio, and vocoder outputs, so you can inspect how alignment and audio quality evolve over time.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Finetune Transformer LM

    Finetune Transformer LM

    Code for "Improving Language Understanding by Generative Pre-Training"

    finetune-transformer-lm is a research codebase that accompanies the paper “Improving Language Understanding by Generative Pre-Training,” providing a minimal implementation focused on fine-tuning a transformer language model for evaluation tasks. The repository centers on reproducing the ROCStories Cloze Test result and includes a single-command training workflow to run the experiment end to end. It documents that runs are non-deterministic due to certain GPU operations and reports a median...
    Downloads: 5 This Week
    Last Update:
    See Project
  • 24
    LUMINOTH

    LUMINOTH

    Deep Learning toolkit for Computer Vision

    ...Luminoth includes support for popular object detection architectures such as Faster R-CNN and SSD, enabling developers to train models on datasets like COCO and Pascal VOC. The toolkit provides command-line utilities for dataset management, training, and inference, making it easier to integrate into research workflows and production systems. Although the project is no longer actively maintained, it remains a useful educational and experimental platform for studying object detection pipelines and deep learning workflows.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    Deepvoice3_pytorch

    Deepvoice3_pytorch

    PyTorch implementation of convolutional neural networks

    An open source implementation of Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning.
    Downloads: 0 This Week
    Last Update:
    See Project