Showing 643 open source projects for "dataset"

View related business solutions
  • 99.99% Uptime for MySQL and PostgreSQL Databases Icon
    99.99% Uptime for MySQL and PostgreSQL Databases

    Sub-second maintenance. 2x read/write performance. Built-in vector search for AI apps.

    Cloud SQL Enterprise Plus delivers near-zero downtime with 35 days of point-in-time recovery. Supports MySQL, PostgreSQL, and SQL Server.
    Try Free
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 1

    ZFS Snapshot Extractor

    Bash script to extract a specific file from ZFS snapshots

    This is a simple Bash script to extract a specific version of a given file from ZFS snapshots. It must be started in the directory where the ZFS dataset is mounted. The relative path for the file must be given as parameter. It then displays a list of all files with a unique modification time time from the snapshots under .zfs/scnapshot and allows to select one (or several) to be copied to the user's home directory. Example: cd /tank/mydata zfs-snapshot-extract.sh documents/report.txt
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    MuseGAN

    MuseGAN

    An AI for Music Generation

    ...The architecture is based on convolutional GAN models that learn temporal musical structure and inter-track relationships from training data. The project was trained using the Lakh Pianoroll Dataset, a large collection of multitrack musical sequences derived from MIDI files.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 3

    ParDRe

    Parallel tool to remove duplicate DNA reads

    ...Duplicate reads can be seen as identical or nearly identical sequences with some mismatches. This tool will let the users to avoid the analysis of not necessary reads, reducing the time of subsequent procedures with the dataset (e.g., assemblies, mappings, etc.). The tool is implemented with MPI in order to exploit the parallel capabilities of multicore clusters. It is faster than multithreaded counterparts (end of 2015) for the same number of cores and, thanks to the message-passing technology, it can be executed on clusters. There also exists a MapReduce counterpart of ParDRe, called MarDRe (see the link above). ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 4

    lmdExplorer

    A midiexplorer clone designed to deal with the lmd_full midi database.

    The lmd_full database is a collection of more than 176,000 midi files occupying 6.2Gb. Finding specific files in this collection is far from straight forward since all the file names are 32 character md5 checksum names. Like midiexplorer, this clone provides various tools to visualize the midi files and to search for files having specific features. In addition, lmdExplorer provides an interface to the midicaps json file. lmdExplorer like midiexplorer is a user interface to numerous...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Get Started Free
  • 5
    PetoronAI-Drug-Discovery

    PetoronAI-Drug-Discovery

    PetoronAI Drug Discovery Experiments

    # PetoronAI Drug Discovery Benchmark https://github.com/01alekseev/PetoronAI PetoronAI was evaluated on the NCI-ALMANAC development dataset. • 2,225,137 experimental records • 602 experimentally measured drug pairs • Blind pair-level benchmark • 20×10 cross-validation • 99.5% validation coverage Validation: Correlation = 0.617327 MAE = 3.971799 Null model MAE = 4.656203 Sign accuracy = 64.69% Permutation p = 0.000100 Top HSA hypotheses: • Dactinomycin + Vinblastine sulfate (9.096) • Cabazitaxel + Vinblastine sulfate (8.224) • Mitoxantrone + Vinblastine sulfate (7.254) • Cabazitaxel + Mitoxantrone (7.128) • Dactinomycin + Daunorubicin HCl (6.965) Generated laboratory validation protocols (8×8 dose matrix, HSA, Bliss, Loewe, ZIP). ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    XMLRAD

    XMLRAD

    Web Application Server Stack for Delphi/FreePascal/Lazarus

    Build HTML5 web apps based on DSL (Domain Specific Language) for Delphi/FreePascal/Lazarus: - XMLGram to create XMLServices using DBExtract to select from DB, DBBatch to update DB, macros using programming patterns such as Multicast and RetroFit, - XQL (eXtensible Query Language) to select records in the embedded database XQLite, - XML to load/store record/dataset, - XTL (eXtensible Template Language) to design HTML5 pages based on templates, includes templates for multiple layouts: desktop, mobile and tablet devices XMLRAD includes an Embedded Web Server running natively on Windows and Linux with Delphi/FreePascal/Lazarus.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    Zylthra

    Zylthra

    Zylthra: A PyQt6 app to generate synthetic datasets with DataLLM.

    Welcome to Zylthra, a powerful Python-based desktop application built with PyQt6, designed to generate synthetic datasets using the DataLLM API from data.mostly.ai. This tool allows users to create custom datasets by defining columns, configuring generation parameters, and saving setups for reuse, all within a sleek, dark-themed interface.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    CPQ Metrics & KPI Dictionary

    CPQ Metrics & KPI Dictionary

    Standard CPQ metrics for quote speed, approvals and commercial control

    The CPQ Metrics & KPI Dictionary is a free, vendor-neutral reference for standardising how quote speed, approval efficiency, commercial control, quote quality and commercial outcomes are defined and measured. Version 1.0.0 contains 14 controlled metrics across four categories: Speed and flow, Approval efficiency, Commercial control, and Quality and outcome. Includes an Excel workbook, CSV and JSON dictionaries, JSON Schema, dashboard specification, data requirements, implementation...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Gemi

    Gemi

    PCR primers / probes design from multiple & degenerate sequences

    ...Gemi accepts multiple aligned and long DNA and RNA sequences with degenerate nucleotide (non-A/C/G/T bases). Gemi can be used for quantitative, real-time and conventional PCR (qPCR, rt-PCR, etc.), and Sanger sequencing. Gemi can parse large dataset of sequences efficiently. Python source code is available upon request. Milestone: The tool reached about 3000 downloads sine 2012. Article Gemi: PCR primers prediction from multiple alignments. Comparative and functional genomics 2012 PMID: https://www.ncbi.nlm.nih.gov/pubmed/23316117 A great review on designing primer, Gemi, and other tools: Designing degenerate primers: Overview, challenges, and computational methods. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 10
    code-act

    code-act

    Official Repo for ICML 2024 paper

    code-act is a research framework for building intelligent language-model agents that interact with their environment through executable code actions. The system proposes a unified action representation where language models produce Python code that can be executed directly, allowing the model to interact with external tools and environments in a structured way. By integrating a Python interpreter with the agent architecture, the system enables the agent to execute code, observe the results,...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Fondant

    Fondant

    Production-ready data processing made easy and shareable

    ...Fondant is designed with reproducibility in mind and supports containerized steps using Docker, making it easy to share and reuse data processing components. It’s built for use in research and production, empowering data scientists to streamline dataset curation and preprocessing workflows efficiently.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    Featuretools

    Featuretools

    An open source python library for automated feature engineering

    ...Featuretools come with a library of low-level functions that can be stacked to create features. You can build and share your own custom primitives to be reused on any dataset. Featuretools works alongside tools you already use to build machine learning pipelines. You can load in pandas' data frames and automatically create meaningful features in a fraction of the time it would take to do so manually.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    NKTgLaw

    NKTgLaw

    Core library & API for the NKTg Law (Nguyen Khanh Tung). Includes core

    Core library & API for the NKTg Law (Nguyen Khanh Tung). Includes core implementation, REST/gRPC API, and 150+ client wrappers
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    DeepSeek Math

    DeepSeek Math

    Pushing the Limits of Mathematical Reasoning in Open Language Models

    DeepSeek-Math is DeepSeek’s specialized model (or dataset + evaluation) focusing on mathematical reasoning, symbolic manipulation, proof steps, and advanced quantitative problem solving. The repository is likely to include fine-tuning routines or task datasets (e.g. MATH, GSM8K, ARB), demonstration notebooks, prompt templates, and evaluation results on math benchmarks. The goal is to push DeepSeek’s performance in domains that require rigorous symbolic steps, calculus, linear algebra, number theory, or multi-step derivations. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Patch-NetVLAD

    Patch-NetVLAD

    Multi-Scale Fusion of Locally-Global Descriptors for Place Recognition

    This repository contains code for the CVPR2021 paper "Patch-NetVLAD: Multi-Scale Fusion of Locally-Global Descriptors for Place Recognition".
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Safety-Prompts

    Safety-Prompts

    Chinese safety prompts for evaluating and improving the safety of LLMs

    ...The prompts are structured to test whether models generate outputs that align with human values and safety guidelines when faced with potentially harmful or sensitive requests. Researchers and developers use the dataset to benchmark how well models avoid unsafe responses and follow alignment constraints. The repository also serves as a training resource for improving model alignment by providing examples of prompts that require safe reasoning and appropriate refusal behavior. In addition to evaluation prompts, the project references related tools and benchmarks for assessing model safety across different contexts.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17

    LakhCleanAnalysis

    Tools to analyze the Lakh Clean Midi Dataset

    The goals of project are described in the web page https://lakhcleananalysis.sourceforge.io/ and in the YouTube video referenced here. Presently, we are analyzing the note onset distribution, the pitch class distribution, and the midi program assignments for the entire dataset. Each of these entities are represented by a separate vector for each midi file. These vectors were clustered using the kmeans and HDBSCAN algorithms. The vectors (midi files) were projected into a two dimensional space using the UMAP algorithm. A user interface, umapPlot.tcl is provided to explore this space.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    scikit-learn-videos

    scikit-learn-videos

    Jupyter notebooks from the scikit-learn video series

    scikit-learn-videos repository accompanies a video tutorial series designed to teach machine learning using Python’s scikit-learn library. It provides the Jupyter notebooks used in each lesson so learners can reproduce the demonstrations and experiment with the code themselves. The series introduces fundamental machine learning concepts such as classification, regression, model evaluation, feature engineering, and cross-validation using clear examples and real datasets. Each video...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 19
    seaborn

    seaborn

    Statistical data visualization in Python

    ...Its plotting functions operate on dataframes and arrays containing whole datasets and internally perform the necessary semantic mapping and statistical aggregation to produce informative plots. Its dataset-oriented, declarative API lets you focus on what the different elements of your plots mean, rather than on the details of how to draw them. Behind the scenes, seaborn uses matplotlib to draw its plots. For interactive work, it’s recommended to use a Jupyter/IPython interface in matplotlib mode, or else you’ll have to call matplotlib.pyplot.show() when you want to see the plot.
    Downloads: 9 This Week
    Last Update:
    See Project
  • 20
    Hiera

    Hiera

    A fast, powerful, and simple hierarchical vision transformer

    ...Documentation emphasizes that model weights may have separate licensing and that the code targets practical experimentation for both research and downstream tasks. Community discussions cover topics like dataset pretrains, integration in other frameworks, and comparisons with related implementations. Security and contribution guidelines follow Meta’s open-source practices, and activity shows ongoing interest and usage across the community.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 21
    GPT-2 Output Dataset

    GPT-2 Output Dataset

    Dataset of GPT-2 outputs for research in detection, biases, and more

    ...While no active development is expected, the dataset remains a useful benchmark for tasks involving text classification, style analysis, and generative model evaluation.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Downloads: 2 This Week
    Last Update:
    See Project
  • 23
    sketch

    sketch

    AI code-writing assistant that understands data content

    Sketch is an open-source AI-powered data analysis assistant designed specifically for pandas users, enabling natural language interaction with tabular datasets to generate code, insights, and transformations. It works by summarizing the structure and statistical properties of a dataset and providing that context to a language model, allowing it to generate highly relevant and accurate responses tailored to the data. The tool integrates directly into pandas dataframes through an extension, making it easy to use within existing Python workflows without requiring additional IDE plugins. Sketch supports a variety of tasks including data cleaning, feature engineering, visualization, and exploratory analysis, all driven by simple natural language prompts. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    DeepSeek LLM

    DeepSeek LLM

    DeepSeek LLM: Let there be answers

    ...According to the evaluation files, DeepSeek LLM 67B Chat achieves strong performance on math benchmarks under both chain-of-thought (CoT) and tool-assisted reasoning modes. The model is trained from scratch, reportedly on a vast multilingual + code + reasoning dataset, and competes with other open or open-weight models. The architecture mirrors established decoder-only transformer families: pre-norm structure, rotational embeddings (RoPE), grouped query attention (GQA), and mixing in languages and tasks. It supports both “Base” (foundation model) and “Chat” (instruction / conversation tuned) variants.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 25
    Bert-VITS2

    Bert-VITS2

    VITS2 backbone with multilingual-bert

    Bert-VITS2 is a neural text-to-speech project that combines a VITS2 backbone with a multilingual BERT front-end to produce high-quality speech in multiple languages. The core idea is to use BERT-style contextual embeddings for text encoding while relying on a refined VITS2 architecture for acoustic generation and vocoding. The repository includes everything needed to train, fine-tune, and run the model, from configuration files to preprocessing scripts, spectrogram utilities, and training...
    Downloads: 1 This Week
    Last Update:
    See Project