Showing 16 open source projects for "dataset"

View related business solutions
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Try It Free
  • 99.99% Uptime for MySQL and PostgreSQL Databases Icon
    99.99% Uptime for MySQL and PostgreSQL Databases

    Sub-second maintenance. 2x read/write performance. Built-in vector search for AI apps.

    Cloud SQL Enterprise Plus delivers near-zero downtime with 35 days of point-in-time recovery. Supports MySQL, PostgreSQL, and SQL Server.
    Try Free
  • 1
    PetoronAI-Drug-Discovery

    PetoronAI-Drug-Discovery

    PetoronAI Drug Discovery Experiments

    # PetoronAI Drug Discovery Benchmark https://github.com/01alekseev/PetoronAI PetoronAI was evaluated on the NCI-ALMANAC development dataset. • 2,225,137 experimental records • 602 experimentally measured drug pairs • Blind pair-level benchmark • 20×10 cross-validation • 99.5% validation coverage Validation: Correlation = 0.617327 MAE = 3.971799 Null model MAE = 4.656203 Sign accuracy = 64.69% Permutation p = 0.000100 Top HSA hypotheses: • Dactinomycin + Vinblastine sulfate (9.096) • Cabazitaxel + Vinblastine sulfate (8.224) • Mitoxantrone + Vinblastine sulfate (7.254) • Cabazitaxel + Mitoxantrone (7.128) • Dactinomycin + Daunorubicin HCl (6.965) Generated laboratory validation protocols (8×8 dose matrix, HSA, Bliss, Loewe, ZIP). ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    coronavirus

    coronavirus

    The coronavirus dataset

    The coronavirus package gives a tidy format dataset of the 2019 Novel Coronavirus COVID-19 (2019-nCoV) epidemic. Relevant and updated information about the virus, such as summary of new cases by country and total number of cases by region can be retrieved from this package. The raw data is pulled and arranged by the Johns Hopkins University Center for Systems Science and Engineering, which is gathered from various leading sources including the World Health Organization, China CDC, US CDC, European Centre for Disease Prevention and Control, and Australia Government Department of Health.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    sRNAWorkbench

    sRNAWorkbench

    The UEA sRNA Workbench

    A suite of tools for analysing small RNA (sRNA) data from Next Generation Sequencing devices. Including expression profiling of known mirco RNA (miRNA), identification of novel miRNA in deep-sequencing data and identification of other interesting landmarks within high-throughput genetic data
    Downloads: 11 This Week
    Last Update:
    See Project
  • 4

    MaChIAto Example Files

    The example files of MaChIAto

    MaChIAto (Microhomology-associated Chromosomal Integration/editing Analysis tools); a comprehensive analysis software that can precisely classify, deeply analyze, correctly align, and thoroughly review the targeted amplicon sequencing analysis data obtained by various CRISPR experiments, including template-free gene knock-out, short homology-based gene knock-in, and even a new-class CRISPR methodology, Prime Editing. In this repository, we provide the example files of MaChIAto. You can...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Veeam Data Platform v13.1 - Get Your Free Trial Icon
    Veeam Data Platform v13.1 - Get Your Free Trial

    Secure by design, portable by default. Recover clean, fast, anywhere. Start a free trial.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try it Free
  • 5

    Calis-p

    Estimates delta13C of species in a microbiome from proteome data

    Calis-p (The CALgary approach to ISotopes in proteomics) is a java application to estimate isotopic composition (e.g. delta13C or delta15N) of individual species in a microbial community from a proteomic dataset. Calis-p 2.0 handles both natural isotope abundances and data from labelling experiments such as stable isotope probing (SIP). It requires a mzIdent (or target spectrum match) and mzML files as the input and requires about 1 min per mzML file with 10 threads and needs <10 Gb of RAM. It has been tested with data from various nano liquid chromatography/Orbitrap platforms. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 6
    metasort

    metasort

    A metagenome assembler by reducing microbial community

    ...MetaSort provides a sorted mini-metagenome approach based on flow cytometry and single-cell sequencing methodologies, and employs new computational algorithms to efficiently recover high-quality genomes from the sorted mini-metagenome by the complementary of the original metagenome. Through extensive evaluations on simulated dataset, salivary and gut microbiomes, we demonstrated that metaSort has an excellent and unbiased performance on genome recovery and assembly.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 7
    PRIMUS

    PRIMUS

    Pedigree Reconstruction and Identification of a Maximum Unrelated Set

    ...Please visit the new website for the complete version of PRIMUS. We present a method adapted from graph theory that always identifies the maximum set of unrelated individuals in any dataset, and allows weighting parameters to be utilized in unrelated sample selection. PRIMUS reads in user-generated IBD estimates and outputs the maximum possible set of unrelated individuals, given a specified threshold of relatedness. Additional information for preferential selection of individuals may also be utilized. For example, when there are two equally sized maximum sets of unrelated individuals in a network, PRIMUS can preferentially select the set with more affected individuals. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 8

    FASTQSim

    NGS data characterization and in silico read generation

    FASTQSim is a tool that provides the dual functionality of Next-Gen Sequencing dataset characterization and metagenomic data generation. FASTQSim is sequencing platform-independent, and computes distributions of read length, quality scores, indel rates, single point mutation rates, indel size, and similar statistics for any sequencing platform. To create training or testing datasets, FASTQSim has the ability to convert target sequences into in silico reads with matching error profiles. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9

    GUDM

    A tool for pre-processing and fusing heterogeneous datasets

    Global Unified Data Modeler (GUDM) is a bioinformatics software tool used for pre-processing and integrating multiple heterogeneous datasets, collected from multi-modal sources, into an integrated dataset. This integrated dataset is supposed to be used for different types of medical analysis and unified decisions, using different machine learning approaches.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Try Free
  • 10
    CIG-P

    CIG-P

    CIG-P is a simple yet flexible data visualization tool

    ...CIG-P can be used to compare a) different AP-MS datasets of various baits or b) a particular bait under various perturbations (lenticular section CIG-P). The output of CIG-P is a simple and intuitively easy to grasp visualization of a complex dataset. Publication: CIG-P: Cicular Interaction Graph for Proteomics http://www.biomedcentral.com/1471-2105/15/344/ Previously known as PIVOT (Protein Interaction Visualization and Observation Tool)
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11

    hiovit-A

    hiovit-A is a simple yet flexible data visualization tool

    ...Hiovit-A can be used to compare a) different AP-MS datasets of various baits or b) a particular bait under various perturbations (lenticular section hiovit-a). The output of hiovit-A is a simple and intuitively easy to grasp visualization of a complex dataset. Previously known as PIVOT (Protein Interaction Visualization and Observation Tool)
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    FOCIS

    FOCIS

    FOCIS finds features for functional follow-up

    FOCIS (Feature Overlapper for Chromosomal Interval Subsets) performs an interval-based screen of a database of genomic features – ChIP-seq peaks, motif matches, and others – for overlap enrichment at a specific subset of genomic regions relative to a dataset-matched background. It was recently used to discover a novel enhancer that mediates drug resistance in melanoma.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 13

    ARDEN

    Specificity Control for Read Alignments Using an Artificial Reference

    We introduce ARDEN (Artificial Reference Driven Estimation of false positives in NGS data), a novel benchmark that estimates error rates based on real experimental reads and an additionally generated artificial reference genome. It allows the computation of error rates specifically for a dataset and the construction of a ROC-curve. Thereby, it can be used to optimize parameters for read mappers, to select read mappers for a specific problem or also to filter alignments based on quality estimation.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14

    CRUMp

    A probabilistic prediction system of protein phosphorylation sites

    ...Given an input set of protein sequences in FASTA format, the system outputs the position, residue type (S, T, or Y), and the estimated probability of each tested site being phosphorylatable. Latest downloadable files: - crump-0.2.0.tar.gz: CRUMp GNU Octave package - crump-0.2.0.zip: CRUMp MATLAB script - crumptestset.fasta: A testing dataset in FASTA format. The sequence headers list the accession number of the protein sequence and the position numbers of known phosphorylation sites. Note that CRUMp may predict additional phosphorylation sites that have not been experimentally verified yet. The testing dataset is from Biswas et al. 2010, http://www.biomedcentral.com/1471-2105/11/273/additional.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    KGP is a program that reveals the KIR (Killer-cell Immunoglobulin-like Receptors (KIR)) genotypic diversity within a dataset using binary coded KIR genotypic patterns generated by the presence and absence of 16 KIR genes on a diploid chromosome.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    FastPval is multiple stage p-value computing software that computes empirical p-values from a large set of permutated/resampled background data.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • Next