Showing 643 open source projects for "dataset"

View related business solutions
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • Build Securely on AWS with Proven Frameworks Icon
    Build Securely on AWS with Proven Frameworks

    Lay a foundation for success with Tested Reference Architectures developed by Fortinet’s experts. Learn more in this white paper.

    Moving to the cloud brings new challenges. How can you manage a larger attack surface while ensuring great network performance? Turn to Fortinet’s Tested Reference Architectures, blueprints for designing and securing cloud environments built by cybersecurity experts. Learn more and explore use cases in this white paper.
    Download Now
  • 1

    Wunderfixer

    Supplies missing data to the Weather Underground

    Given a weewx or wview SQLITE dataset, this utility compares timestamps against the Weather Underground. If any are missing, it republishes the data to the WU.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    KACST Arabic Phonetics Database This dataset is provided by KACST http://www.kacst.edu.sa
    Downloads: 1 This Week
    Last Update:
    See Project
  • 3
    ExData Plotting1

    ExData Plotting1

    Plotting Assignment 1 for Exploratory Data Analysis

    This repository explores household energy usage over time using the “Individual household electric power consumption” dataset from the UC Irvine Machine Learning Repository. The dataset covers nearly four years of minute-level measurements, including power consumption, voltage, current intensity, and detailed sub-metering values for different household areas. For analysis, focus is placed on a two-day period in February 2007, highlighting short-term consumption trends. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4

    pmutualinformation

    Computes pairwise genes mutual information using GEPs

    pmutualinformation (parallel mutual information) computes the pairwise Mutual Information for all pairs of genes from a potentially massive and heterogeneous dataset containing GEPs. GEPs can be part of different experiments collected from public repositories. pmutualinformation runs on MIMD systems and has been implemented in C using the MPI standard.
    Downloads: 0 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 5
    phpMyForms is a php script which allows the user to create simple database listings (done in 3 lines of configuration), or simple dataset manipulation formulars (~10 lines) by adjusting small php config files.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6

    Pepgrep

    Tool for peptide MS2 pattern matching in MS tandem dataset

    pepgrep offers a quick and friendly way to examine MS tandem output data by finding a spectral pattern matching selected peptide sequence. It is database-free command-line utility that uses Peptide-to-MS2 scoring algorithm. The Peptide-to-MS2 pattern matching can accumulate dozens of mass offsets corresponding to a variety of Post-Translational Modifications. Decoy peptide sequences are used with the tested peptide sequence to reduce false-positive results. Read article at:...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    Calculates Natural Product(NP)-likeness of a molecule, i.e. the similarity of the molecule to the structure space covered by known natural products. NP-likeness is a useful criterion to screen compound libraries and to design new lead compounds. Maven dependancy: <dependency> <groupId>uk.ac.ebi.cheminformatics</groupId> <artifactId>NP-Likeness</artifactId> <version>2.1</version> </dependency> Required repository: <repositories> ...
    Downloads: 5 This Week
    Last Update:
    See Project
  • 8

    hiovit-A

    hiovit-A is a simple yet flexible data visualization tool

    ...Hiovit-A can be used to compare a) different AP-MS datasets of various baits or b) a particular bait under various perturbations (lenticular section hiovit-a). The output of hiovit-A is a simple and intuitively easy to grasp visualization of a complex dataset. Previously known as PIVOT (Protein Interaction Visualization and Observation Tool)
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9

    Cost-sensitive Classifiers

    Adaboost extensions for cost-sentive classification

    ...Minimum expected cost criteria Input also requires to load an arff file and a cost matrix (sample arff and cost files are uploaded for users' reference) This extension uses weka for classification and generates the classification model along with confusion matrix. For given dataset and cost matrix
    Downloads: 0 This Week
    Last Update:
    See Project
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Try It Free
  • 10

    QSAR Dataset Division GUI

    Dataset division GUI is a user friendly QSAR dataset division tool

    The purpose of this application tool is to perform rational selection of training and test set using Kennard Stone algorithm, diversity based and activity based division method.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    miRStress

    miRStress

    A desktop viewer for the miRStress database

    A python based program to allow users to interrogate the miRStress dataset off-line. miRStress project conceived and overseen by Dr. Dave RF Carter and Laura A Jacobs. miRStress python file written by Dr. Mark Poolman and Findlay Copley (this is also what powers the web site). Tkinter scripts written by Findlay Copley Publications - Jacobs LA, et al. Meta-analysis using a novel database, miRStress, reveals miRNAs that are frequently associated with the radiation and hypoxia stress-responses. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    FOCIS

    FOCIS

    FOCIS finds features for functional follow-up

    FOCIS (Feature Overlapper for Chromosomal Interval Subsets) performs an interval-based screen of a database of genomic features – ChIP-seq peaks, motif matches, and others – for overlap enrichment at a specific subset of genomic regions relative to a dataset-matched background. It was recently used to discover a novel enhancer that mediates drug resistance in melanoma.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 13
    Integrated Pipeline for Genome-Wide Association Studies
    Downloads: 1 This Week
    Last Update:
    See Project
  • 14

    Multilabel Dataset GUI

    An application for evaluating metric in Multilabel dataset

    Multilabel Dataset GUI is a software for evaluating metric in Multilabel dataset. In this project we try to collect as much information as possible for characterizing multilabel dataset. This application is closely based upon the Weka and MULAN framework.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    ogiga | Hydrail is an Object-Relational-Mapping for PHP 5.x
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    VCL.JS

    VCL.JS

    TypeScript component based framework for enterprise web application

    ...//Simple dbgrid bounded to a query import V = require("VCL/VCL"); export class PageHome extends V.TPage { constructor() { super(); //create a backend query var qur = new V.TQuery(this); qur.SQL = "SELECT CustomerKey, FirstName, LastName FROM Customers"; qur.open(); //create a grid on the screen var grd = new V.TDBGrid(this, "grid"); grd.Dataset = qur; //bind the grid to the dataset grd.PageSize = 15; var col = grd.createColumn(“FirstName”); var col = grd.createColumn(“Lastname”,”Last Name”); } }
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17

    TextProcessor

    A Java package to preprocess text datasets for posterior text analysis

    The TextProcessor Java package is a text processing toolkit, which provides some frequently used text processing functions such as stemming, removing stop-words, generating a term vocabulary, and calculating the term-doc frequency matrix. Basic topic mining models such as LDA and sparse NMF are also supported. The package can also generate feature files from a given text dataset with LDA and LIBSVM format for posterior procedures such as classification or clustering. The toolkit is also being extended for more advanced text analysis tasks based on natural language processing techniques.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 18

    Qudaich

    Performs local sequence alignment for NGS data

    Qudaich (queries and unique database alignment inferred by clustering homologs) is a software package for aligning sequences. Qudaich generates the pairwise local alignments between a query dataset against a database. The main design purpose of qudaich is to focus on datasets from next generation sequencing. These the datasets generally have hundreds of thousand sequences or more, and so, the input database should contain large number of sequences. Qudaich is flexible and its algorithmic structure imposes no restriction on the absolute limit of the acceptable read length, but the current version of qudaich allow read length <2000 bp. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 19
    cd-hit

    cd-hit

    Ultrafast program for clustering large set of biological sequences

    ...CD-HIT helps to significantly reduce the computational and manual efforts in many sequence analysis tasks and aids in understanding the data structure and correct the bias within a dataset. The CD-HIT package has cd-hit, cd-hit-2d, cd-hit-est, cd-hit-est-2d, cd-hit-454, psi-cd-hit, cd-hit-otu, cd-hit-lap, cd-hit-dup and over a dozen scripts for various clustering needs.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 20

    X-Freq

    X-Ray Reflectivity Frequency Analysis

    Analyze the thicknesses of a sample from an x-ray reflectivity dataset using Fourier Transform techniques.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21

    MS-AND

    Identifies differentially methylated regions (DMRs)

    Our approach, methylation status mask AND (MS-AND), uses bit operations and masking and can be applied to any microarray dataset in General Feature Format (GFF).
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    This dataset deals with finding potential retweeters given a particular tweet. It contains 500 tweets and 40574 judged followers.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Document Analysis and Exploitation
    The Document Analysis and Exploitation Platform is a Drupal based web interface to a cloud enabled Document Analysis resource set.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24

    PRICE de novo (meta)genome assembler

    Targeted de novo assembly of (meta)genomic components

    ...Its name describes the strategy that it implements for genome assembly: PRICE uses paired-read information to iteratively increase the size of existing contigs. Initially, those contigs can be individual reads from a subset of the paired-read dataset, non-paired reads from sequencing technologies that provide non-paired data, or contigs that were output from a prior run of PRICE or any other assembler. PRICE was designed to address the challenge of assembling viral genomes that comprise a small minority of the reads within ultra-deep, short-read, shotgun metagenomic datasets. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25

    ARDEN

    Specificity Control for Read Alignments Using an Artificial Reference

    We introduce ARDEN (Artificial Reference Driven Estimation of false positives in NGS data), a novel benchmark that estimates error rates based on real experimental reads and an additionally generated artificial reference genome. It allows the computation of error rates specifically for a dataset and the construction of a ROC-curve. Thereby, it can be used to optimize parameters for read mappers, to select read mappers for a specific problem or also to filter alignments based on quality estimation.
    Downloads: 0 This Week
    Last Update:
    See Project