Showing 102 open source projects for "fasta"

View related business solutions
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Start Free
  • 1

    mfsizes

    Multi-FASTA sequence (DNA or protein) statistics calculator.

    A simple command-line utility to calculate biological sequence (DNA or protein) sizes in a (multi) FASTA file. It gives averages, GC (or methionine) content, N50, N90, N95, number of N's, and total bases, and can also report by codon if requested.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2

    Genome Downloader

    Downloads genome data from NCBI based on search terms.

    GenomeDownloader is a command-line Perl program to download genomic data (using wget) from NCBI. It has been recently (2017-10) completely rewritten to work with the "new" data organization structure at NCBI. Assembly completion level (i.e., Contig, Scaffold, Chromosome or Complete Genome) can also be selected as a criterion for downloading data. Genomic data can be downloaded from all organisms belonging to a certain taxon (e.g., Mammalia or 40674), and downloads can be limited to...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3

    Genome Profiler

    a wgMLST analysis tool for bacterial WGS data

    ...Please download the latest version from: https://github.com/jizhang-nz Update 4th Oct. 2017: version: 2.1 bugs fixed. Update 8th Sep 2015: When using a multi-Fasta file of the the allele sequences (nt) as reference (switch -n), please use only CAPITAL letters A, T, G and C for the sequences. This bug will be fixed in the next version.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4

    CRISPR-offinder-v1-2

    A CRISPR tool for user-defined protospacer adjacent motif

    ...However, Cas9 from different types of bacteria or variant recognizes different PAM sequences. To meet the needs of different CRISPR system with specific and efficient sgRNA design, CRISPR-offinder was developed. Given an input FASTA file of the target sites and queries the reference genome as well as a CRISPR system with a defined spacer length and PAM sequence, this standalone tool will identify putative sites and assign a predicted activity based on support vector machine model which conducted by sgRNA Scorer 2.0. In addition, sgRNAs with minimal off-target activity were predicted by Cas-OFFinder, and score with Off-Target Cutting Frequency Determination (CFD).
    Downloads: 0 This Week
    Last Update:
    See Project
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • 5
    ...It's simple linux program which evaultes the genome assembly with high speed and accuracy. Please read instuctions.md [Options] Argument 1 -> Name of the input fasta/fastaq file Argument 2 -> Sequence Limit (optional)(Default: 999999999) Argument 3 -> Usable scaffold length (optional)(Default: 2500) Argument 4 -> Number of Nucliotide to spit into contigs (optional)(Default: 25) Argument 5 -> Genome size (optional)
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    BioFace is simple software for editing and analyzing DNA, RNA, and protein sequences that is written in Java using SWT and JFace as libraries. Opening GenBank, FASTA, EMBL, or simple sequence files and analizing these sequences can be done.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7

    BioUtils Perl Library

    A collection of Perl modules for handling fasta/q sequences and files.

    WARNING: BioUtils has been migrated to Github (Nov 2017). For the most up-to-date versions and info please visit: https://github.com/islandhopper81/BioUtils BioUtils are a collection of Perl modules for DNA sequence analysis in bioinformatics. BioUtils is a significantly faster and more memory efficient alternative to BioPerl. However, it's functionality is currently limited to the features listed below.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8

    MSTgold

    Estimate minimum spanning trees with statistical bootstrap support

    ...The MSTgold package includes Mac OS X, Linux, and Windows executables of the MSTgold program, a detailed Manual, example data and results, and executables of the program Fasta2MSTG which converts Fasta sequence files to the MSTgold input format.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 9
    PBSuite

    PBSuite

    Software for Long-Read Sequencing Data from PacBio

    .... ----- PBJelly ----- Read The Paper http://www.plosone.org/article/info%3Adoi%2F10.1371%2Fjournal.pone.0047768 PBJelly is a highly automated pipeline that aligns long sequencing reads (such as PacBio RS reads or long 454 reads in fasta format) to high-confidence draft assembles. PBJelly fills or reduces as many captured gaps as possible to produce upgraded draft genomes. ----- PBHoney ----- Read The Paper http://www.biomedcentral.com/1471-2105/15/180/abstract PBHoney is an implementation of two variant-identification approaches designed to exploit the high mappability of long reads (i.e., greater than 10,000 bp). ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Train ML Models With SQL You Already Know Icon
    Train ML Models With SQL You Already Know

    BigQuery automates data prep, analysis, and predictions with built-in AI assistance.

    Build and deploy ML models using familiar SQL. Automate data prep with built-in Gemini. Query 1 TB and store 10 GB free monthly.
    Start Free
  • 10
    Kinannote

    Kinannote

    Protein Kinase Identification and Classification

    Kinannote identifies and classifies protein kinases in a user-provided fasta file using an HMM derived from serine/threonine protein kinases, a position specific scoring matrix derived from the HMM, and comparison with a local version of the curated kinase database from kinase.com. If the user inputs a complete proteome, additional modules evaluate the completeness of the kinome and place it in context with reference kinomes.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11

    PrAS

    predict protein amidation sites

    This predictor is developed to predict amidation sites based on support vector machine (SVM) classifier. It is supplied in source code form along with the required data files and run under the linux. The input is a protein sequence file (fasta format) by Tong Wang and Wei Zheng (tongwang.scu@gmail.com and jlspzw139@sina.com) Notice:You should download all zip file in this project!
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12

    GenomeDatabase

    For creating a local database of reference genomes

    Genome Database - A tool to create a local database of reference genome sequences Usage: java path/to/GenomeDatabase.jar [options] By Marc Strous, 2016 This tool enables you to download fasta files of protein and RNA sequences encoded in reference genomes at NCBI. You can select relevant genomes with a set of queries. Each query has four fields, separated by comma's. Example of a queries are: superkingdom,Bacteria,genus,ftp superkingdom,Archaea,genus,ftp superkingdom,Eukaryota,phylum,ftp superkingdom,Viruses,family,elink The first query would download (with ftp) all available reference genomes of the superkingdom Bacteria, limited to one genome per genus. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    metawatt

    metawatt

    Binner for assembled metagenomes

    The Metawatt binner is a graphical binning tool that makes use of multivariate statistics of tetranucleotide frequencies and differential coverage based binning. It also performs taxonomic assessment of binning quality (via diamond BLASTx). Created bins can be edited and exported as fasta. The Metawatt is implemented in Java SWING and minimally depends on Diamond, HMMer3.1, BBMap, Prodigal and the Batik library for the export of SVG graphics. Citation: Strous M, Kraft B, Bisdorf R, TegetMeyer H (2012) The binning of metagenomic contigs for microbial physiology of mixed cultures. Frontiers in Microbial Physiology and Metabolism 3:410. doi: 10.3389/fmicb.2012.00410
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    converts a SAM file to fasta file. SAM file is a file output from bwa alignment software. It outputs aligned fasta file.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15

    KA-predictor

    lysine acetylation site prediction

    This predictor is developed to predict species-specific lysine acetylation sites based on support vector machine (SVM) classifier. It is supplied in source code form along with th e required data files and run under the linux. The input is a protein sequence file (fasta format).
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16

    RGAAT

    Reference based genome assembly and annotation for new genome

    This program can assemble and/or annotate genome for new genome and known genome upgrade using sequence alignment file (SAM or BAM format), sequence variant file (VCF format or five coloum table (tab-delimited, including chromosome, position, id, reference allele and alternative allele)) or new genome sequence file (FASTA format) based on reference genome sequence file (FASTA format) and annotation file (TBL, GTF, GFF, GFF3 or BED format).
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17

    gDNA-Prot

    DNA-binding protein prediction software

    This program will predict the DNA-binding proteins when input a protein sequence (Fasta format), it is supplied in source code form along with the required data files and run under the windows.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Mauve computes and interactively visualizes genome sequence comparisons. Using FastA or GenBank sequence data, Mauve constructs multiple genome alignments that identify large-scale rearrangement, gene gain, gene loss, indels, and nucleotide substutit
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19

    JFASTA

    Java implementation for the FASTA file format.

    JFASTA is a lightweight framework for handling FASTA files. It supports reading, writing and parsing of single- or multi-FASTA files. Files may be read/ written at once, or stream-based (memory efficient). It's stable, very intuitive and good integrated with Java 1.5 SDK and later.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 20

    CATO Clone Alignment Tool

    Identifies clone sequences corresponding to set of reference sequences

    Specialized sequence alignment software that helps associate the most likely matches of clone sequences to a set of reference sequences. Both sets of sequences are assumed to be nucleotides.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Protospacer Workbench

    Protospacer Workbench

    CRISPR/Cas9 guide-RNA design

    Protospacer Workbench helps to design, analyze, and share CRISPR target-sites for any organism or set of FASTA formatted sequences. Design of guide-RNAs for the CRISPR/Cas genome editing system is intuitively easy, but computationally difficult. The main difficulty arises from the need to identify potential off-targets that may be quite different from the intended target. Current online tools for guide-RNA design provide a user friendly interface to sequence mapping software such as Bowtie or BLAST. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22

    PanCGP

    PanCGP: Pangenome and Comparative Genome Analysis Pipeline tool

    ...Those results are interpreted in the form of pan-genome, core genome/proteins, dispensable genes/proteins, unique/strain-specific genes/proteins, new gene/protein families and total number of genes per strain of the specie. It also extracts the corresponding protein sequences against pan-genome, core protein and dispensable protein sequences in separate FASTA formatted text files. The tabular text format can be used to render the final results in the form of visual graphs representing the trends in gene/protein sequence analysis.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23

    ArtificialFastqGenerator

    Ouputs artificial FASTQ files derived from a reference genome.

    ArtificialFastqGenerator takes the reference genome (in FASTA format) as input and outputs artificial FASTQ files in the Sanger format. It can accept Phred base quality scores from existing FASTQ files, and use them to simulate sequencing errors. Since the artificial FASTQs are derived from the reference genome, the reference genome provides a gold-standard for calling variants (Single Nucleotide Polymorphisms (SNPs) and insertions and deletions (indels)).
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    LociMapViz - multi loci visualization
    Software for visualization of point variability in pairwise sequence alignment FASTA format. Featuring GUI interface, this simple application enables insight into variation of nucleic and amino acids on specific loci. Current 1.1 version supports amino and nucleic acid alignments. There are 2 variations of this software. One is CLI based and the other one is GUI based. Read corresponding README files in order to get familiar with the software.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    iMSAT
    iMSAT represents a full update of the vcf2MSAT program. This command line python program allows for a user to use the polymorphism data generated using SAM- and BAM-tools and a .fasta alignment file to search for polymorphic microsatellite markers (MSATs or STRs). By identifying polymorphic makers, rather than simple repeat regions as previous programs have done, iMSAT greatly increases the speed at which polymorphic MSATs that can be identified -- saving researchers precious time and money. Visit http://www.biomedcentral.com/1471-2164/15/858/abstract for the pdf article describing the utility of iMSAT and if you use the program please cite this article as: Andersen and Mills: iMSAT: a novel approach to the development of microsatellite loci using barcoded Illumina libraries. ...
    Downloads: 0 This Week
    Last Update:
    See Project