SimplyTheBlast: a small perl tool to build genes presence/absence matrices over a set of Fasta formatted genomes.

This code requires:
Bio::SeqIO;
Bio::Perl;
Bio::Tools::Run::StandAloneBlast;
Bio::Seq;
Bio::Tools::Blast;
Bio::DB::GenBank;
Bio::DB::WebDBSeqI;

and BLAST 2.2.28 (blastall and formatcmd) installed and reachable from your command line

Usage: perl SimplyTheBlast-Align.pl <fasta formatted seeds file> <path to genomes folder> <Alignment length threshold in %> <Alignment identity threshold in %>

OR

Usage: perl SimplyTheBlast-Evalue.pl <fasta formatted seeds file> <path to genomes folder> <Evalue threshold>

Genomes files names must end with *.faa

Output files:
TABULAR_FBH_OUTPUT.xls is an Excel readable file with the identifier of the best hits found
TABULAR_FBH_OUTPUT.csv is a file with the number of the best hits found
query_n* files are fasta formatted files with the sequences of the best hits found
bugs /comments:
marco.fondi@unifi.it

Project Activity

See All Activity >

Follow SimplyTheBlast

SimplyTheBlast Web Site

Other Useful Business Software
$300 Free Credits for Your Google Cloud Projects Icon
$300 Free Credits for Your Google Cloud Projects

Start building on Google Cloud with $300 in free credits. No commitment, no credit card required until you're ready to scale.

Launch your next project with $300 in free Google Cloud credits—no strings attached. Test, build, and deploy without risk. Use your credits across the entire Google Cloud platform to find what works best for your needs. After your credits are used, continue with always-free tier services. Only pay when you're ready to scale. Sign up in minutes and start exploring.
Start Free Trial
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of SimplyTheBlast!

Additional Project Details

Registered

2014-11-14