We evaluated six tools including SOAPfuse (v1.22) based on both actual and simulated datasets. This comparison work is also mentioned in SOAPfuse Method Paper.
-
Released datasets
We used released RNA-Seq data, downloaded from NCBI SRA, from two published researches (see References) as actual datasets. This two studies discovered some validated gene fusions based on their RNA-Seq data, which are specified with Sanger sequences in their supplementaries. One is concerned with melanoma and CML (dataset A, 15 fusions), and another one is breast cancer (dataset B, 27 fusions). All information of validated gene fusions are based on Ensemble Release59 of hg19.
We applied SOAPfuse (based on Ensemble Release59, hg19-GRCh37.59) to analyze the downloaded RNA-Seq data, and compared the results of other five published tools. Comparison is shown as below.

For dataset A, which contains ~111 million paired-end reads, SOAPfuse consumed the least CPU time (~5.2 hours) and the second least memory (~7.1 Gigabytes) to complete the data analysis (including the alignment of reads against reference), and was able to detect all the 15 fusion events. DeFuse and FusionHunter detected comparable number of known fusion events (12~13 of the 15 fusions), but took 82.1 and 21.3 CUP hours, respectively, at least four times as much as SOAPfuse. The computational resource cost of SnowShoes-FTD was comparable with SOAPfuse, but SnowShoes-FTD only identified eight of 15 events. The remaining two tools, chimerascan and TopHat-Fusion, detected four confirmed fusion events but used significantly more CPU hours or memory usage. For dataset B containing ~55 million paired-end reads, SOAPfuse detected 26 of the 27 reported fusion events with 4.1 CPU hours and 6.3 Gigabytes memory. The other five tools were able to identify comparable numbers of reported fusions (15~21) and cost at least 6.4h CPU time. As we can see, SOAPfuse shows the best performance in all three aspects.
To get configs (parameters) and results of all tools about downloaded datasets, please click here.
-
Simulated datasets
For simulated datasets, we generated a set of paired-end reads (2 x 75 nt) based on the transcriptome of human (hg19, Homo_sapiens, Ensemble Release59). We simulated 150 fusions based on some criterions, and generated PE-reads at 5-, 10-, 20-, 50-, 80-, 100-, 150x and 200x fold sequencing depth (to imitate different expression levels) using the short-read simulator provided by MAQ (Li et al., 2008).
And then, we mixed simulated reads of each fold with cleaned background data (BG). BG is downloaded from NCBI Sequence Read Ar-chive (SRA) under accession NO. SRR065491 and SRR066679, which were generated by the ENCODE Caltech RNA-Seq project (Birney, et al., 2007; Raney, et al., 2010). It is RNA-Seq data from embryonic stem cells, and also used as background by FusionMap (H. Ge, et al., 2011).
Chimerascan, FusionHunter and SnowShoes-FTD only detect cases fused at edge of exon, considering not all simulated cases are exon-edge type, we abandoned comparing this three tools. Several strategies are applied to achieve fair and conservative comparison. Combining all results, 149 (99%) are detected, and 142 (94%) are confirmed by at least two tools, proving our simulation is available. Further to be prudent, compares are operated based on these 142 simulated cases for their ratification by at least two algorithms. Comparision is shown as below.

As expected, FN rates decreased with increasing expression levels of fusion transcripts (a). SOAPfuse and deFuse achieved the lowest FN rates at 5% with fusion transcript expression levels of 30-fold or greater. TopHat-Fusion had higher FN rates, especially at low fusion transcript expression levels (5~20-fold). For FP rate (b), only SOAPfuse achieved < 5% at different fusion transcript expression levels, while deFuse and TopHat-Fusion had higher FP rates at lower fusion transcript expression levels. SOAPfuse (v1.22) missed 3 simulated fusions which are detected by both deFuse and TopHat-Fusion (c), revealing a weakness in analysis of homologous gene sequences and short fusion transcripts of long genes. We have fixed it from v1.24.
Generally, lower FN rates and lower FP rates are contradictory for detection of fusions, however, SOAPfuse and deFuse are good at reducing FN and FP rates during fusion transcript identification. In summary, SOAPfuse showed optimal performance with low FN and FP rates at different expression levels of fusion transcripts.
To get simulated datasets, click here.
To get configs and results of all tools, click here.
-
Cell lines datasets
We also sequenced paired-end RNA-Seq reads for two bladder cancer cell lines, and applied SOAPfuse on it with some criterions. SOAPfuse identified a total of 16 fusions, all of which are intrachromosomal and fused at exon-edge. We designed primers for RT-PCR experimental validation of all predicted fusions, and Sanger sequences confirmed 15 genuine fusion events (93% validation rate), in which 6 pairs are novel and shared by this two cell lines. There are some validated fusions that may be caused by chromosomal rearrangements on genome with strong signals. Some are formed by genes from different strands that imply potential inversions, and some fusions are formed by same strand genes with their reversed genomic orientation. RNA-Seq data from the two bladder cancer cell lines has been submitted to NCBI SRA and is available under accession number [SRA052960].