遇见数据集

TABLE 2 in BlasTax-a user-friendly stand-alone tool to leverage the BLAST + program for molecular taxonomy

收藏
Zenodo2026-07-13 更新2026-08-01 收录
官方服务:

资源简介:

TABLE 2. Overview of the modes available in BlasTax (each available from a tile in the home window of the program). Program modeMain functionOptionsCommentsMake BLAST DatabaseMakes a local BLAST database from a FASTA file, based on the makeblastdb executable.Protein or nucleotide. Batch mode: multiple databases can be consecutively created.GUI-driven use of basic local BLAST function.Regular BLASTPerforms regular BLAST searches.Implements the blastn, blastp, tblastn, blastx, tblastx executables.GUI-driven use of basic local BLAST function.BLAST-AppendSearches for sequences similar to query and appends the matches to the query FASTA file.Batch mode for reference and query. Single best matches or multiple matches can be appended.Targeted to add homolog sequences of closely related taxa, from new transcriptome assemblies or annotated genomes, to pre-existing phylogenomic alignments.BLAST-Append-XSearches for sequences similar to query based on protein sequences and appends the nucleotide sequences corresponding to the matches to the query FASTA file.Batch mode for query. Similar to BLAST-Append but BLAST performed on additional query file with translated protein sequences.Experimental mode. Useful to add sequences of less closely related taxa where nucleotide BLAST may not find reliably matching sequencesDecontaminateBLAST of query sequences to an ingroup and an outgroup database and sorting into two new FASTA files depending on best match.Batch mode for query. Comparison can be done based on sequence similarity (pident), sequence length, or bitscore.Useful to decontaminate sequence files that may contain contaminants (e.g., from prey, symbionts, bacteria, human).MuseoscriptCompares a FASTQ or FASTA file with reference and writes matches into new FASTA file.Minimum sequence identity for matches to be parsed can be specified. Either only the aligned BLAST match is parsed, or the entire sequence containing the match.For analysis of high-throughput data from archival DNA sequencing, to extract the (often few) matches to the target sequence for subsequent assembly/alignment.Assign taxonomyUses a BLAST database containing a NCBI taxID and the NCBI taxDB database to assign taxonomic identity to query sequences based on BLAST matches.Can output a FASTA file with query sequences annotated based on best BLAST matches, as well as a table of best matches per query sequences, and a summary table with counts of matches per reference sequence.Summary table represents a simple version of the standard output of specialized programs for DNA metabarcoding analysis.Database operationsVarious data extractions and conversions related to BLAST databases.Extraction of a FASTA file or a taxID mapping file from a BLAST database and other conversions.Useful to prepare input files for downstream analyses.FastMergeMerges several sequence files into one file.Can process FASTA and FASTQ, as well as gzip compressed files.Basic utility for preparation of files to be used as query or reference in BLAST searches.FastSplitSplits large sequences or text files into smaller files.Can process FASTA and FASTQ, as well as gzip compressed files.Basic utility for preparation of files to be used as query or reference in BLAST searches.GroupMergeMerges FASTQ or FASTA files based on a part of their filename, and subsequently merges all sequences of each group into a single FASTA file.Upon deduplicating sequence identifiers, can either keep all of them or only the first occurrence.Sorts the output produced by subsequent runs of other program modes.FastPrepareSanitizes and trims sequence identifiers in FASTA files.Batch mode is available. Search-replace, trimming, sanitizing (removal of special characters) in all sequence identifiers of a FASTA file.Important to prepare FASTA files for making BLAST databases which do only accept sequence identifiers up to 50 characters and no special characters.Stop codon removalSearches for stop codons in FASTA files with coding sequences and removes them.Batch mode is available. Sequences are either trimmed after and including the stop codon, or the entire sequence removed, or the entire FASTA file deleted.Detects and/or removes erroneous sequences or stretches of sequences before analysis.Codon trimmingTakes a FASTA file with coding sequences and trims all sequences to start on first codon position.Autodetects reading frame based on the lowest number of stop codons. Different options of how to deal with remaining stop codons, if these are found.Preparation step for codon-aware multiple sequence alignment, stop codon removal and other downstream applications.Protein translatorTranslates nucleotide sequences into protein (amino acid) sequences.Batch mode is available. Translation table and reading frame user defined. Auto-detect reading frame is available. Transcript mode extract longest open reading frame.Utility for preparation of files to be used as query or reference in BLAST searches.Codon-aware alignmentTranslates coding sequences, then aligns the protein sequence, and adjusts the nucleotide alignment accordingly.Requires all sequences to start with the first codon position.More accurate alignment of coding sequences, especially when sequences are very variable and including indels.SCaFoS-PySelects and/or fuses sequences belonging to the same species in one FASTA file.Batch mode is available.Detects species based on a part of the sequence identifier. Either selects the longest sequence per species to keep, or the sequence with highest similarity to other sequences in file, or fuses sequences using IUPAC ambiguity codes where overlaps do not match.Downstream processing step of sequence files obtained with BLAST-Append. Most functions require sequences to be aligned. ......continued on the next page

提供机构:
Zenodo
创建时间:
2026-07-13
二维码
社区交流群
二维码
科研交流群
商业服务