# Reference Alignments: Best Way to Download Many Sequences from NCBI GenBank

**URL:** https://forum.qiime2.org/t/reference-alignments-best-way-to-download-many-sequences-from-ncbi-genbank/19824
**Category:** General Discussion
**Tags:** taxonomy, rescript, feature-table
**Created:** [June 7, 2021, 9:40pm UTC](https://forum.qiime2.org/t/reference-alignments-best-way-to-download-many-sequences-from-ncbi-genbank/19824 "2021-06-07T21:40:29Z")
**Posts on this page:** 1
**Showing post:** 2

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [June 8, 2021, 6:32am UTC](https://forum.qiime2.org/t/reference-alignments-best-way-to-download-many-sequences-from-ncbi-genbank/19824/2 "2021-06-08T06:32:47Z")

</div>

Welcome to the forum @alexkrohn !

Sounds like you are on the right track with RESCRIPt!

> [@alexkrohn](#):
>
> Would it be better/faster to just run one giant search for each taxon separated by OR in RESCRIPt?

I'd go for the latter. Would be slow to loop 400 times, also slow to merge. Also your provenenance graph would look really ugly! 😝

> [@alexkrohn](#):
>
> Or faster to just use entrez, then format the sequence/taxonomy data to import into Qiime2?

Probably not faster.

> [@alexkrohn](#):
>
> It seems like merge and merge-seqs will be the best tools to concatenate the Feature[Sequence] and Feature[Taxonomy] files together. Is that right?

merge-seqs is correct, but merge will merge feature tables, not FeatureData[Taxonomy] artifacts. There is a separate `merge-taxa` action (one in RESCRIPt with more complex functionality, one in q2-taxa that has simpler functionality).

> [@alexkrohn](#):
>
> it might just be best to download all animal sequences for the gene of interest (e.g. one for COI, one for 16S etc.), then filter, and align accordingly?

Depending on your use case, this might be the best approach overall (and RESCRIPt can be used to download the entire query), minus the filtering part (if you must filter, I recommend the restricted query approach above). For most applications you would want a complete database, not only the 400 taxa of interest, so that you can get more reliable estimates of confidence for taxonomic classification, clustering, etc, and avoid issues with misclassification. Also I realize that you are focusing on animals, but it helps to include in the database any non-targets that are amplified by the same primers.

On the other hand, I also realize that for eDNA surveys it can be useful to take geographic range information into account, but building a really restricted database always makes me nervous since it is making the bold assumption that _only those 400 species exist in your geographic range_. I have been curious about using geographic [species distribution data](https://www.nature.com/articles/s41467-019-12669-6) with [q2-clawback](https://forum.qiime2.org/t/using-q2-clawback-to-assemble-taxonomic-weights/5859) to improve taxonomic classification with COI and other eDNA markers... the question is where/how to get species frequency information for those markers. You could put together this frequency information artificially, so that you give all species you expect to find a 1, and all other species in the database a 0, then convert to relative frequency and pass that table to q2-clawback for training your classifier. The disadvantage is that this would be an arduous process, but the advantage is that you would shed this assumption that _only those 400 species can be observed_... rather you would say that _these 400 species are much much more likely to be observed_ but you are keeping an open mind to the strange things that happen in the world 😁 so in other words: your classifier would not be prone to misclassification or other issues.

To save you some time, here are some tutorials using RESCRIPt for SILVA, NCBI, and for COI genes. The SILVA and COI tutorials also link to downloads for the pre-compiled sequences so that you can save yourself a few days in compiling (and in the case of COI, aligning) the entire databases.

> [@Using RESCRIPt to compile sequence databases and taxonomy classifiers from NCBI Genbank](https://forum.qiime2.org/t/using-rescript-to-compile-sequence-databases-and-taxonomy-classifiers-from-ncbi-genbank/15947):
>
> This tutorial will describe how to create custom reference databases from NCBI Genbank using [RESCRIPt](https://github.com/bokulich-lab/RESCRIPt). Read more about RESCRIPt here: [Processing, filtering, and evaluating the SILVA database (and other reference sequence data) with RESCRIPt - #7](https://forum.qiime2.org/t/processing-filtering-and-evaluating-reference-sequence-data-with-rescript/15494/7) If using NCBI Genbank data, please be aware of the [NCBI disclaimer and copyright notice](https://www.ncbi.nlm.nih.gov/home/about/policies/). Citation: If you use RESCRIPt or any RESCRIPt-processed data in your research, please cite the following pre-print: Michael S Robeson II, Devon R O'Rourke, Benj…

> [@Processing, filtering, and evaluating the SILVA database (and other reference sequence data) with RESCRIPt](https://forum.qiime2.org/t/processing-filtering-and-evaluating-the-silva-database-and-other-reference-sequence-data-with-rescript/15494):
>
> construction Please consider this tutorial a living document, which may change based upon community feedback and ongoing plugin development. RESCRIPt [RESCRIPt](https://github.com/bokulich-lab/RESCRIPt) (REference Sequence annotation and CuRatIon Pipeline) is a python package and QIIME 2 plugin for formatting, managing, and manipulating sequence reference databases. This package was designed for compiling, manipulating, and evaluating sequence reference databases from SILVA, NCBI, Greengenes, GTDB, and other sources, and for construc…

> [@Building a COI database from NCBI references](https://forum.qiime2.org/t/building-a-coi-database-from-ncbi-references/16500):
>
> construction stop_signconstruction stop_signconstruction stop_signconstruction stop_signconstruction stop_signCitation: If you use the following COI resources or RESCRIPt for COI database preparation, please cite the following: Michael S Robeson II, Devon R O’Rourke, Benjamin D Kaehler, Michal Ziemski, Matthew R Dillon, Jeffrey T Foster, Nicholas A Bokulich. RESCRIPt: Reproducible sequence taxonomy reference database management for the masses. bioRxiv 2020.10.05.326504; d…

Good luck!

---

_[View the full topic](https://forum.qiime2.org/t/reference-alignments-best-way-to-download-many-sequences-from-ncbi-genbank/19824)._
