# q2-phylogeny for vertebrates

**URL:** https://forum.qiime2.org/t/q2-phylogeny-for-vertebrates/23955
**Category:** User Support
**Tags:** phylogeny, fragment-insertion
**Created:** [August 22, 2022, 10:13pm UTC](https://forum.qiime2.org/t/q2-phylogeny-for-vertebrates/23955 "2022-08-22T22:13:51Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![DeannaB](https://forum.qiime2.org/user_avatar/forum.qiime2.org/deannab/32/9387_2.png) [@DeannaB](https://forum.qiime2.org/u/DeannaB)
#### Post date: [August 22, 2022, 10:13pm UTC](https://forum.qiime2.org/t/q2-phylogeny-for-vertebrates/23955/1 "2022-08-22T22:13:51Z")

</div>

Hello, I'm working with Illumina short reads of fish environmental DNA (using 12S MiFish primers). I've successfully generated a de novo phylogenetic tree of OTU clusters using **qiime phylogeny fasttree** after generating a Mafft alignment and masking noisy positions. However, I'd like to include some reference sequences either using the [mitohelper qza formatted reference database](https://zenodo.org/record/6336244#.YwP81uzML9E) or by adding some reference sequences to my rep\_seqs.qza file. Is there a way to achieve this? Or is this best achieved outside of qiime2?

From my understanding, I won't be able to use q2-fragment-insertion because there isn't a SeppReferenceDatabase available or validated for the reference database I'm using, according to the qiime2 forum post [here](https://forum.qiime2.org/t/fragment-insertion-sepp-error/15965).

Thank you for any guidance.  
Deanna

---

<div class="post-metadata">

### Author: ![SoilRotifer](https://forum.qiime2.org/user_avatar/forum.qiime2.org/soilrotifer/32/21071_2.png) [@SoilRotifer](https://forum.qiime2.org/u/SoilRotifer)
#### Post date: [August 23, 2022, 5:26pm UTC](https://forum.qiime2.org/t/q2-phylogeny-for-vertebrates/23955/2 "2022-08-23T17:26:55Z")

</div>

Hi @DeannaB,

I think the easiest approach would be the following:

1. First extract the amplicon region from the `12S-seqs-derep-uniq.qza`, and save for later. Like this:

```auto
qiime feature-classifier extract-reads \
  --i-sequences 12S-seqs-derep-uniq.qza \
  --p-f-primer forward-primer-sequence \
  --p-r-primer reverse-primer-sequence \
  --p-trunc-len xx \
  --p-min-length yy \
  --p-max-length zz \
  --o-reads 12S-seqs-derep-extract-uniq.qza

```

1. Generate a taxonomy visualization for the `12S-tax-derep-uniq.qza`:

```auto
qiime metadata tabulate \
  --m-input-file 12S-tax-derep-uniq.qza \
  --o-visualization 12S-tax-derep-uniq.qzv

```

1. Then using [QIIME 2 View](https://view.qiime2.org/) search (upper right of visualization) or scroll through the taxonomy presented in the `12S-tax-derep-uniq.qzv` file. Write down a list of IDs that you'd like to use into a text file. We'll call it `seq-ids-to-keep.txt`. Using the following format:

> feature-id  
> AY846380.1.2583  
> AY909584.1.2313  
> AY929372.1.1770  
> ...  
> ...

1. Then you can run the following command to extract those reference sequences and write them to file:

```auto
qiime feature-table filter-seqs \
     --i-data 12S-tax-derep-uniq.qza \
     --m-metadata-file seq-ids-to-keep.txt \
     --o-filtered-data 12S-tax-derep-uniq-subset.qza 

```

1. Now you can simply merge your reference sequences (the amplicon region of the mitohelper database we extracted earlier) with your and OTUs/ESVs (here called `my-otus.qza`) with the merge command:

```auto
qiime feature-table merge-seqs \
	--i-data 12S-tax-derep-uniq-subset.qza my-otus.qza \
	--o-merged-data merged-12S-refs-and-my-otus.qza

```

1. Then you should be able to build your phylogeny how you'd like.

-Cheers!  
-Mike

---

<div class="post-metadata">

### Author: ![DeannaB](https://forum.qiime2.org/user_avatar/forum.qiime2.org/deannab/32/9387_2.png) [@DeannaB](https://forum.qiime2.org/u/DeannaB)
#### Post date: [August 24, 2022, 4:45am UTC](https://forum.qiime2.org/t/q2-phylogeny-for-vertebrates/23955/3 "2022-08-24T04:45:30Z")

</div>

This is a great suggestion. Thank you for this.

I've added one additional step, which is to merge the `FeatureData[Taxonomy]` associated with my OTUs with the `FeatureData[Taxonomy]` associated with the reference database:

`qiime feature-table merge-taxa --i-data rep_seqs_mitofish_blast_taxonomy.qza 12S-tax-derep-uniq.qza --o-merged-data rep_seqs_mitofish_blast_taxonomy_12S-tax-derep-uniq.qza`

I couldn't find a way to filter the `FeatureData[Taxonomy]` artifact of the reference database (12S-tax-derep-uniq.qza) prior to merging. I did find this [post](https://forum.qiime2.org/t/filtering-taxonomy-tables/2030/10) on the topic.

However, in the end it doesn't seem to matter because I am importing the tree, feature-table, and taxonomy into a phyloseq object in R to draw a tree with the plot\_tree function. And according to phyloseq [documentation](https://joey711.github.io/phyloseq/import-data.html): "OTUs and samples are included in the combined object only if they are present in all components. For instance, extra “leaves” on the tree will be trimmed off when that tree is added to a phyloseq object."

If you have any suggestions on how to circumvent this problem, I'd love to know. The end goal is to draw a phylogenetic tree of my OTUs that contains some reference sequences (with taxonomic labels or accession numbers for reference sequences).

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [August 24, 2022, 5:35am UTC](https://forum.qiime2.org/t/q2-phylogeny-for-vertebrates/23955/4 "2022-08-24T05:35:42Z")

</div>

Hi @DeannaB ,  
I just want to chime in to add to @SoilRotifer 's excellent advice.

One amendment to this step:

> [@SoilRotifer](#):
>
> Then using [QIIME 2 View](https://view.qiime2.org/) search (upper right of visualization) or scroll through the taxonomy presented in the `12S-tax-derep-uniq.qzv` file. Write down a list of IDs that you'd like to use into a text file. We'll call it `seq-ids-to-keep.txt`. Using the following format:

Making a text file of IDs should not be necessary — there should be programmatic ways of accomplishing the same thing, whether [you want IDs that are in a feature table](https://forum.qiime2.org/t/filtering-sequences-by-feature-table/6767/2), or to [filter based on IDs found in a taxonomy](https://docs.qiime2.org/2022.2/tutorials/filtering/#filtering-sequences).

> [@DeannaB](#):
>
> I couldn't find a way to filter the `FeatureData[Taxonomy]` artifact of the reference database (12S-tax-derep-uniq.qza) prior to merging. I did find this [post](https://forum.qiime2.org/t/filtering-taxonomy-tables/2030/10) on the topic.

The `RESCRIPt` plugin (not currently installed as part of QIIME 2, but installation instructions are available on the forum) has an action to filter a taxonomy based on a list of IDs or search term. See this tutorial:

> [@Using RESCRIPt to compile sequence databases and taxonomy classifiers from NCBI Genbank](https://forum.qiime2.org/t/using-rescript-to-compile-sequence-databases-and-taxonomy-classifiers-from-ncbi-genbank/15947#creating-a-classifier-from-refseqs-data-4):
>
> This tutorial will describe how to create custom reference databases from NCBI Genbank using [RESCRIPt](https://github.com/bokulich-lab/RESCRIPt). Read more about RESCRIPt here: [Processing, filtering, and evaluating the SILVA database (and other reference sequence data) with RESCRIPt - #7](https://forum.qiime2.org/t/processing-filtering-and-evaluating-reference-sequence-data-with-rescript/15494/7) If using NCBI Genbank data, please be aware of the [NCBI disclaimer and copyright notice](https://www.ncbi.nlm.nih.gov/home/about/policies/). Citation: If you use RESCRIPt or any RESCRIPt-processed data in your research, please cite the following pre-print: Michael S Robeson II, Devon R O'Rourke, Benj…

RESCRIPt, by the way, could also be used to programmatically download reference sequences and taxonomies directly from NCBI based on an entrez search query... so that could also be an option if you only want to grab a limited number of accessions vs. all of mitofish.

Good luck!

---

<div class="post-metadata">

### Author: ![DeannaB](https://forum.qiime2.org/user_avatar/forum.qiime2.org/deannab/32/9387_2.png) [@DeannaB](https://forum.qiime2.org/u/DeannaB)
#### Post date: [August 25, 2022, 5:35pm UTC](https://forum.qiime2.org/t/q2-phylogeny-for-vertebrates/23955/7 "2022-08-25T17:35:18Z")

</div>

@Nicholas_Bokulich thanks so much for these suggestions.

And I agree on your point above "there should be programmatic ways of accomplishing the same thing." For this particular reference database, I've used a python tool called [mitohelper](https://github.com/aomlomics/mitohelper) for this purpose (functions get record and get alignment).

I didn't know about these functionalities of the RESCRIPt plugin! Thank you for this information. This looks like another viable option to obtain references & associated taxonomies.

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [September 25, 2022, 11:35pm UTC](https://forum.qiime2.org/t/q2-phylogeny-for-vertebrates/23955/8 "2022-09-25T23:35:33Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
