# SILVA taxonomy-classifier clarification

**URL:** https://forum.qiime2.org/t/silva-taxonomy-classifier-clarification/5087
**Category:** User Support
**Created:** [July 17, 2018, 1:46pm UTC](https://forum.qiime2.org/t/silva-taxonomy-classifier-clarification/5087 "2018-07-17T13:46:26Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![Kara](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/k/e480ec/32.png) [@Kara](https://forum.qiime2.org/u/Kara)
#### Post date: [July 17, 2018, 1:46pm UTC](https://forum.qiime2.org/t/silva-taxonomy-classifier-clarification/5087/1 "2018-07-17T13:46:27Z")

</div>

We had to dig around a little to find an example of how to import and use the rep\_set and taxonomy files from the SILVA\_132\_release classifier in QIIME2. We downloaded the 132 release files and then imported the 99% rep\_set and 99% taxonomy files (majority\_ or consensus\_taxonomy\_7\_levels) into QIIME2. We were then able to run the command below in QIIME2. There are so many variables to change in this final command, and we wanted to double-check that we understand all of the options (there are a lot of "behind-the-scenes" steps going on!):

The first variable seems to be the rep\_set. We get to choose the rep\_set % as either 90, 94, 97, or 99 for 16S only data. It sounds like the standard is 97% or above, and from what we understand, this determines the percent sequence similarity for clustering. For example, it starts with one long sequence as a "seed" in the cluster and compares a second sequence to it. If the second sequence is 97%+ similar to the first, they will be clustered together (they'll be assigned a taxonomy later). If it's less than 97% similar it will become it's own "seed" in a cluster. This continues with the 3rd, 4th, etc. sequences until all sequences have been clustered based on this rep\_set percentage threshold. Is that correct?

Now that the clusters are made, taxonomy can be assigned. We get to choose the taxonomy % as either 90, 94, 97, or 99 as well. Within those options we can then choose our desired consensus\_taxonomy\_7\_levels or majority\_taxonomy\_7\_levels (for us, the choice between consensus and majority didn't impact how many reads came up as unassigned at the end). We understand consensus to mean that all potential taxa strings for a cluster are identical (100% the same), whereas majority means they are 90% similar. We aren't quite sure then, how you could pick for example the 94 taxonomy folder and then choose the consensus\_taxonomy option inside. What do these 90, 94, 97, 99 %s mean if it's not related to the taxa string similarities within a cluster? Do we need to choose the same % for the taxonomy file as the rep\_set file? Can they be different?

The last variable seems to be the --p-perc-identity 0.98. This seems the easiest to understand. It's the minimum percent similarity we want between OUR unknown sequences and the BLAST sequences. Yes?

Phew---long question......sorry! Hoping for any clarification of our questions above!

 ![19%20AM](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/8/87a358682b7b52c94c3c028595ca4646fe4ef741.png)

---

<div class="post-metadata">

### Author: ![thermokarst](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/t/8e7dd6/32.png) [@thermokarst](https://forum.qiime2.org/u/thermokarst)
#### Post date: [July 17, 2018, 2:05pm UTC](https://forum.qiime2.org/t/silva-taxonomy-classifier-clarification/5087/2 "2018-07-17T14:05:08Z")

</div>



---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [July 17, 2018, 3:54pm UTC](https://forum.qiime2.org/t/silva-taxonomy-classifier-clarification/5087/3 "2018-07-17T15:54:05Z")

</div>

> [@Kara](#):
>
> We had to dig around a little to find an example of how to import and use the rep\_set and taxonomy files from the SILVA\_132\_release classifier in QIIME2

You could have looked in the [overview tutorial](https://docs.qiime2.org/2018.6/tutorials/overview/#taxonomy-classification-and-taxonomic-analyses) or [classifier training](https://docs.qiime2.org/2018.6/tutorials/feature-classifier/) tutorial, which both give examples (the latter shows a workflow for importing and getting to blast, the latter has actual files or greengenes in the expected formats and import examples).

> [@Kara](#):
>
> There are so many variables to change in this final command, and we wanted to double-check that we understand all of the options

Sounds like you are conflating the SILVA database construction with actual taxonomy assignment.

The `classify-consensus-blast` command uses blast+ for database searching, followed by LCA taxonomy consensus assignment in QIIME2. You can read more about how that method works and its parameters in the [publication for q2-feature-classifier](https://microbiomejournal.biomedcentral.com/articles/10.1186/s40168-018-0470-z)

The alignment-related parameters come directly from blastn — you can read the [blastn manual](https://www.ncbi.nlm.nih.gov/books/NBK279690/) for additional details.

> [@Kara](#):
>
> The first variable seems to be the rep\_set

Sounds like you are talking about which SILVA `rep_set` to use, not QIIME 2 parameters.

Sounds like you've read the readme that comes with the SILVA qiime-compatible release. I'd recommend the 99% OTUs rep set, as this has the greatest specificity. I prefer the majority taxonomy, since some labels can be incorrect and that will cause problems if you are looking for 100% consensus.

> [@Kara](#):
>
> Now that the clusters are made, taxonomy can be assigned.

Again, you are conflating database building with taxonomy assignment. You are just making a selection of database OTU clusters (and their matching taxonomies), there is not really a complicated decision to make at this stage involving multiple parameters (see below). There is one decision to make: what level of OTU clustering do I wish to use? Higher will mean more specific OTUs, more specific taxonomies, but also many more reference sequences (leading to longer runtime and memory requirements). So the decision is easy: always use 99% unless if you are unable to do so due to computational limitations.

> [@Kara](#):
>
> Do we need to choose the same % for the taxonomy file as the rep\_set file? Can they be different?

Yes, you need to choose the matching files. The taxonomy files are just the taxonomy labels that correspond to the different rep sets — they have not been clustered independently. The consensus vs. majority taxonomies are based on the raw taxonomies of sequences that are clustered together. So the different taxonomy files contain different IDs and potentially different taxonomy labels for any IDs that are shared between these files (because OTU clusters will be tighter at 99% than 94%, for example, and lead to shallower consensus/majority taxonomies). But most importantly you need the IDs to match or the taxonomy classification will not work.

I hope that clarifies!

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [July 17, 2018, 3:54pm UTC](https://forum.qiime2.org/t/silva-taxonomy-classifier-clarification/5087/4 "2018-07-17T15:54:08Z")

</div>



---

<div class="post-metadata">

### Author: ![Kara](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/k/e480ec/32.png) [@Kara](https://forum.qiime2.org/u/Kara)
#### Post date: [July 18, 2018, 7:06pm UTC](https://forum.qiime2.org/t/silva-taxonomy-classifier-clarification/5087/5 "2018-07-18T19:06:31Z")

</div>

Yes, that absolutely clarifies! We had no idea that we had to choose the same "99" number for the taxonomy folder as the "99" number for the rep\_set. We couldn't understand what the "99" would even mean for taxonomy! This helps a LOT! Thanks very much for deciphering what we were actually trying to ask, and then answer it 🙂. We are VERY new to this---so it's hard to even put our questions into the correct phrasing sometimes. Working on it.....

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [August 19, 2018, 1:06am UTC](https://forum.qiime2.org/t/silva-taxonomy-classifier-clarification/5087/6 "2018-08-19T01:06:33Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
