# Unclassified at the phylum level

**URL:** https://forum.qiime2.org/t/unclassified-at-the-phylum-level/31260
**Category:** General Discussion
**Tags:** taxonomy, feature-classifier, greengenes2
**Created:** [August 28, 2024, 11:01am UTC](https://forum.qiime2.org/t/unclassified-at-the-phylum-level/31260 "2024-08-28T11:01:26Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![microbiome\_25](https://forum.qiime2.org/user_avatar/forum.qiime2.org/microbiome_25/32/16450_2.png) [@microbiome\_25](https://forum.qiime2.org/u/microbiome_25)
#### Post date: [August 28, 2024, 11:01am UTC](https://forum.qiime2.org/t/unclassified-at-the-phylum-level/31260/1 "2024-08-28T11:01:26Z")

</div>

I would like to ask about unclassified bacteria at the phylum level.

I am using Greengenes2 (2022.10 full length sequences) for taxonomic classification, using "qiime feature-classifier classify-sklearn" on QIIME2 version 2023.9.  
I got `d __Bacteria;__ ` and `d __Bacteria;p__ ` at the phylum level.  
They seem to be treated as different phyla in the classification, but should I merge them into one phylum category (unclassified bacteria)?  
I have seen discussions in this forum that it is better to discard bacteria that are not classified at the phylum level.  
However, I could not find published papers mentioning this in their method sections.  
Could you please give me some insight on this as well?

Thank you very much!

---

<div class="post-metadata">

### Author: ![salias](https://forum.qiime2.org/user_avatar/forum.qiime2.org/salias/32/18594_2.png) [@salias](https://forum.qiime2.org/u/salias)
#### Post date: [August 28, 2024, 4:28pm UTC](https://forum.qiime2.org/t/unclassified-at-the-phylum-level/31260/5 "2024-08-28T16:28:01Z")

</div>

Hello!

> [@microbiome\_25](#):
>
> They seem to be treated as different phyla in the classification

Functionally, they are the same (lack of) phylum. The difference is that, for `d __Bacteria;__ ` the classifier couldn't assign any taxonomy beyond the domain level; whereas for `d __Bacteria;p__ ` the classifier found a match but that match is not annotated at phylum level. So, long story short: yes, they can be understood as the same on a practical level. [This post](https://forum.qiime2.org/t/what-is-the-differences-between-k-bacteria-and-k-bacteria-p/10572) addresses the same issue.

> [@microbiome\_25](#):
>
> should I merge them into one phylum category (unclassified bacteria)

> [@microbiome\_25](#):
>
> I have seen discussions in this forum that it is better to discard bacteria that are not classified at the phylum level

Yes you could merge them into one "Unclassified" phylum, although I personally prefer to get rid off these too general taxonomic annotations. I came to this conclusion when I asked [this question](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526) a couple of months ago.

> [@microbiome\_25](#):
>
> I could not find published papers mentioning this in their method sections.

Sadly there are a lot of bioinformatic work out there with a Methods section that does not allow replication due to too shallow explanations (not to mention those with no shared code at all...). I suppose this is because a lot of people still think of bioinformatic tools as a black box, and they follow default steps with default variables that they assume they don't need to mention in Methods. Anyway, if you want an example of a Methods section where this filtering is stated, [here you have one](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10269740/). From its Methods section:

> _The table and sequences were filtered to exclude any ASV without phylum-level annotation or which could not be inserted into the phylogenetic tree._

I hope this is useful for you.

Best wishes!

Sergio

---

<div class="post-metadata">

### Author: ![microbiome\_25](https://forum.qiime2.org/user_avatar/forum.qiime2.org/microbiome_25/32/16450_2.png) [@microbiome\_25](https://forum.qiime2.org/u/microbiome_25)
#### Post date: [September 2, 2024, 5:13am UTC](https://forum.qiime2.org/t/unclassified-at-the-phylum-level/31260/6 "2024-09-02T05:13:56Z")

</div>

Hello @salias  
Thank you very much for your answer.  
I have five more questions.

> [@salias](#):
>
> I personally prefer to get rid off these too general taxonomic annotations.

1. Could you clarify what "too general taxonomic annotations" are?
2. Does this mean taxa that are not annotated at the phylum level when analyzing 16S rRNA genes at the phylum and genus level?

> [@salias](#):
>
> I came to this conclusion when I asked [this question](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526) a couple of months ago.

> [@salias](#):
>
> Anyway, if you want an example of a Methods section where this filtering is stated, [here you have one](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10269740/). From its Methods section:
> 
> > _The table and sequences were filtered to exclude any ASV without phylum-level annotation or which could not be inserted into the phylogenetic tree._

Thanks for providing this information.  
3. It appears that the taxa with too general annotations were excluded before the relative abundance of each taxon was calculated, and these taxa were not included in the taxa abundance table. Is this correct?  
4. When you construct a phylogenetic tree and calculate alpha and beta diversity metrics on QIIME2, do you use the filtered table and sequences?  
5. Is it correct that filtering sequences based on annotations is important when comparing taxonomic abundances, not when calculating diversity metrics?

Thank you very much.

---

<div class="post-metadata">

### Author: ![salias](https://forum.qiime2.org/user_avatar/forum.qiime2.org/salias/32/18594_2.png) [@salias](https://forum.qiime2.org/u/salias)
#### Post date: [September 2, 2024, 2:37pm UTC](https://forum.qiime2.org/t/unclassified-at-the-phylum-level/31260/7 "2024-09-02T14:37:38Z")

</div>

Hello again!

> [@microbiome\_25](#):
>
> 1. Could you clarify what "too general taxonomic annotations" are?

> [@microbiome\_25](#):
>
> 1. Does this mean taxa that are not annotated at the phylum level when analyzing 16S rRNA genes at the phylum and genus level?

They are basically poorly classified sequences, those who are annotated only until e.g. phylum level. For example, things like these (Fungi example because I'm a Fungi guy):

`Unassigned; __;__ ; __;__ `

`k __Fungi;p__ Ascomycota; __;__ ;__`

The most likely explanation for these is that they are non-target sequences. You can BLAST a few of them if you want to make sure before discarding them.

> [@microbiome\_25](#):
>
> 1. It appears that the taxa with too general annotations were excluded before the relative abundance of each taxon was calculated, and these taxa were not included in the taxa abundance table. Is this correct?

Yes it is.

> [@microbiome\_25](#):
>
> 1. When you construct a phylogenetic tree and calculate alpha and beta diversity metrics on QIIME2, do you use the filtered table and sequences?

Yes.

> [@microbiome\_25](#):
>
> 1. Is it correct that filtering sequences based on annotations is important when comparing taxonomic abundances, not when calculating diversity metrics?

I would filter for both differential abundance and diversity, otherwise you may get misleading results.

Cheers,

Sergio

---

<div class="post-metadata">

### Author: ![roachjm-unc](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/r/73ab20/32.png) [@roachjm-unc](https://forum.qiime2.org/u/roachjm-unc)
#### Post date: [September 3, 2024, 12:35pm UTC](https://forum.qiime2.org/t/unclassified-at-the-phylum-level/31260/8 "2024-09-03T12:35:50Z")

</div>

It depends a little bit on what community you are actually looking at the sequencing results for, but in my experience, these classifications tends to correspond to eukaryotes / host (i.e. human, mouse, plant, whatever). The 16S universal primers will amplify mitochondrial and chloroplastic DNA, so depending on your situation you may be getting some (or possibly a lot) of reads corresponding to one or more of those.

Check out the BLAST results in the representative sequences qzv for the OTUs / ASVs / sOTUs that get the unclassified / poorly resolved taxa classifications. Then, if they do correspond to human or mouse or plant or whatever, you can do an alignment against that reference (bowtie2 / bwa whatever) to remove those. Then process the remaining reads as you ordinarily would for 16S.

I've found that Kraken2 does a bit better job at identifying and removing the host reads, but depending on your host, building a new Kraken2 database that includes the host may be more work than you are looking for.

---

<div class="post-metadata">

### Author: ![salias](https://forum.qiime2.org/user_avatar/forum.qiime2.org/salias/32/18594_2.png) [@salias](https://forum.qiime2.org/u/salias)
#### Post date: [September 4, 2024, 2:29pm UTC](https://forum.qiime2.org/t/unclassified-at-the-phylum-level/31260/13 "2024-09-04T14:29:37Z")

</div>

Hello @roachjm-unc and @microbiome_25

> [@roachjm-unc](#):
>
> I've found that Kraken2 does a bit better job at identifying and removing the host reads

While I have little experience with Kraken2, I know that other mods like @colinbrislawn also suggest using Kraken2 for this purpose (see line 0083 [here](https://patents.google.com/patent/US20240127907A1/en)).

Also (thanks @SoilRotifer for noting that!), if your number of poorly classified sequences is relatively big, you may want to check if your reads are in mixed orientation. You can [look for more info on mixed orientation in the forum](https://forum.qiime2.org/search?q=mixed%20orientation). While `q2-feature-classifier` needs the reads in the same orientation as the classifier, tools like vsearch or kraken are fine with any read orientation, so they will provide you with a "correct" taxonomy either case. And if you are moving to do something else with your sequences (e.g. phylogenetic tree), you don't want your reads to be in mixed orientation.
