# Taxa collapse shows too general taxonomic assignations

**URL:** https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526
**Category:** User Support
**Tags:** taxonomy, its, unite
**Created:** [June 18, 2024, 11:14am UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526 "2024-06-18T11:14:10Z")
**Posts on this page:** 13
**Page:** 1

<div class="post-metadata">

### Author: ![salias](https://forum.qiime2.org/user_avatar/forum.qiime2.org/salias/32/18594_2.png) [@salias](https://forum.qiime2.org/u/salias)
#### Post date: [June 18, 2024, 11:14am UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/1 "2024-06-18T11:14:10Z")

</div>

I'm collapsing my feature table by family taxonomic level prior to running ANCOM-BC. I generated the QZV from the QZA collapsed table, and I found that there is a too general taxonomic assignation with a really high frecuency - in fact, the second highest frequency: `k __Fungi;__ ; __;__ ;__`. You can see it here:

 ![level_5_table_qzv](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/0/0/008e0713906ed1417492dd70014d29069b745a0c.png)

(To clarify: I'm using ITS sequencing data, and I assigned taxonomy with UNITE database).

This is not the only too general taxonomic assignment appearing here: I also spotted things like `Unassigned; __;__ ; __;__ ` or `k __Fungi;p__ Ascomycota; __;__ ;__` (but those have less frequency so I'm less worried about them).

I'm not sure about what would be the best way of work with these general assignments. For now, the options I'm considering are:

1. Perform ANCOM-BC with the collapsed table directly.
2. Rule out all too general taxonomic assignation I find.
3. Rule out too general taxonomic assignations only if they have low frequency. In my case, that would mean that I keep `k __Fungi;__ ; __;__ ;__`, but I exclude the rest.

Any suggestions? I believe there is a consensus of how to deal with too general taxonomy but I don't know why I cannot find that exact issue in the forum.

Thank you in advance

Best wishes,

Sergio

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [June 18, 2024, 12:53pm UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/2 "2024-06-18T12:53:02Z")

</div>

Good morning, Sergio,

I think you have outlined three good options. I would like to suggest a 4th.

> [@salias](#):
>
> I'm collapsing my feature table by family taxonomic level prior to running ANCOM-BC.

1. run ANBOM-BC at the ASV level (and add taxonomy information later)

This may be more difficult bioinformatically, but means the ANCOM results will not be biased by taxonomy. Because taxonomy can be hard for ITS, this is super helpful.

---

<div class="post-metadata">

### Author: ![salias](https://forum.qiime2.org/user_avatar/forum.qiime2.org/salias/32/18594_2.png) [@salias](https://forum.qiime2.org/u/salias)
#### Post date: [June 18, 2024, 2:03pm UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/3 "2024-06-18T14:03:31Z")

</div>

Hello @colinbrislawn ,

Thank you so much. Yes, in fact my workflow runs ANCOM-BC at ASV, species, genus and family levels. For ASV level I'll load the taxonomy in R and map hashes to taxonomies prior to plotting.

For now I'll try that without ruling out general taxonomies and I'll see what happens. Maybe those are really abundant in all samples but not differentially abundant so they are not going to bother me - if they do, it is quite likely that I come back here if I find any issue when running ANCOM-BC on filtered tables.

Again, thank you so much

Best wishes,

Sergio

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [June 18, 2024, 3:40pm UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/4 "2024-06-18T15:40:11Z")

</div>

Hi @salias ,  
I would go with your option 2 (removing poorly classified sequences) after checking a few with NCBI BLAST to see what else they may hit.

Most ITS primers can amplify non-fungal eukaryotes, and if these are not represented in your database you will often get these poor classifications. So it is often a good idea to use the UNITE fungi +\_eukaryote database to detect these non-target hits.

Presumably you don't want non-fungi in your survey, so I would just remove them (after confirming that they are non-fungal or junk etc)

> [@colinbrislawn](#):
>
> run ANBOM-BC at the ASV level (and add taxonomy information later)

This might not be a good idea either, as you might then still include these in alpha and beta diversity and other measurements where having non-target sequences present could lead to misleading results. But that depends on your biological question and methods...

---

<div class="post-metadata">

### Author: ![salias](https://forum.qiime2.org/user_avatar/forum.qiime2.org/salias/32/18594_2.png) [@salias](https://forum.qiime2.org/u/salias)
#### Post date: [June 18, 2024, 4:59pm UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/5 "2024-06-18T16:59:44Z")

</div>

Hello @Nicholas_Bokulich ,

Thank you so much for your comment.

> [@Nicholas\_Bokulich](#):
>
> So it is often a good idea to use the UNITE fungi +\_eukaryote database to detect these non-target hits.

Yes, I use the UNITE version with all eukaryotes and remove all the non-fungi I come across (`k __Alveolata` , `k__ Metazoa`, `k__Viridiplantae`, etc). I was specially worried about the Unassigned and the too general Fungi annotations, so I kept those.

> [@Nicholas\_Bokulich](#):
>
> Presumably you don't want non-fungi in your survey, so I would just remove them (after confirming that they are non-fungal or junk etc)

Okay, I'll check with BLAST and if I find too general fungal matches I'll remove them. If I find something specific maybe it's time to RESCRIPt.

> [@Nicholas\_Bokulich](#):
>
> you might then still include these in alpha and beta diversity and other measurements where having non-target sequences present could lead to misleading results.

Yes, currently my diversity analyses include those too general fungal results. I do diversity analysis directly on ASVs (because as far as I understood ASVs, collapsing before diversity would mean losing the advantages of using ASVs). So I suppose what I should do is use the `taxa barplot` visualization to make sure the uncollapsed table is okay before running `diversity`.

Apart from those too general kingdom-level assignments, I also found entries that only have information until phylum level (e.g. `k __Fungi;p__ Ascomycota;__`), class level, order level and so on. Should I keep those e.g. when I want to do ANCOM-BC on a family-collapsed table? Or should I be permissive until one category above (e.g. allow no more general than order-level assignments when using family-collapsed table)?

Again, thank you so much

Best,

Sergio

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [June 19, 2024, 5:12am UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/6 "2024-06-19T05:12:50Z")

</div>

> [@salias](#):
>
> Apart from those too general kingdom-level assignments, I also found entries that only have information until phylum level (e.g. `k __Fungi;p__ Ascomycota;__`), class level, order level and so on. Should I keep those e.g. when I want to do ANCOM-BC on a family-collapsed table? Or should I be permissive until one category above (e.g. allow no more general than order-level assignments when using family-collapsed table)?

In the past I have found that anything that does not classify to at least class or order level is usually also junk sequences, e.g., non-target etc. So I would check the ASVs that only classify to phylum level to confirm, too, but they are probably something that should be removed as well...

good luck!

---

<div class="post-metadata">

### Author: ![salias](https://forum.qiime2.org/user_avatar/forum.qiime2.org/salias/32/18594_2.png) [@salias](https://forum.qiime2.org/u/salias)
#### Post date: [June 19, 2024, 9:35am UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/7 "2024-06-19T09:35:44Z")

</div>

Hello @Nicholas_Bokulich ,

Okay, I will follow your advice and check the general assignments to confirm they are junk and filter them. Thank you!

Sergio

---

<div class="post-metadata">

### Author: ![salias](https://forum.qiime2.org/user_avatar/forum.qiime2.org/salias/32/18594_2.png) [@salias](https://forum.qiime2.org/u/salias)
#### Post date: [June 20, 2024, 8:47am UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/8 "2024-06-20T08:47:28Z")

</div>

Hi,

Sorry for bringing this back again, but I've been thinking about this:

> [@Nicholas\_Bokulich](#):
>
> having non-target sequences present could lead to misleading results. But that depends on your biological question and methods...

I think that, **maybe** , those sequencies with only a fungal phylum associated (or simply `k __Fungi;__ ; __;__ ;__`) that does not match with anything meaningful in NCBI BLAST could be actually uncultured / uncharacterised fungi not present either in UNITE or BLAST. So if the biological cuestion is related with "search for as yet undescribed fungi", we should keep them. And when performing e.g. ANCOM-BC it should be done on the ASV level (because if we taxa collapse, maybe two uncultured fungi annotated as `k__Fungi;p __Ascomycota;__ ; __;__ ` are considered as the same one).

Maybe I'm overthinking it 🤯 but in my head this approach sounds nice.

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [June 20, 2024, 9:00am UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/9 "2024-06-20T09:00:54Z")

</div>

> [@salias](#):
>
> I think that, **maybe** , those sequencies with only a fungal phylum associated (or simply `k __Fungi;__ ; __;__ ;__`) that does not match with anything meaningful in NCBI BLAST could be actually uncultured / uncharacterised fungi not present either in UNITE or BLAST.

well, it depends on your samples and question. How likely is it to find a novel fungal phylum in your samples? Discovering a totally novel phylum is certainly possible, but unlikely, with the degree of unlikelihood depending on where you look.

And even if it were a new phylum, you would still expect this to hit other fungi in the nr database with NCBI BLAST, just with a poor alignment. So inspect the alignments, and if it hits _nothing_ it's questionable (more likely junk than a new phylum). Make sure to search the full nr database with these, do not restrict to a specific group.

> [@salias](#):
>
> And when performing e.g. ANCOM-BC it should be done on the ASV level (because if we taxa collapse, maybe two uncultured fungi annotated as `k __Fungi;p__ Ascomycota; __;__ ;__` are considered as the same one).

Yes indeed, I would always do differential abundance testing on ASV level (even if taxonomic levels are tested separately), for the reason that some ASVs could be differentially abundant and that is interesting. But I would only do this on valid ASVs, as differential abundance of, e.g., non-target DNA is probably not useful (but this also depends on your experimental goals and question and system and etc)

---

<div class="post-metadata">

### Author: ![salias](https://forum.qiime2.org/user_avatar/forum.qiime2.org/salias/32/18594_2.png) [@salias](https://forum.qiime2.org/u/salias)
#### Post date: [June 20, 2024, 9:56am UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/10 "2024-06-20T09:56:03Z")

</div>

Hi @Nicholas_Bokulich ,

Okay, I understand now. I was thinking that somehow e.g. `k __Fungi;p__ Ascomycota; __;__ ;__` could be assigned by the classifier because the sequence is equally likely to belong to two species of different classes (e.g. class `c__A` and class `c__B`), but in real life what is happening is that the species belong to class A but it is not sufficiently well described. But you are right, in those cases the most likely scenario is that the sequence is junk (or a novel class, but that would be very very very rare).

Thank you so much and sorry for cluttering the forum with my questions 😅

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [June 20, 2024, 1:59pm UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/11 "2024-06-20T13:59:20Z")

</div>

> [@salias](#):
>
> I was thinking that somehow e.g. `k __Fungi;p__ Ascomycota; __;__ ;__` could be assigned by the classifier because the sequence is equally likely to belong to two species of different classes (e.g. class `c__A` and class `c__B`), but in real life what is happening is that the species belong to class A but it is not sufficiently well described.

Your first inference is one possible scenario; the other possibility is that the sequence simply does not resemble _any_ class with a sufficiently high degree of probability.

> [@salias](#):
>
> sorry for cluttering the forum with my questions

On the contrary, thanks for the lively discussion! These are great questions, and you are not alone in searching for answers...

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [June 20, 2024, 2:20pm UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/12 "2024-06-20T14:20:18Z")

</div>

Sometimes I forget that Nick literally wrote the book on [Optimizing amplicon taxonomic classification](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5956843/) and [ITS primer design](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3623200/).

I usually start with a naive approach because I am also naive.

The advice from resident experts like Nick helps me update my priors. 🧑‍🎓

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [July 21, 2024, 9:15pm UTC](https://forum.qiime2.org/t/taxa-collapse-shows-too-general-taxonomic-assignations/30526/14 "2024-07-21T21:15:48Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
