# Taxonomy table looked different in R than when it was opened in QIIME (missing taxon level prefix)?

**URL:** https://forum.qiime2.org/t/taxonomy-table-looked-different-in-r-than-when-it-was-opened-in-qiime-missing-taxon-level-prefix/23205
**Category:** Other Bioinformatics Tools
**Tags:** qiime2r, phyloseq, jupyter, r
**Created:** [May 31, 2022, 5:36am UTC](https://forum.qiime2.org/t/taxonomy-table-looked-different-in-r-than-when-it-was-opened-in-qiime-missing-taxon-level-prefix/23205 "2022-05-31T05:36:36Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![Diki](https://forum.qiime2.org/user_avatar/forum.qiime2.org/diki/32/18419_2.png) [@Diki](https://forum.qiime2.org/u/Diki)
#### Post date: [May 31, 2022, 5:36am UTC](https://forum.qiime2.org/t/taxonomy-table-looked-different-in-r-than-when-it-was-opened-in-qiime-missing-taxon-level-prefix/23205/1 "2022-05-31T05:36:36Z")

</div>

Hi,

I ran QIIME2 for my raw reads via jupyter using `%%bash` in the shared Mac for heavy analysis (also a little Python for data viz). After that, we copied several outputs to build a `phyloseq` object in our personal computer (Win).

Previously, I also did that but found an interesting [topic here](https://github.com/dmil/jupyter-quickstart#jupyter-quickstart) to mix run Python and R in one Jupyter, so I tried it. Because I am not comfortable with jupyter yet, I only stopped at building phyloseq (also for others to use the shared Mac).

```auto
# in bash
# install brew first, check brew.sh
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"

# then install R and libgit2, either from jupyter or shell
brew install r
brew install libgit2 # this is for the R devtools package

```

```auto
# in Jupyter
%load_ext rpy2.ipython

```

```auto
%%R
if (!requireNamespace("devtools", quietly = TRUE)){install.packages("devtools", repos = "https://cloud.r-project.org")} 
devtools::install_github("jbisanz/qiime2R", quiet = TRUE, upgrade = FALSE)

```

```auto
%%R
library(qiime2R)

```

```auto
%%R
physeq <- qza_to_phyloseq(
    features="table.qza",
    tree="rooted-tree.qza",
    taxonomy="taxonomy.qza",
    metadata = "metadata.tsv"
    )

```

```auto
%%R
saveRDS(physeq, file = "ps_mother.rds")

```

It runs without any error notice, but no output. Is this normal?

So it went well and we continued the analysis with RStudio by `ps0 <- readRDS("ps_mother.rds")` to do some filtering. After finishing, we export it back to make both new `taxonomy2.(qza|qzv)`s.

Only to realize that there is a small difference in our taxonomy files. We are missing the taxonomy level prefixes.

| Filenames | Different taxa names |
| --- | --- |
| taxonomy.qzv | d\_\_Bacteria; p\_\_Bacteroidota; c\_\_Bacteroidia; o\_\_Bacteroidales; f\_\_Rikenellaceae; g\_\_Alistipes; s\_\_uncultured\_bacterium |
| taxonomy2.qzv | d\_\_Bacteria;Bacteroidota;Bacteroidia;Bacteroidales;Rikenellaceae;Alistipes;uncultured\_bacterium |

We then tried to investigate at what step does this change occurred by open `ps_mother.rds` first time in R.

```auto
ps0 <- readRDS("ps_mother.rds")
ps0 %>% tax_table() %>% data.frame() %>% filter(Species != "<NA>") %>% head(1)

                                     Kingdom Phylum Class Order
f876bfbfd29d664d606243d9f95edf84 d__Bacteria Proteobacteria Alphaproteobacteria Rickettsiales
                                       Family Genus Species
f876bfbfd29d664d606243d9f95edf84 Mitochondria Mitochondria uncultured_bacterium

```

We found that it is already like this from the first phyloseq object. Do you perhaps know what is wrong here? We worried that other invisible changes might have also occurred.

Our current solution here is to just go back to our regular steps (i.e. build our phyloseq inside R), but we are curious about what happened here.

Any thoughts are appreciated.

Thank you very much.

> **R \`sessionInfo()\`**
>
> R version 4.1.3 (2022-03-10) Platform: x86\_64-w64-mingw32/x64 (64-bit) Running under: Windows 10 x64 (build 19044)
> 
> Matrix products: default
> 
> locale:  
> [1] LC\_COLLATE=English\_United States.1252 LC\_CTYPE=English\_United States.1252  
> [3] LC\_MONETARY=English\_United States.1252 LC\_NUMERIC=C  
> [5] LC\_TIME=English\_United States.1252
> 
> attached base packages:  
> [1] stats graphics grDevices utils datasets methods base
> 
> other attached packages:  
> [1] rlang\_1.0.2 ggrepel\_0.9.1 biomformat\_1.22.0  
> [4] RColorBrewer\_1.1-3 ggstatsplot\_0.9.3 cowplot\_1.1.1  
> [7] decontam\_1.12.0 speedyseq\_0.5.3.9018 phyloseq\_1.38.0  
> [10] qiime2R\_0.99.6 forcats\_0.5.1 stringr\_1.4.0  
> [13] dplyr\_1.0.9 purrr\_0.3.4 readr\_2.1.2  
> [16] tidyr\_1.2.0 tibble\_3.1.7 ggplot2\_3.3.6  
> [19] tidyverse\_1.3.1 here\_1.0.1
> 
> loaded via a namespace (and not attached):  
> [1] readxl\_1.4.0 backports\_1.4.1 Hmisc\_4.7-0  
> [4] plyr\_1.8.7 igraph\_1.3.1 splines\_4.1.3  
> [7] TH.data\_1.1-1 GenomeInfoDb\_1.30.1 digest\_0.6.29  
> [10] foreach\_1.5.2 htmltools\_0.5.2 fansi\_1.0.3  
> [13] magrittr\_2.0.3 checkmate\_2.1.0 paletteer\_1.4.0  
> [16] cluster\_2.1.3 tzdb\_0.3.0 Biostrings\_2.62.0  
> [19] modelr\_0.1.8 vroom\_1.5.7 sandwich\_3.0-1  
> [22] jpeg\_0.1-9 colorspace\_2.0-3 rvest\_1.0.2  
> [25] haven\_2.5.0 xfun\_0.31 crayon\_1.5.1  
> [28] RCurl\_1.98-1.6 jsonlite\_1.8.0 zeallot\_0.1.0  
> [31] survival\_3.3-1 zoo\_1.8-10 iterators\_1.0.14  
> [34] ape\_5.6-2 glue\_1.6.2 gtable\_0.3.0  
> [37] zlibbioc\_1.40.0 emmeans\_1.7.4-1 XVector\_0.34.0  
> [40] statsExpressions\_1.3.2 Rhdf5lib\_1.16.0 BiocGenerics\_0.40.0  
> [43] scales\_1.2.0 mvtnorm\_1.1-3 DBI\_1.1.2  
> [46] Rcpp\_1.0.8.3 performance\_0.9.0 xtable\_1.8-4  
> [49] htmlTable\_2.4.0 bit\_4.0.4 foreign\_0.8-82  
> [52] Formula\_1.2-4 stats4\_4.1.3 DT\_0.23  
> [55] truncnorm\_1.0-8 datawizard\_0.4.1 htmlwidgets\_1.5.4  
> [58] httr\_1.4.3 ellipsis\_0.3.2 farver\_2.1.0  
> [61] pkgconfig\_2.0.3 NADA\_1.6-1.1 nnet\_7.3-17  
> [64] dbplyr\_2.1.1 utf8\_1.2.2 labeling\_0.4.2  
> [67] tidyselect\_1.1.2 reshape2\_1.4.4 munsell\_0.5.0  
> [70] cellranger\_1.1.0 tools\_4.1.3 cli\_3.3.0  
> [73] generics\_0.1.2 ade4\_1.7-19 broom\_0.8.0  
> [76] evaluate\_0.15 fastmap\_1.1.0 yaml\_2.3.5  
> [79] rematch2\_2.1.2 bit64\_4.0.5 knitr\_1.39  
> [82] fs\_1.5.2 nlme\_3.1-157 xml2\_1.3.3  
> [85] correlation\_0.8.1 compiler\_4.1.3 rstudioapi\_0.13  
> [88] png\_0.1-7 zCompositions\_1.4.0-1 reprex\_2.0.1  
> [91] stringi\_1.7.6 parameters\_0.18.1 lattice\_0.20-45  
> [94] Matrix\_1.4-1 vegan\_2.6-2 permute\_0.9-7  
> [97] multtest\_2.50.0 vctrs\_0.4.1 pillar\_1.7.0  
> [100] lifecycle\_1.0.1 rhdf5filters\_1.6.0 estimability\_1.3  
> [103] data.table\_1.14.2 bitops\_1.0-7 insight\_0.17.1  
> [106] patchwork\_1.1.1 R6\_2.5.1 latticeExtra\_0.6-29  
> [109] gridExtra\_2.3 IRanges\_2.28.0 codetools\_0.2-18  
> [112] MASS\_7.3-57 assertthat\_0.2.1 rhdf5\_2.38.1  
> [115] rprojroot\_2.0.3 withr\_2.5.0 multcomp\_1.4-19  
> [118] S4Vectors\_0.32.4 GenomeInfoDbData\_1.2.7 mgcv\_1.8-40  
> [121] bayestestR\_0.12.1 parallel\_4.1.3 hms\_1.1.1  
> [124] grid\_4.1.3 rpart\_4.1.16 coda\_0.19-4  
> [127] rmarkdown\_2.14 Biobase\_2.54.0 lubridate\_1.8.0  
> [130] base64enc\_0.1-3

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [May 31, 2022, 1:34pm UTC](https://forum.qiime2.org/t/taxonomy-table-looked-different-in-r-than-when-it-was-opened-in-qiime-missing-taxon-level-prefix/23205/2 "2022-05-31T13:34:03Z")

</div>

> [@Diki](#):
>
> ```auto
> physeq <- qza_to_phyloseq(
> features="table.qza",
> tree="rooted-tree.qza",
> taxonomy="taxonomy.qza",
> metadata = "metadata.tsv"
> )
> 
> ```

I'm not 100% sure what's going on here. I wonder if something is being lost during import.

For all the Phyloseq import functions, including `qza_to_phyloseq()`, all the input data is joined based on SampleIDs and FeatureIDs that overlap. This means that if some ASVs are missing from the tree, for example, then they would be dropped from the feature table and taxonomy table.

**This 'inner join' during import silently drops all non-matching data!! 🙀**

I wonder if that's what happened here. Try importing these one at a time and check to see if all SampleIDs and FeatureIDs match.

---

<div class="post-metadata">

### Author: ![Diki](https://forum.qiime2.org/user_avatar/forum.qiime2.org/diki/32/18419_2.png) [@Diki](https://forum.qiime2.org/u/Diki)
#### Post date: [June 1, 2022, 7:08am UTC](https://forum.qiime2.org/t/taxonomy-table-looked-different-in-r-than-when-it-was-opened-in-qiime-missing-taxon-level-prefix/23205/5 "2022-06-01T07:08:59Z")

</div>

Hi @colinbrislawn,

Thank you for your response.

I think I know now why the taxon prefix is missing, it is because of `parse_taxonomy()` function inside the `qza_to_phyloseq()`. 😅

However, to my understanding, this [line in `parse_taxonomy()`](https://github.com/jbisanz/qiime2R/blob/2a3cee181513e7451d9b50fdb455cb71fe30fab8/R/parse_taxonomy.R#L24), should also corrects leading `d__` in taxonomy2.qza there. Is it something else?

> [@colinbrislawn](#):
>
> This 'inner join' during import silently drops all non-matching data!! 🙀

I will try it.

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [June 1, 2022, 1:00pm UTC](https://forum.qiime2.org/t/taxonomy-table-looked-different-in-r-than-when-it-was-opened-in-qiime-missing-taxon-level-prefix/23205/6 "2022-06-01T13:00:15Z")

</div>

> [@Diki](#):
>
> However, to my understanding, this [line in `parse_taxonomy()`](https://github.com/jbisanz/qiime2R/blob/2a3cee181513e7451d9b50fdb455cb71fe30fab8/R/parse_taxonomy.R#L24), should also corrects leading `d__` in taxonomy2.qza there.

You are correct. As Mike reminded me:

> there is a set of options within phyloseq to parse the taxonomic information. By default it expects GreenGenese labels: `k __...; p__...; c __...; ...`, this is why all of the labels but the `d__ ` are missing.

This should not change featureIDs, so investigating the inner join is still a good idea!
