# how to convert a trained taxonomy classifier to a gz file?

**URL:** https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021
**Category:** Other Bioinformatics Tools
**Created:** [January 11, 2021, 12:02pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021 "2021-01-11T12:02:28Z")
**Posts on this page:** 15
**Page:** 1

<div class="post-metadata">

### Author: ![farhad1990](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/f/f05b48/32.png) [@farhad1990](https://forum.qiime2.org/u/farhad1990)
#### Post date: [January 11, 2021, 12:02pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/1 "2021-01-11T12:02:28Z")

</div>

Hi everyone,

I have decided to use R for my own continence as my pipeline recently. However, I really couldn't find a method to train a region-specific classifier (based on my primers) according to SILVA 138 database in R. However, back then I have already trained my own classifier based on my primer-set in QIIME2 through this [tutorial](https://forum.qiime2.org/t/processing-filtering-and-evaluating-the-silva-database-and-other-reference-sequence-data-with-rescript/15494), which its format is .qza, while the format which assignTaxonomy() function in R knows is a .gz as the input for the classifier library. Is there any way to convert my silva138-classifier-341f-805r.qza to silva138-classifier-341f-805r.gz?  
Much appreciated in advance.

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [January 11, 2021, 12:28pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/2 "2021-01-11T12:28:49Z")

</div>

> [@farhad1990](#):
>
> However, I really couldn’t find a method to train a region-specific classifier (based on my primers) according to SILVA 138 database in R.

Perhaps because this method and workflow are quite specific to QIIME 2. There are other taxonomy classifiers available in R, but they have their own workflows and input formats.

> [@farhad1990](#):
>
> Is there any way to convert my silva138-classifier-341f-805r.qza to silva138-classifier-341f-805r.gz?

No. That file does not contain a gzipped set of DNA sequences, it contains a trained scikit-learn classifier, which would be unreadable by anything in R. So there is no way to export it and use that classifier in R.

note: I edited the title to make it more specific to your question. QZA is a vague extension (just as gz can contain any gzipped contents, a QZA can contain any QIIME 2 results).

Good luck!

---

<div class="post-metadata">

### Author: ![farhad1990](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/f/f05b48/32.png) [@farhad1990](https://forum.qiime2.org/u/farhad1990)
#### Post date: [January 11, 2021, 2:30pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/3 "2021-01-11T14:30:41Z")

</div>

Thanks Nicholas,

I already downloaded the silva138 classifier but it is for the whole gene, while I want to have it region-specific to my primer sets. Since couldn't find a workflow to make the classifier region-specific to my primer, do you think it would be alright if I just go with the whole classifier?

Kinds

---

<div class="post-metadata">

### Author: ![SoilRotifer](https://forum.qiime2.org/user_avatar/forum.qiime2.org/soilrotifer/32/21071_2.png) [@SoilRotifer](https://forum.qiime2.org/u/SoilRotifer)
#### Post date: [January 11, 2021, 3:15pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/4 "2021-01-11T15:15:05Z")

</div>

Hi @farhad1990,

> [@farhad1990](#):
>
> Since couldn’t find a workflow to make the classifier region-specific to my primer, do you think it would be alright if I just go with the whole classifier?

Earlier in this thread you said you worked through the RESCRIPt tutorial. Did you not try [this part](https://forum.qiime2.org/t/processing-filtering-and-evaluating-the-silva-database-and-other-reference-sequence-data-with-rescript/15494#heading--sixth-header) of the tutorial?

-Mike

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [January 11, 2021, 3:26pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/5 "2021-01-11T15:26:33Z")

</div>

Ah thanks for clarifying @SoilRotifer — I think I understand your question now @farhad1990

Are you asking rather " **can RESCRIPt be used to create a classifier that can be exported and used in R?**"

**You could use RESCRIPt (following that tutorial) to compile and format a custom reference sequence database, then export those formatted _sequences_** (prior to training the classifier) to fasta format and use them in R (e.g., for taxonomy classification). However, you cannot export a trained classifier and use it in R because it is in a very special format.

---

<div class="post-metadata">

### Author: ![farhad1990](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/f/f05b48/32.png) [@farhad1990](https://forum.qiime2.org/u/farhad1990)
#### Post date: [January 11, 2021, 4:15pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/6 "2021-01-11T16:15:25Z")

</div>

Hi Mike,  
Yes I did [this part](https://forum.qiime2.org/t/processing-filtering-and-evaluating-the-silva-database-and-other-reference-sequence-data-with-rescript/15494#heading--sixth-header) for sure and trained my classifier. However, now I would like to use this classifier which I trained and specified it to my primers, to be used in my R workflow. Since I couldn't find a way in R for training and making an amplicon-specific classifier, I was wondering if there is any ways to convert use my in-qiime-trained classifier in R.

Kinds,  
Farhad

---

<div class="post-metadata">

### Author: ![SoilRotifer](https://forum.qiime2.org/user_avatar/forum.qiime2.org/soilrotifer/32/21071_2.png) [@SoilRotifer](https://forum.qiime2.org/u/SoilRotifer)
#### Post date: [January 11, 2021, 5:06pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/7 "2021-01-11T17:06:27Z")

</div>

Okay, then you can proceed as @Nicholas_Bokulich suggested above. Just take the formatted taxonomy and sequence files (the ones you'd input into the classifier) and import them into R instead. Then use your favorite R tools to train your reference database and classify. For example, you can likely use the approach from [this pipeline](https://benjjneb.github.io/dada2/assign.html).

---

<div class="post-metadata">

### Author: ![farhad1990](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/f/f05b48/32.png) [@farhad1990](https://forum.qiime2.org/u/farhad1990)
#### Post date: [January 11, 2021, 5:10pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/8 "2021-01-11T17:10:00Z")

</div>

Thanks Mike,  
I will go for it 🙂

Kinds,  
Farhad

---

<div class="post-metadata">

### Author: ![farhad1990](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/f/f05b48/32.png) [@farhad1990](https://forum.qiime2.org/u/farhad1990)
#### Post date: [January 11, 2021, 7:35pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/9 "2021-01-11T19:35:06Z")

</div>

That was the exact question I asked and thanks for the answer. I am currently working on it.

---

<div class="post-metadata">

### Author: ![farhad1990](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/f/f05b48/32.png) [@farhad1990](https://forum.qiime2.org/u/farhad1990)
#### Post date: [January 18, 2021, 9:17pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/10 "2021-01-18T21:17:05Z")

</div>

Hi Mike,

I have sued the formatted taxonomy and sequence files as the input for the dada2:::makeTaxonomyFasta\_SilvaNR() function in R and got the region-specific classifier based on my primers. However, when I used it for assigning the taxonomy I can see that compared to the non-region-specific classifier ([silva\_nr99\_v138\_wSpecies\_train\_set.fa.gz](https://www.arb-silva.de/fileadmin/silva_databases/release_138/Exports/SILVA_138_SSURef_tax_silva.fasta.gz)), I got a lot of NAs in different taxa levels.  
This pic is for the results of region-specific classifier

 ![image](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/c/ceadea90fd26a26711b517cebe4f20915ef4c8b1.png)

and this is for the full seq classifier

 ![image](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/f/f1d69bbc09aa257918fbc7b7435e62630ba8c5d5.png)  
Is it expected or something is wrong?

Kinds,  
Farhad

---

<div class="post-metadata">

### Author: ![SoilRotifer](https://forum.qiime2.org/user_avatar/forum.qiime2.org/soilrotifer/32/21071_2.png) [@SoilRotifer](https://forum.qiime2.org/u/SoilRotifer)
#### Post date: [January 18, 2021, 11:17pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/11 "2021-01-18T23:17:36Z")

</div>

Difficult to say.

Is there a reason why you link to the fasta file from the SILVA web site? The name you provide does not match the file name in the link. Was this intended?

I guess I'd need more details on how the reference data is handled and/or further processed within R prior to classification. I am unfamiliar with how this tool and its commands (e.g. `makeTaxonomyFasta_SilvaNR`) works. So, I'd suggest classifying through QIIME 2 for a comparison and sanity-check. 🤷‍♂️

I assume you followed all of the "Make amplicon-region specific classifier" parts of the RESCRIPt tutorial? That is, the sequence and taxonomy dereplication steps, prior to importing them into R? Just asking to make sure I understand all the steps you've taken. 🙂

Check out this thread: [training classifiers: performance of full-length vs. extract-reads](https://forum.qiime2.org/t/training-classifiers-performance-of-full-length-vs-extract-reads/14138) for some additional insights.

Do you have anything to add, @Nicholas_Bokulich ?

-Mike

---

<div class="post-metadata">

### Author: ![TurboQiimer](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/t/46a35a/32.png) [@TurboQiimer](https://forum.qiime2.org/u/TurboQiimer)
#### Post date: [January 19, 2021, 1:12am UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/12 "2021-01-19T01:12:35Z")

</div>

Hi,  
Regarding the photo you shared, can I ask what environment it is?  
Thanks  
Qiimer

---

<div class="post-metadata">

### Author: ![farhad1990](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/f/f05b48/32.png) [@farhad1990](https://forum.qiime2.org/u/farhad1990)
#### Post date: [January 19, 2021, 3:40pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/13 "2021-01-19T15:40:58Z")

</div>

Hi Mike,

Following your [tutorial](https://forum.qiime2.org/t/processing-filtering-and-evaluating-the-silva-database-and-other-reference-sequence-data-with-rescript/15494#heading--sixth-header) I have made two artifacts, dereplicated sequences and taxa trimmed for 341f and 805r primer sets (ready to be used for training my classifier). Then I've converted them to fasta (the sequence) and tsv (the taxa) files by [qiime exporter](https://docs.qiime2.org/2020.11/tutorials/exporting/#).  
Then I tried to use them as the inputs for the "[dada2:::makeTaxonomyFasta\_SilvaNR(trimmed-seqs.fasta, full-taxa.txt, output=classifier.gz)](https://rdrr.io/github/benjjneb/dada2/src/R/taxonomy.R)". This function uses naiev-bayes method for training. However, reading the taxa.txt file, it kept giving me a format error. Then I tried to compare it with the standard one (full length) from the silva138 website, I realized that the taxa output of qiime has only two columns, Feature ID and Taxon:

 ![image](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/3/3c4b411c0334857a314fc1203d193d756ab756ae.png)

it seems compeletely different than the full taxa file I got from silva138 database:

 ![image](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/b/b7283d52985c340c4b5abb33226545b5d0b9396a.png)

To solve this error, I only used my trimmed (based on the primers) seqs and used the full taxa file instead of the trimmed one. And it ran successfully, but the results I got from this was different than when I used the silva138 full length classifier. Could that be the reason?

---

<div class="post-metadata">

### Author: ![farhad1990](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/f/f05b48/32.png) [@farhad1990](https://forum.qiime2.org/u/farhad1990)
#### Post date: [January 19, 2021, 3:46pm UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/14 "2021-01-19T15:46:37Z")

</div>

Hi,

It is Rstudio in Jupyter notebook.

Kinds,  
Farhad

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [February 20, 2021, 2:17am UTC](https://forum.qiime2.org/t/how-to-convert-a-trained-taxonomy-classifier-to-a-gz-file/18021/17 "2021-02-20T02:17:11Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
