# How to train the classifier for V3-V4 region with 99% identity using full length seuqnces from new relase of GreenGenes-2022??

**URL:** https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991
**Category:** Library Support
**Tags:** taxonomy, feature-classifier, greengenes
**Created:** [October 18, 2023, 6:37am UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991 "2023-10-18T06:37:45Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![Rashmi\_Ira](https://forum.qiime2.org/user_avatar/forum.qiime2.org/rashmi_ira/32/10224_2.png) [@Rashmi\_Ira](https://forum.qiime2.org/u/Rashmi_Ira)
#### Post date: [October 18, 2023, 6:37am UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991/1 "2023-10-18T06:37:45Z")

</div>

Hello All,

I want to perform functional annotation using **Picrust2 plugin**.

> Blockquote

First of all I want to train my **classifier for V3-V4 region with 99%** identity for latest release of **greengenes-2022**

> Blockquote. I tried to check files from [Index of /greengenes\_release/2022.10](http://ftp.microbio.me/greengenes_release/2022.10/) new release of GG-2022 but confused that which file is of my use!!!  
> Also, I am facing problem to use my QIIME2 outputs as a input file file for PICRUSt2. Please help me out for this, so that i can start my further data processing and analysis.

Any help will be appreciated.

Thanks and Regards,

Rashmi Ira

---

<div class="post-metadata">

### Author: ![wasade](https://forum.qiime2.org/user_avatar/forum.qiime2.org/wasade/32/2317_2.png) [@wasade](https://forum.qiime2.org/u/wasade)
#### Post date: [October 18, 2023, 8:22pm UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991/2 "2023-10-18T20:22:02Z")

</div>

Hi @Rashmi_Ira,

I think what you would want to do is use `q2-feature-classifier` to `extract-reads` based on your primers from the Greengenes2 backbone sequences, and then train a Naive Bayes classifier on the result

Best,  
Daniel

---

<div class="post-metadata">

### Author: ![Rashmi\_Ira](https://forum.qiime2.org/user_avatar/forum.qiime2.org/rashmi_ira/32/10224_2.png) [@Rashmi\_Ira](https://forum.qiime2.org/u/Rashmi_Ira)
#### Post date: [October 19, 2023, 9:49am UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991/3 "2023-10-19T09:49:27Z")

</div>

Hi @wasade

Thanks for your response.  
Okay I will check it out.

> [@wasade](#):
>
> use `q2-feature-classifier` to `extract-reads` based on your primers from the Greengenes2

Can you please guide me with few commands and also which file is actually of my use to start with read extraction!??!

Thanks in advance.

Best Regards,  
Rashmi Ira

---

<div class="post-metadata">

### Author: ![buzic](https://forum.qiime2.org/user_avatar/forum.qiime2.org/buzic/32/16494_2.png) [@buzic](https://forum.qiime2.org/u/buzic)
#### Post date: [October 23, 2023, 2:31pm UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991/7 "2023-10-23T14:31:26Z")

</div>

Hi,

the readme file looks like you'd need the following files:

2022.10.backbone.full-length.fna.qza  
2022.10.backbone.tax.qza

The following method should help point you in the right direction. First take the sequence files and trim them based on your primers (obviously I've just added a random sequence here!). You can also add truncation and min/max lengths of sequences based on your experimental design, for example:

```auto
qiime feature-classifier extract-reads \
  --i-sequences 2022.10.backbone.full-length.fna.qza \
  --p-f-primer GTGGTGGTGGTGGTGGTG \
  --p-r-primer GGACTGGACTGGACTGGA \
  --p-min-length 100 \
  --p-max-length 600 \
  --o-reads gg_12_10_ref_primer_region_seqs.qza

```

then use your newly trimmed sequence file along with the backbone taxonomy to train your classifier:

```auto
qiime feature-classifier fit-classifier-naive-bayes \
  --i-reference-reads gg_12_10_ref_primer_region_seqs.qza \
  --i-reference-taxonomy 2022.10.backbone.tax.qza \
  --o-classifier gg_12_10_primer_region-classifier.qza

```

I hope that helps, there are lots of walkthroughs and helpful documents in the qiime2 forum and docs, for example [here](https://docs.qiime2.org/2023.9/tutorials/feature-classifier/)

---

<div class="post-metadata">

### Author: ![Rashmi\_Ira](https://forum.qiime2.org/user_avatar/forum.qiime2.org/rashmi_ira/32/10224_2.png) [@Rashmi\_Ira](https://forum.qiime2.org/u/Rashmi_Ira)
#### Post date: [October 23, 2023, 4:06pm UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991/8 "2023-10-23T16:06:10Z")

</div>

Hi @buzic

Thank you so much for your response. I will follow the same as suggested.

Best Regards,  
Rashmi Ira

---

<div class="post-metadata">

### Author: ![stephhhhanniee](https://forum.qiime2.org/user_avatar/forum.qiime2.org/stephhhhanniee/32/20911_2.png) [@stephhhhanniee](https://forum.qiime2.org/u/stephhhhanniee)
#### Post date: [June 14, 2024, 10:49pm UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991/9 "2024-06-14T22:49:33Z")

</div>

Hi Victoria,

Thank you so much for this example! I used this for my samples, but I was wondering how you know the % identity. Does this code default to 99% identity?

I read elsewhere in the forum that you can specify the % identity, but it was using a different code?

Any clarification would be greatly appreciated. Thank you!

Stephanie

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [June 15, 2024, 2:25pm UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991/10 "2024-06-15T14:25:28Z")

</div>

> [@stephhhhanniee](#):
>
> but I was wondering how you know the % identity.

You don't! These two commands use the same features from the input database. Perhaps the input database was clustered at a set identify, or not!

Also, the act of selecting an internal region will make previous calculations of identity invalid.

> Does this code default to 99% identity?

No.

There are tools designed to assist with database curation, including RESCRIPt:  
[https://library.qiime2.org/plugins/rescript/27/](https://library.qiime2.org/plugins/rescript/27/)

---

<div class="post-metadata">

### Author: ![stephhhhanniee](https://forum.qiime2.org/user_avatar/forum.qiime2.org/stephhhhanniee/32/20911_2.png) [@stephhhhanniee](https://forum.qiime2.org/u/stephhhhanniee)
#### Post date: [June 17, 2024, 1:31am UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991/12 "2024-06-17T01:31:11Z")

</div>

Hi Colin,

Thanks so much for the explanation! So the "act of selecting an internal region" is like specifying the primers for the V3-V4 region, for example?

Also, after I commented I was looking at my taxonomy output file and the "Confidence" column ranges from 0.72-0.99 so is that in some way connected to the % identity? I ran a different code (qiime greengenes2 non-v4-16s) which has 99% identity and my "Confidence" column was all 1.0 so I was wondering if that's connected to the % identity.

Thanks again!

Stephanie

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [June 17, 2024, 1:05pm UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991/13 "2024-06-17T13:05:06Z")

</div>

> [@stephhhhanniee](#):
>
> Thanks so much for the explanation! So the "act of selecting an internal region" is like specifying the primers for the V3-V4 region, for example?

That's right. And that's what these commands do:

```auto
qiime feature-classifier extract-reads ...
and
qiime rescript extract-seq-segments

```

* * *

> [@stephhhhanniee](#):
>
> the "Confidence" column ranges from 0.72-0.99 so is that in some way connected to the % identity? I ran a different code (qiime greengenes2 non-v4-16s) which has 99% identity and my "Confidence" column was all 1.0 so I was wondering if that's connected to the % identity.

That column is the confidence (think 'confidence interval') of a query sequence's annotation, not the database sequence's identity.

Different methods report confidence differently, so it depends on the program. It's never database pre-cluster distance, though.

---

<div class="post-metadata">

### Author: ![stephhhhanniee](https://forum.qiime2.org/user_avatar/forum.qiime2.org/stephhhhanniee/32/20911_2.png) [@stephhhhanniee](https://forum.qiime2.org/u/stephhhhanniee)
#### Post date: [June 20, 2024, 7:20pm UTC](https://forum.qiime2.org/t/how-to-train-the-classifier-for-v3-v4-region-with-99-identity-using-full-length-seuqnces-from-new-relase-of-greengenes-2022/27991/14 "2024-06-20T19:20:57Z")

</div>

Oh gotcha, thank you again for your help, I appreciate it!
