# Issues with choosing classifiers on feature-classifier classify-sklearn

**URL:** https://forum.qiime2.org/t/issues-with-choosing-classifiers-on-feature-classifier-classify-sklearn/28948
**Category:** User Support
**Tags:** greengenes2
**Created:** [January 18, 2024, 5:20pm UTC](https://forum.qiime2.org/t/issues-with-choosing-classifiers-on-feature-classifier-classify-sklearn/28948 "2024-01-18T17:20:29Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![UnevenCuttlefish](https://forum.qiime2.org/user_avatar/forum.qiime2.org/unevencuttlefish/32/17842_2.png) [@UnevenCuttlefish](https://forum.qiime2.org/u/UnevenCuttlefish)
#### Post date: [January 18, 2024, 5:20pm UTC](https://forum.qiime2.org/t/issues-with-choosing-classifiers-on-feature-classifier-classify-sklearn/28948/1 "2024-01-18T17:20:29Z")

</div>

Hi all!  
I am currently working with an environmentally collected dataset run with SE Illumina Miseq 515F 926R primers and am trying to create the classifiers to use but they are showing quite wildly different results and I would like some clarification on a few things.

firstly I had issues with the greengenes2 on my first time through with this dataset as my samples were kept at 250bp and gg2 doesn't have tips out that far, only around the 150bp length. So I went ahead and went through my analysis with my old classifier and did just fine.

Now I'm trying to make sure that my old classifier didn't run into any issues so I'm creating two new ones to compare my results with - BOTH at the 150bp length (since gg2 can't go out further than that). below is my process for creating the two classifiers.

**Greengenes2 classifier**  
files:  
[150bp\_gg\_classified\_taxonomy.qza](https://forum.qiime2.org/uploads/short-url/dXRZsj5tfovMCGnWLeHDAZJaJuD.qza) (412.6 KB)  
[visualized\_150bpgg2\_taxonomy.qzv](https://forum.qiime2.org/uploads/short-url/cPdmwP1gH2i6Z9d61V2uYInh8rI.qzv) (2.3 MB)  
(unfortunately the classifier itself was too large to upload here)

qiime feature-classifier extract-reads  
--i-sequences 2022.10.seqs.fna.qza  
--p-f-primer GTGCCAGCMGCCGCGGTAA  
--p-r-primer CCGYCAATTYMTTTRAGTTT  
--p-min-length 0  
--p-max-length 0  
--p-trunc-len 150  
--p-n-jobs 8  
--o-reads gg2-ref-seqs

qiime feature-classifier fit-classifier-naive-bayes  
--i-reference-reads gg2-ref-seqs.qza  
--i-reference-taxonomy reference-taxonomy.qza  
--o-classifier TBJ\_150bp\_gg2-classifier

qiime feature-classifier classify-sklearn  
--i-classifier TBJ\_150bp\_gg2-classifier.qza  
--i-reads gg2-ref-seqs.qza  
--o-classification test\_classification

qiime feature-classifier classify-sklearn  
--i-classifier taxonomy/TBJ\_150bp\_gg2-classifier.qza  
--i-reads taxonomy/rep\_seqs\_deblur\_150nt.qza  
--p-n-jobs -1  
--o-classification taxonomy/150bp\_gg\_classified\_taxonomy

qiime metadata tabulate  
--m-input-file taxonomy/150bp\_gg\_classified\_taxonomy.qza  
--o-visualization visualized\_150bpgg2\_taxonomy

For brevity, here is what the assigned taxonomy visualizes to for gg2

 ![Screenshot from 2024-01-18 09-13-49](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/d/7/d75c5f976903ae7726e5cb12b91a81e929cbf1c6.png)

**Silva\_138**  
Files:  
[TBJ\_Silva\_138\_classifier.qza](https://forum.qiime2.org/uploads/short-url/3tDlDXCEcIRexQNAm1dLMnqalMR.qza) (124.5 KB)  
[visualized\_silva138\_actual\_taxonomy.qzv](https://forum.qiime2.org/uploads/short-url/xlDp9JouYxw2fbXZYfpeLbKB2br.qzv) (2.0 MB)  
[silva150bp\_assigned\_taxonomy.qza](https://forum.qiime2.org/uploads/short-url/7xROiVEjutudFQYJ13LHspjfHbd.qza) (360.4 KB)

qiime feature-classifier extract-reads  
--i-sequences silva-138-99-seqs-515-806.qza  
--p-f-primer GTGCCAGCMGCCGCGGTAA  
--p-r-primer CCGYCAATTYMTTTRAGTTT  
--p-min-length 50  
--p-max-length 250  
--p-n-jobs 10  
--o-reads extracted\_silva\_138\_reads

qiime feature-classifier fit-classifier-naive-bayes  
--i-reference-reads extracted\_silva\_138\_reads.qza  
--i-reference-taxonomy silva-138-99-tax-515-806.qza  
--p-classify--chunk-size 30000  
--o-classifier TBJ\_Silva\_138\_classifier

qiime feature-classifier classify-sklearn  
--i-classifier TBJ\_Silva\_138\_classifier.qza  
--i-reads extracted\_silva\_138\_reads.qza  
--p-n-jobs -1  
--o-classification test\_classification

qiime metadata tabulate  
--m-input-file test\_classification.qza  
--o-visualization test\_taxonomyy\_silva138\_vis

here is the visualization for the Silva\_138 classifier

 ![Screenshot from 2024-01-18 09-15-29](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/3/3/33f51ae1c2144ae023a246f0ee7d3e27b2a1c6b7.png)

of most obvious note is the different parameters with the classifier creation itself, I changed them when I had such long computation time originally, but I don't think that should have that much of a difference other than possible the truncate length on gg2, correct? I can redo these to match exactly if needed, that is no issue. but I have a feeling something else isn't correct for what I've done. The only issue I can think of was my original files I took to create these classifiers was from the [QIIME2 docs page](https://docs.qiime2.org/2023.9/data-resources/) using the 515F806R that were available there.

I can remove the unassigned in the Silva workflow in subsequent steps, but I am wanting to confirm there is an error here before I move on. Thank you in advance!

-UC

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [January 18, 2024, 6:47pm UTC](https://forum.qiime2.org/t/issues-with-choosing-classifiers-on-feature-classifier-classify-sklearn/28948/2 "2024-01-18T18:47:52Z")

</div>

Hi @UnevenCuttlefish ,

The issue is coming from this command:

> [@UnevenCuttlefish](#):
>
> qiime feature-classifier extract-reads  
> --i-sequences silva-138-99-seqs-515-806.qza  
> --p-f-primer GTGCCAGCMGCCGCGGTAA  
> --p-r-primer CCGYCAATTYMTTTRAGTTT  
> --p-min-length 50  
> --p-max-length 250  
> --p-n-jobs 10  
> --o-reads extracted\_silva\_138\_reads

Why are you setting `max-length` to 250? This is shorter than the V4 amplicon that is amplified by the 515-806 primers is in most species. Anything longer than that will be dropped — i.e., most sequences.

So you are creating a tiny classifier with only a few sequences in it, hence why you are only hitting Streptococcus or Unassigned.

You did this with the SILVA classifier but not with the GG classifier, so you are sort of comparing 🍎 s and 🟠 or maybe more like 🍎 s and 🐙 s.

Anyway, increasing the max-length to an appropriate value in that action (like \> 300 nt) should do the trick.

Good luck!

---

<div class="post-metadata">

### Author: ![UnevenCuttlefish](https://forum.qiime2.org/user_avatar/forum.qiime2.org/unevencuttlefish/32/17842_2.png) [@UnevenCuttlefish](https://forum.qiime2.org/u/UnevenCuttlefish)
#### Post date: [January 18, 2024, 7:10pm UTC](https://forum.qiime2.org/t/issues-with-choosing-classifiers-on-feature-classifier-classify-sklearn/28948/3 "2024-01-18T19:10:28Z")

</div>

Thank you! I figured that's where the error was coming from. I had misread the max-length parameter when I had originally done it. I wanted to make sure that was the error before I did it over again.

---

<div class="post-metadata">

### Author: ![wasade](https://forum.qiime2.org/user_avatar/forum.qiime2.org/wasade/32/2317_2.png) [@wasade](https://forum.qiime2.org/u/wasade)
#### Post date: [January 18, 2024, 8:08pm UTC](https://forum.qiime2.org/t/issues-with-choosing-classifiers-on-feature-classifier-classify-sklearn/28948/4 "2024-01-18T20:08:26Z")

</div>

Hi @UnevenCuttlefish,

Using all ~20 million sequences in Greengenes2 2022.10, where most are fragments, for training the classifier could have unexpected effects. Why not use either the existing full length model, or train on just the full length backbone sequences?

While it's true we primarily placed 90, 100 and 150bp sequences in Greengenes2 2022.10, the full length classifier isn't constrained to those lengths. The full length classifier can be obtained [here](https://docs.qiime2.org/2023.9/data-resources/). It's plausible the V4 classifier would just work too though the rev primer is a little different.

All the best,  
Daniel

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [February 19, 2024, 2:09am UTC](https://forum.qiime2.org/t/issues-with-choosing-classifiers-on-feature-classifier-classify-sklearn/28948/5 "2024-02-19T02:09:10Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
