# Training feature classifier: values for --p-trunc-len; --p-min-length; and --p-max-length

**URL:** https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596
**Category:** User Support
**Created:** [February 13, 2020, 5:07am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596 "2020-02-13T05:07:47Z")
**Posts on this page:** 20
**Page:** 1

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 13, 2020, 5:07am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/1 "2020-02-13T05:07:47Z")

</div>

In the code below:

qiime feature-classifier extract-reads   
--i-sequences 85\_otus.qza   
--p-f-primer GTGCCAGCMGCCGCGGTAA   
--p-r-primer GGACTACHVGGGTWTCTAAT   
--p-trunc-len 120   
--p-min-length 100   
--p-max-length 400   
--o-reads ref-seqs.qza

If we trimmed the forward and reverse sequences differently at the dada2 step, what --p-trunc-len should we use here? So for me, the forward read was not cut and stayed at 250nts, while the reverse was lower quality and so was cut at 244.

Also, how do we choose values for the --p-min-length and --p-max-length?

Thanks

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [February 13, 2020, 4:57pm UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/2 "2020-02-13T16:57:38Z")

</div>

Hi @Negin,  
Please see the tutorial, which has answers to all these questions and more! See the "notes" in this section:  
[https://docs.qiime2.org/2019.10/tutorials/feature-classifier/#extract-reference-reads](https://docs.qiime2.org/2019.10/tutorials/feature-classifier/#extract-reference-reads)

---

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 13, 2020, 5:52pm UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/3 "2020-02-13T17:52:53Z")

</div>

Hi Nicholas,

Thank you for your response. So does the sentence below from the notes of the tutorial mean that I would need to use --p-trunc-len of 250 since my original sequences are 250nts, so basically I should use the size before trimming?

"For classification of paired-end reads and untrimmed single-end reads, we recommend training a classifier on sequences that have been extracted at the appropriate primer sites, but are not trimmed."

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [February 13, 2020, 5:54pm UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/4 "2020-02-13T17:54:23Z")

</div>

> [@Negin](#):
>
> So does the sentence below from the notes of the tutorial mean that I would need to use --p-trunc-len of 250 since my original sequences are 250nts, so basically I should use the size before trimming?

No — since you are using paired-end sequences you should not truncate the extracted reference sequences.

---

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 13, 2020, 6:00pm UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/5 "2020-02-13T18:00:32Z")

</div>

oh okay thank you! So this means that I should probably go with the default options for –p-trunc-len, –p-min-length and –p-max-length, so just leave these arguments out.

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [February 13, 2020, 6:11pm UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/6 "2020-02-13T18:11:53Z")

</div>

for paired-end reads, do not truncate the reference sequences with `extract-reads`

min and max length are another story, though. You should check the literature to see what the expected size range is for your primer set — or just switch these off and then check the length of the extracted sequences to see what the length distribution is, and decide for yourself if there are abnormally short or long sequences that need to be winnowed out.

Good luck!

---

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 13, 2020, 6:16pm UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/7 "2020-02-13T18:16:06Z")

</div>

Thank you for your help!

---

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 13, 2020, 7:52pm UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/8 "2020-02-13T19:52:38Z")

</div>

I extracted the reads using my primers with default options and used the code below to visualize, but when I try to visualize it, it gives me a blank page. I thought these were sequences so I could use qiime feature-table tabulate-seqs but probably not. How else can I look at the file to see what length the extracted sequences are?

qiime feature-table tabulate-seqs   
--i-data qza/silva\_132\_99\_v3v4\_eub-euf\_extracted.qza   
--o-visualization qzv/silva\_132\_99\_v3v4\_eub-euf\_extracted.qzv &

The file seem to be too big for me to upload here.

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [February 14, 2020, 2:32pm UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/9 "2020-02-14T14:32:08Z")

</div>

> [@Negin](#):
>
> I thought these were sequences so I could use qiime feature-table tabulate-seqs but probably not.

You can use `tabulate-seqs` — I am not sure why the page will not load, maybe a browser issue? Or the file is too large to display?

---

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 14, 2020, 10:32pm UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/10 "2020-02-14T22:32:19Z")

</div>

> [@Negin](#):
>
> –p-max-length

Yes, the file is large indeed!

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [February 14, 2020, 10:48pm UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/11 "2020-02-14T22:48:44Z")

</div>

Okay that makes sense... also makes sense that it would be really large if you are extracting sequences from a reference database (as opposed to a collection of ASVs or OTUs from a real dataset). So I think this is effectively a browser issue, the file may be too large to load, which we occasionally see e.g., with really large emperor plots.

Try this: just extract the QZV file and grab the length distribution summary like this:

```bash
$ qiime tools extract --input-path rep-seqs.qzv --output-path .
Extracted rep-seqs.qzv to directory 789ea3c6-8ac4-442a-adbd-d80738359b71
$ head 789ea3c6-8ac4-442a-adbd-d80738359b71/data/seven_number_summary.tsv 
Quantile	Value
0.02	120
0.09	120
0.25	120
0.5	120
0.75	120
0.91	120
0.98	120

```

Note: you will need to modify the filepath to reflect the ID that is printed to the screen; so see how I got this message: `Extracted rep-seqs.qzv to directory 789ea3c6-8ac4-442a-adbd-d80738359b71` and then used that ID as the directory name in the following line.

---

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 15, 2020, 12:03am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/12 "2020-02-15T00:03:34Z")

</div>

Hi Nicholas,

I ran these codes and there isn't a seven\_number\_summary in the data folder. I checked there physically in addition to running the code. Strangely, there is one in my downloads folder from 7 days ago and I don't remember generating that. But anyways, there isn't one related to the task I just ran. Maybe there is an issue with the file I created.

I was trying to find the normal range for the amplicon for my primer (EUBF-EUBR) in the literature too and I was not very successful. I found this link that seems to show between 100-500 for v3v4 which should work for me although my primer is a bit longer than the one shown here:  
[https://help.ezbiocloud.net/comparison-between-v3v4-and-full-length-sequencing-of-16s-rrna-genes/](https://help.ezbiocloud.net/comparison-between-v3v4-and-full-length-sequencing-of-16s-rrna-genes/)

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [February 15, 2020, 12:41am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/13 "2020-02-15T00:41:57Z")

</div>

> [@Negin](#):
>
> I ran these codes and there isn’t a seven\_number\_summary in the data folder. I checked there physically in addition to running the code.

Sounds like you running an older release of QIIME 2. This length summary was added a couple releases ago, I believe.

> [@Negin](#):
>
> found this link that seems to show between 100-500 for v3v4 which should work for me although my primer is a bit longer

As long as the primers are hitting the same site, you can go off that info

Another good place to get info like this is the forum! Here is a recent topic describing expected length for V3V4, though the length range is not stated, only (presumably) the mean:

> [@DADA2, truncation parameters question](https://forum.qiime2.org/t/dada2-truncation-parameters-question/13627):
>
> Hello, I have a question about truncation parameters for a 2x300 bp paired end run using the V3-V4 region. These results are with the primers removed (341F and 805R). Here is the visualization of demux.qzv after trimming primers using cutadapt. [Picture1] My first data point is with reads that have been truncated at --p-trunc-len-f 281 and -p--trunc-len-r 202. Theoretically, these limits should be fine given that the expected amplicon size is around ~460 bp. This should leave enough for the ~…

Based on these findings, 300-600 is probably a fine, permissive range for you to use (though in practice the variance is probably much less, since most 16S regions don't have that much length variation)

---

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 15, 2020, 12:50am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/14 "2020-02-15T00:50:47Z")

</div>

I am using qiime2/2019.10 which should be the latest version?

Regarding the amplicon size, I have reads that are below 300 in my data. My average read length was 320 nts. What would happen to those that are smaller. Would they get removed in the further taxonomy assignment?

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [February 15, 2020, 12:57am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/15 "2020-02-15T00:57:28Z")

</div>

> [@Negin](#):
>
> Regarding the amplicon size, I have reads that are below 300 in my data. My average read length was 320 nts. What would happen to those that are smaller.

okay, so maybe set 200 nt as a lower bound. Sounds like you may be using different primers compared to that other topic.

---

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 15, 2020, 1:18am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/16 "2020-02-15T01:18:28Z")

</div>

I am using these primers:

Forward: EUB\_F 5'-TCCTACGGGAGGCAGCAGT (19 nts) ​  
Reverse: EUB\_R 5'-GGACTACCAGGGTATCTAATCCTGTT (26 nts)

---

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 15, 2020, 1:33am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/17 "2020-02-15T01:33:45Z")

</div>

Actually, I had used the wrong file. Sorry about that. Here is the range that seems reasonable. Does that mean that I don't need to redo the extract reads with max and min parameters cause I had them at default?

 ![image](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/6/64994ef975c06e5012b61212551f34d330fcfd0c.png)

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [February 15, 2020, 1:47am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/18 "2020-02-15T01:47:03Z")

</div>

> [@Negin](#):
>
> Does that mean that I don’t need to redo the extract reads with max and min parameters cause I had them at default?

That is correct!

Good luck!

---

<div class="post-metadata">

### Author: ![Negin](https://forum.qiime2.org/user_avatar/forum.qiime2.org/negin/32/5030_2.png) [@Negin](https://forum.qiime2.org/u/Negin)
#### Post date: [February 15, 2020, 2:17am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/19 "2020-02-15T02:17:16Z")

</div>

Great thanks! 🙂

---

<div class="post-metadata">

### Author: ![Daryl](https://forum.qiime2.org/user_avatar/forum.qiime2.org/daryl/32/7293_2.png) [@Daryl](https://forum.qiime2.org/u/Daryl)
#### Post date: [February 16, 2020, 3:22am UTC](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596/20 "2020-02-16T03:22:09Z")

</div>

> [@Nicholas\_Bokulich](#):
>
> You can use `tabulate-seqs` — I am not sure why the page will not load, maybe a browser issue? Or the file is too large to display?

yes,this is a really browser problem . The Google browser can works with it

[Next page](https://forum.qiime2.org/t/training-feature-classifier-values-for-p-trunc-len-p-min-length-and-p-max-length/13596.md?page=2)
