# Picking values for --p-min-length and --p-max-length in qiime feature-classifier extract-reads

**URL:** https://forum.qiime2.org/t/picking-values-for-p-min-length-and-p-max-length-in-qiime-feature-classifier-extract-reads/20912
**Category:** User Support
**Tags:** feature-classifier
**Created:** [September 28, 2021, 8:24am UTC](https://forum.qiime2.org/t/picking-values-for-p-min-length-and-p-max-length-in-qiime-feature-classifier-extract-reads/20912 "2021-09-28T08:24:18Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![KQUB](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/k/5f9b8f/32.png) [@KQUB](https://forum.qiime2.org/u/KQUB)
#### Post date: [September 28, 2021, 8:24am UTC](https://forum.qiime2.org/t/picking-values-for-p-min-length-and-p-max-length-in-qiime-feature-classifier-extract-reads/20912/1 "2021-09-28T08:24:18Z")

</div>

Hi there,

We’ve done some 16S rRNA amplicon sequencing with primers that Illumina says are used to sequence the V3 and V4 variable regions of the 16S rRNA gene.

Forward primer:  
5’-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG **CCTACGGGNGGCWGCAG** -3’

Reverse primer:  
5’-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG **GACTACHVGGGTATCTAATCC** -3’

Here's my `qiime dada2 denoise-paired` command:

```
qiime dada2 denoise-paired \
--i-demultiplexed-seqs demux.qza \
--p-trim-left-f 17 \
--p-trim-left-r 21 \
--p-trunc-len-f 294 \
--p-trunc-len-r 216 \
--o-table table.qza \
--o-representative-sequences rep-seqs.qza \
--o-denoising-stats stats.qza \
--p-n-threads 8 \
--verbose

```

And here's my `qiime feature-classifier extract-reads` command:

```
qiime feature-classifier extract-reads \
  --i-sequences silva-138-99-seqs.qza \
  --p-f-primer CCTACGGGNGGCWGCAG \
  --p-r-primer GACTACHVGGGTATCTAATCC \
  --p-min-length 400 \
  --p-max-length 500 \
  --o-reads ref-seqs.qza

```

My understanding is that the primers we're using produce amplicons of ~464 bp in length (see [this forum post](https://forum.qiime2.org/t/questions-about-v3-v4-primers-for-16s-rrna-amplicon-sequencing-and-calculating-overlap/20250)). So, by using `--p-trim-left-f 17` and `--p-trim-left-r 21` in the `qiime dada2 denoise-paired` step, I'd end up with amplicons of length (464 − 17 − 21) = **426 bp**. Is that correct? If so, what would be the best values to use for `--p-min-length` and `--p-max-length` in `qiime feature-classifier extract-reads`?

I notice that [someone used `--p-min-length 400` and `--p-max-length 450` in a similar situation (with the same primers and similar trimming)](https://forum.qiime2.org/t/v3v4-trim-length-and-length-parameters-for-extract-reads/10119/4), and got a thumbs up from @Mehrbod_Estaki. When I ran my analysis initially, I used `--p-min-length 400` and `--p-max-length 500` in the `qiime feature-classifier extract-reads` command, but I guess I'm wondering if there's any significant difference between using, say ...

- a **tight** interval like: `--p-min-length 420` and `--p-max-length 430`
- or, a **wider** interval like: `--p-min-length 400` and `--p-max-length 450`
- or, an **even wider** interval like: `--p-min-length 400` and `--p-max-length 500`

Is there any 'rule of thumb' people use for this? Or does it even matter very much?

Some relevant info (from the `qiime feature-classifier extract-reads` [usage page](https://docs.qiime2.org/2021.8/plugins/available/feature-classifier/extract-reads/)):

```
--p-min-length INTEGER Minimum amplicon length. Shorter amplicons are
    Range(0, None) discarded. Applied after trimming and truncation, so
                          be aware that trimming may impact sequence
                          retention. Set to zero to disable min length
                          filtering. [default: 50]
--p-max-length INTEGER Maximum amplicon length. Longer amplicons are
    Range(0, None) discarded. Applied before trimming and truncation,
                          so plan accordingly. Set to zero (default) to
                          disable max length filtering. [default: 0]

```

Thanks as always for the help! 😊

Kevin

---

<div class="post-metadata">

### Author: ![Mehrbod\_Estaki](https://forum.qiime2.org/user_avatar/forum.qiime2.org/mehrbod_estaki/32/4001_2.png) [@Mehrbod\_Estaki](https://forum.qiime2.org/u/Mehrbod_Estaki)
#### Post date: [September 28, 2021, 9:00pm UTC](https://forum.qiime2.org/t/picking-values-for-p-min-length-and-p-max-length-in-qiime-feature-classifier-extract-reads/20912/2 "2021-09-28T21:00:56Z")

</div>

Hi @KQUB,

The min/max parameters are used to trim the "extracted" primer-specific region from your reference database, in your case SILVA. So, as long as your reference sequences encompass the whole region extracted by your primers it should be fine. While it's been shown that a classifier trained on a specific region can improve classification a bit, I'm not sure anyone has benchmarked the effect of that additional trimming. My gut feeling is as long as the reference is equal or longer than your query sequence, those small length differences wouldn't really affect classification. My recommendation for a region such as v3-v4 that has variable lengths, is to just not do any additional trimming.

---

<div class="post-metadata">

### Author: ![KQUB](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/k/5f9b8f/32.png) [@KQUB](https://forum.qiime2.org/u/KQUB)
#### Post date: [September 29, 2021, 5:54am UTC](https://forum.qiime2.org/t/picking-values-for-p-min-length-and-p-max-length-in-qiime-feature-classifier-extract-reads/20912/3 "2021-09-29T05:54:47Z")

</div>

Hi, @Mehrbod_Estaki.

Thanks for your response!

> [@Mehrbod\_Estaki](#):
>
> My gut feeling is as long as the reference is equal or longer than your query sequence, those small length differences wouldn't really affect classification.

So, you think that, when running the `qiime feature-classifier extract-reads` command, the difference between something like `--p-min-length 400` and `--p-max-length 450` or `--p-min-length 400` and `--p-max-length 500` isn't very significant?

> [@Mehrbod\_Estaki](#):
>
> My recommendation for a region such as v3-v4 that has variable lengths, is to just not do any additional trimming.

Do you mean that your current recommendation, when using primers targeting the V3–V4 region, is to set both `--p-min-length` and `--p-max-length` to zero?

Thanks again for your time!

---

<div class="post-metadata">

### Author: ![Deni\_Ribicic](https://forum.qiime2.org/user_avatar/forum.qiime2.org/deni_ribicic/32/17065_2.png) [@Deni\_Ribicic](https://forum.qiime2.org/u/Deni_Ribicic)
#### Post date: [September 29, 2021, 12:43pm UTC](https://forum.qiime2.org/t/picking-values-for-p-min-length-and-p-max-length-in-qiime-feature-classifier-extract-reads/20912/4 "2021-09-29T12:43:36Z")

</div>

Hi Kevin,

I think your assumption regarding the amplicon length is correct. But for training the classifier I wouldn't go with too tight interval params. You may be risking removing biologically relevant sequences due to differences in 16S rRNA gene in different prokaryotic groups (as you mention also in your previous post).

I am pretty comfortable with `--p-min-length 400` and `--p-max-length 500` as this gives me enough error margin, and certainty that I am recapturing relevant sequences and excluding some larger artefacts 🙂

---

<div class="post-metadata">

### Author: ![KQUB](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/k/5f9b8f/32.png) [@KQUB](https://forum.qiime2.org/u/KQUB)
#### Post date: [September 29, 2021, 12:51pm UTC](https://forum.qiime2.org/t/picking-values-for-p-min-length-and-p-max-length-in-qiime-feature-classifier-extract-reads/20912/5 "2021-09-29T12:51:42Z")

</div>

Thanks for your input, Deni! 😊

---

<div class="post-metadata">

### Author: ![Mehrbod\_Estaki](https://forum.qiime2.org/user_avatar/forum.qiime2.org/mehrbod_estaki/32/4001_2.png) [@Mehrbod\_Estaki](https://forum.qiime2.org/u/Mehrbod_Estaki)
#### Post date: [September 30, 2021, 6:26am UTC](https://forum.qiime2.org/t/picking-values-for-p-min-length-and-p-max-length-in-qiime-feature-classifier-extract-reads/20912/6 "2021-09-30T06:26:06Z")

</div>

Hi @KQUB,

> [@KQUB](#):
>
> something like `--p-min-length 400` and `--p-max-length 450` or `--p-min-length 400` and `--p-max-length 500` isn't very significant?

I agree with @Deni_Ribicic here, I think in your scenario you shouldn't use such strict parameters. Either don't trim at all (what I would do) or use something relaxed such as @Deni_Ribicic's recommended 400/500. This is especially important when targeting a variable region like the V3-V4, the last thing you want is to introduce bias of a certain clade that is either longer or shorter than your arbitrary parameters.

> [@KQUB](#):
>
> Do you mean that your current recommendation, when using primers targeting the V3–V4 region, is to set both `--p-min-length` and `--p-max-length` to zero?

That's what I would, to be play it safe 🤷‍♂️ .

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [October 31, 2021, 12:26pm UTC](https://forum.qiime2.org/t/picking-values-for-p-min-length-and-p-max-length-in-qiime-feature-classifier-extract-reads/20912/7 "2021-10-31T12:26:37Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
