# Using RESCRIPt's 'extract-seq-segments' to extract reference sequences without PCR primer pairs.

**URL:** https://forum.qiime2.org/t/using-rescripts-extract-seq-segments-to-extract-reference-sequences-without-pcr-primer-pairs/23618
**Category:** Tutorials
**Tags:** taxonomy, feature-classifier, trnl, lsu, its, ssu, co1, tutorial, rescript
**Created:** [July 13, 2022, 4:18pm UTC](https://forum.qiime2.org/t/using-rescripts-extract-seq-segments-to-extract-reference-sequences-without-pcr-primer-pairs/23618 "2022-07-13T16:18:36Z")
**Posts on this page:** 1
**Showing post:** 12

<div class="post-metadata">

### Author: ![SoilRotifer](https://forum.qiime2.org/user_avatar/forum.qiime2.org/soilrotifer/32/21071_2.png) [@SoilRotifer](https://forum.qiime2.org/u/SoilRotifer)
#### Post date: [October 26, 2023, 4:30pm UTC](https://forum.qiime2.org/t/using-rescripts-extract-seq-segments-to-extract-reference-sequences-without-pcr-primer-pairs/23618/12 "2023-10-26T16:30:37Z")

</div>

> [@John](#):
>
> This region does not vary significantly, so it seems worrisome. I'm curious if that is what you would expect? Do you think if I filter seqs based on what I would anticipate the max length to be in this database that I would be OK? I'm just worried at how large the sequences have gotten. Perhaps I don't clearly understand how it's working and that I shouldn't worry about these results. In fact, that's what I'm hoping!

Great question! You are right to be concerned about this. This is exactly why I state in the tutorial:

> [@SoilRotifer](#):
>
> Keep in mind you want to balance the acquisition of as many reference sequence segments as you can, against the potential of extracting too many spurious segments that may decrease the effectiveness of your reference database!

This can become more of a problem if you are not cleaning the output after each iteration. The more ambiguous IUPAC bases you have the more spurious things can become. I've built databases for these amplicons (rbcL, ITS2, CO1) too, and have had to filter based on length etc.. **Thus, I'd recommend checking a few of these spuriously long sequences via online BLAST, etc.**

That being said...

It is okay if there is variation, as some of the data in GenBank might not be of high quality. Also, it could be, after these iterations, that you are at the point where you are extracting the fuller-length sequences of the marker genes you are searching for. If this is the case, then it might not be an issue.

Again, as you've noted, the initial PCR primer pair extracts only the amplicon region, the later iterations could expand your reference pool to sequences that contain a longer portion of your marker gene, in which the primer sequences are not a great match (which is why they were not extracted initially), as expected given primer biases etc...

Finally, you can consider starting with a 90% cutoff, then increase by 2-5 % after each iteration, _e.g._ 90, 95, 97, ...

Hopefully, this helps. 🙂

---

_[View the full topic](https://forum.qiime2.org/t/using-rescripts-extract-seq-segments-to-extract-reference-sequences-without-pcr-primer-pairs/23618)._
