# Multi-gene amplicon sequences with dada2

**URL:** https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402
**Category:** User Support
**Tags:** dada2
**Created:** [October 3, 2017, 8:53pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402 "2017-10-03T20:53:55Z")
**Posts on this page:** 13
**Page:** 1

<div class="post-metadata">

### Author: ![Stream\_biofilm](https://forum.qiime2.org/user_avatar/forum.qiime2.org/stream_biofilm/32/535_2.png) [@Stream\_biofilm](https://forum.qiime2.org/u/Stream_biofilm)
#### Post date: [October 3, 2017, 8:53pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/1 "2017-10-03T20:53:55Z")

</div>

Hello,

We amplified four gene markers (16S, 18S, 23S, rbcL) for each sample and combined all four markers before adding the barcode and sequencing. I now have demultiplexed sequences but each sample still contains sequences of all four amplicons. Do I need to create individual files for each sample and marker before processing with dada2? Each amplicon is a different size but I figure I can use the primer to parse them out. Is there anything in place to assist me with this?

Thanks,  
Danny

---

<div class="post-metadata">

### Author: ![ebolyen](https://forum.qiime2.org/user_avatar/forum.qiime2.org/ebolyen/32/11_2.png) [@ebolyen](https://forum.qiime2.org/u/ebolyen)
#### Post date: [October 9, 2017, 5:55pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/3 "2017-10-09T17:55:17Z")

</div>

Hi @Stream_biofilm,

Sorry for the delayed response. I'm afraid I can't speak to whether DADA2 can handle multiple amplicons (of different regions) at once, but hopefully @benjjneb can tell us. In any-case, supposing you did denoise multiple amplicons, you'd still need to separate the results as those features aren't directly comparable. So you are probably best served by separating your data first, then processing each amplicon separately.

Unfortunately, we don't have any functionality to separate mixed amplicons (with shared barcordes) at this time (and we don't have any immediate plans either). You may be able to do something with `cutadapt` or similar tools, but we can't really help you with that here.

Sorry I couldn't be of more help!

---

<div class="post-metadata">

### Author: ![Stream\_biofilm](https://forum.qiime2.org/user_avatar/forum.qiime2.org/stream_biofilm/32/535_2.png) [@Stream\_biofilm](https://forum.qiime2.org/u/Stream_biofilm)
#### Post date: [October 9, 2017, 7:39pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/5 "2017-10-09T19:39:54Z")

</div>

Thanks for the response, I have a script that allows me to parse out each gene for each sample and make separate fastq files. However the result consists of similar but different numbers of forward and reverse reads for each gene within a sample. When I put the individual forward and reverse reads of a gene within a sample into dada2 I got an error -\> "(1) Filtering Error in fastqPairedFilter(c(unfiltsF[[i]], unfiltsR[[i]]), c(filteredFastqF, : Mismatched forward and reverse sequence files: 59521, 59049. ". I assume this error is due to the unequal number of forward and reverse reads using my script.

---

<div class="post-metadata">

### Author: ![benjjneb](https://forum.qiime2.org/user_avatar/forum.qiime2.org/benjjneb/32/1602_2.png) [@benjjneb](https://forum.qiime2.org/u/benjjneb)
#### Post date: [October 9, 2017, 7:43pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/6 "2017-10-09T19:43:01Z")

</div>

> I assume this error is due to the unequal number of forward and reverse reads using my script.

Yes that is the problem. Is there a way to modify the script so that it only keeps F/R reads if they both get assigned to the same locus?

There is an option to check read IDs and match F/R pairs based on that (rather than read order) in the dada2 R package, but its not accessible through the Q2 plugin right now. Right now the plugin requires that the F/R reads match each other (same number, same order).

---

<div class="post-metadata">

### Author: ![Stream\_biofilm](https://forum.qiime2.org/user_avatar/forum.qiime2.org/stream_biofilm/32/535_2.png) [@Stream\_biofilm](https://forum.qiime2.org/u/Stream_biofilm)
#### Post date: [October 10, 2017, 3:51pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/7 "2017-10-10T15:51:03Z")

</div>

> [@benjjneb](#):
>
> match F/R pairs based on that (rather than read order) in the dada2 R package, but its not accessible through the Q2 plugin right now. Right now the plugin requires that the F/R reads match each other (same number, same order).

Would I be able to do this using the dada2 R package and then transfer the files back into qiime2? Also, would there be an issue on running dada2 with multiple genes not parsed and then separating them downstream?

---

<div class="post-metadata">

### Author: ![benjjneb](https://forum.qiime2.org/user_avatar/forum.qiime2.org/benjjneb/32/1602_2.png) [@benjjneb](https://forum.qiime2.org/u/benjjneb)
#### Post date: [October 10, 2017, 5:22pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/9 "2017-10-10T17:22:49Z")

</div>

> [@Stream\_biofilm](#):
>
> Would I be able to do this using the dada2 R package and then transfer the files back into qiime2?

Yes\* using the filterAndTrim(..., matchIDs=TRUE) function. I could help you with that at the dada2 R package forum if you'd like to post an issue there. I also wonder if one of the Q2 masters might know if a method that can do this is already included in Q2?

^ If they are in a normal format, e.g. w/ Illumina IDs. Also we do plan on bringing a lot of these functions into Q2 once Pipelines are implemented.

> [@Stream\_biofilm](#):
>
> Also, would there be an issue on running dada2 with multiple genes not parsed and then separating them downstream?

If they were sequenced together, then that would probably be just fine. The algorithm should easily separate the different gene regions from each other.

---

<div class="post-metadata">

### Author: ![jairideout](https://forum.qiime2.org/user_avatar/forum.qiime2.org/jairideout/32/9_2.png) [@jairideout](https://forum.qiime2.org/u/jairideout)
#### Post date: [October 10, 2017, 5:54pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/10 "2017-10-10T17:54:47Z")

</div>

> [@benjjneb](#):
>
> I also wonder if one of the Q2 masters might know if a method that can do this is already included in Q2?

Unfortunately I don't know of a method in qiime2 that handles paired-end data that's not in the same record order. That would be great to expose in `q2-dada2` when qiime2 is able to support pipelines (I think that's blocking this feature right now?).

---

<div class="post-metadata">

### Author: ![ebolyen](https://forum.qiime2.org/user_avatar/forum.qiime2.org/ebolyen/32/11_2.png) [@ebolyen](https://forum.qiime2.org/u/ebolyen)
#### Post date: [October 10, 2017, 7:34pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/11 "2017-10-10T19:34:58Z")

</div>

> [@Stream\_biofilm](#):
>
> Would I be able to do this using the dada2 R package and then transfer the files back into qiime2?

> [@benjjneb](#):
>
> I also wonder if one of the Q2 masters might know if a method that can do this is already included in Q2?

I wrote [a post](https://forum.qiime2.org/t/dada2-back-to-qiime-2/1316/11) for importing your data after following the DADA2 tutorial. Getting the `FeatureData[Sequence]` artifact is pretty easy, but the `FeatureTable[Frequency]` takes a bit of massaging.

> [@benjjneb](#):
>
> The algorithm should easily separate the different gene regions from each other.

That is super cool!

> [@Stream\_biofilm](#):
>
> Also, would there be an issue on running dada2 with multiple genes not parsed and then separating them downstream?

I can't think of a good way to separate them after the fact in QIIME 2. So splitting them, then using DADA2 directly (with `matchIDs=TRUE`) may be your best option.

In the future we might be able to frame this problem as removing "contamination" (@Nicholas_Bokulich is working/thinking about plugins for that), but I couldn't say for sure.

---

<div class="post-metadata">

### Author: ![Stream\_biofilm](https://forum.qiime2.org/user_avatar/forum.qiime2.org/stream_biofilm/32/535_2.png) [@Stream\_biofilm](https://forum.qiime2.org/u/Stream_biofilm)
#### Post date: [October 10, 2017, 7:48pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/13 "2017-10-10T19:48:19Z")

</div>

> [@benjjneb](#):
>
> If they were sequenced together, then that would probably be just fine. The algorithm should easily separate the different gene regions from each other.

Yes they are sequenced together. I tried to run the reads through q2-dada2 without separating out the genes. It produced a warning message: "Self-consistency loop terminated before convergence" after it hit selfConsist step 10. I believe I read somewhere on this forum about that warning being okay depending on the error rate

> [@ebolyen](#):
>
> I wrote a post for importing your data after following the DADA2 tutorial. Getting the FeatureData[Sequence] artifact is pretty easy, but the FeatureTable[Frequency] takes a bit of massaging.

Thanks! It sounds like this might be the best option to try now. Ill report back with my progress.

I did some more research and found using the Illumina bcl2fastq program I might be able to take the raw multiplexed sequences and demultiplex them using their software but it looks a little confusing.

---

<div class="post-metadata">

### Author: ![Stream\_biofilm](https://forum.qiime2.org/user_avatar/forum.qiime2.org/stream_biofilm/32/535_2.png) [@Stream\_biofilm](https://forum.qiime2.org/u/Stream_biofilm)
#### Post date: [October 11, 2017, 5:40pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/14 "2017-10-11T17:40:53Z")

</div>

Is there any reason why I cant just demultiplex the samples twice? First time using the sample index and then take those and demultiplex again using the primer sequence for each gene within the sample?

---

<div class="post-metadata">

### Author: ![ebolyen](https://forum.qiime2.org/user_avatar/forum.qiime2.org/ebolyen/32/11_2.png) [@ebolyen](https://forum.qiime2.org/u/ebolyen)
#### Post date: [October 12, 2017, 10:36pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/16 "2017-10-12T22:36:42Z")

</div>

> [@Stream\_biofilm](#):
>
> Is there any reason why I cant just demultiplex the samples twice? First time using the sample index and then take those and demultiplex again using the primer sequence for each gene within the sample?

If I'm understanding correctly, you are suggesting using the primer as a barcode for the second demultiplexing step?

It'll probably be more work than it's worth to do it with QIIME 2 since `demux emp-single/paired` assumes that your "barcodes" are the same length, and that there is a barcodes.fastq.gz file with the "barcodes" in exactly the same order. It seems like you would need to script something to meet those assumptions, and at that point you may as well do the demultiplexing yourself. Not to mention all of the importing/exporting you would need to do.

---

<div class="post-metadata">

### Author: ![Stream\_biofilm](https://forum.qiime2.org/user_avatar/forum.qiime2.org/stream_biofilm/32/535_2.png) [@Stream\_biofilm](https://forum.qiime2.org/u/Stream_biofilm)
#### Post date: [October 17, 2017, 7:23pm UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/18 "2017-10-17T19:23:27Z")

</div>

Just an update on this issue. I was able to use `repair.sh` from BBTools to fix my forward and reverse fastq files putting them in the same order and making sure only sequences found in both files remain.  
Thanks for the help everyone!  
Danny

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [November 18, 2017, 1:23am UTC](https://forum.qiime2.org/t/multi-gene-amplicon-sequences-with-dada2/1402/19 "2017-11-18T01:23:35Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
