# DADA2 filtering stats

**URL:** https://forum.qiime2.org/t/dada2-filtering-stats/3735
**Category:** User Support
**Tags:** pending-development
**Created:** [April 11, 2018, 4:56pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735 "2018-04-11T16:56:17Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![nerdynella](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nerdynella/32/21728_2.png) [@nerdynella](https://forum.qiime2.org/u/nerdynella)
#### Post date: [April 11, 2018, 4:56pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735/1 "2018-04-11T16:56:18Z")

</div>

Continuing the discussion from [Summary statistics after dada2](https://forum.qiime2.org/t/summary-statistics-after-dada2/1860/11):

> [@Summary statistics after dada2](https://forum.qiime2.org/t/summary-statistics-after-dada2/1860/11):
>
> The DADA2 plugin in QIIME 2 2017.12 now prints how many reads were filtered, denoised, merged, and non-chimeric to stdout! (You can capture that via the --verbose flag.)

@ebolyen does this mean we need to rerun DADA2 in v 2017.12 or later to get the stats? Is there a way to still obtain these stats from outputs that were obtained using an earlier q2 version? DADA2 took forever to run on my data and we'd rather not try to rerun it to obtain the stats - please tell me there's another way

P.S. could you please update the tutorial page to show that this option is now available?

Nsa

---

<div class="post-metadata">

### Author: ![nerdynella](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nerdynella/32/21728_2.png) [@nerdynella](https://forum.qiime2.org/u/nerdynella)
#### Post date: [April 11, 2018, 6:53pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735/2 "2018-04-11T18:53:51Z")

</div>

while i await a response, I have gone ahead and downloaded the freq/sample csv from my resulting DADA2 table.

I am multiplying the reported seq/sample by 2 (since my seqs were PE) to obtain the reads retained per sample after DADA2. is this approach correct? does the seq/sample reported in the csv file account for reads that were combined into unique sequences during the dereplication stage ?

---

<div class="post-metadata">

### Author: ![thermokarst](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/t/8e7dd6/32.png) [@thermokarst](https://forum.qiime2.org/u/thermokarst)
#### Post date: [April 12, 2018, 2:10pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735/3 "2018-04-12T14:10:37Z")

</div>



---

<div class="post-metadata">

### Author: ![nerdynella](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nerdynella/32/21728_2.png) [@nerdynella](https://forum.qiime2.org/u/nerdynella)
#### Post date: [April 12, 2018, 6:09pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735/4 "2018-04-12T18:09:29Z")

</div>

i think i got one answer.  
using the moving pictures tutorial data, i ran DADA2 and compared the output obtained by passing `--verbose` to those in the resulting DADA2 table (freq/sample csv), and it appears that the data presented in the freq/sample csv is the output of the final DADA2 step - filtered, dereplicated, non-chimeric reads.

Now, `--verbose` only shows stats for the first 6 samples, how do you visualize stats for the remaining samples?

I am still hoping that there's a way to obtain these stats from the outputs of a previous DADA2 run (ours took about a week to complete on our largest high mem node) and I don't think I'll be lucky again to have access to this node for that long.

Thanks,  
Nsa

---

<div class="post-metadata">

### Author: ![ebolyen](https://forum.qiime2.org/user_avatar/forum.qiime2.org/ebolyen/32/11_2.png) [@ebolyen](https://forum.qiime2.org/u/ebolyen)
#### Post date: [April 16, 2018, 10:30pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735/5 "2018-04-16T22:30:42Z")

</div>

Hi @nerdynella!

Sorry for the very delayed response.

> [@nerdynella](#):
>
> Is there a way to still obtain these stats from outputs that were obtained using an earlier q2 version? DADA2 took forever to run on my data and we’d rather not try to rerun it to obtain the stats - please tell me there’s another way

☹ I'm afraid not.

> [@nerdynella](#):
>
> P.S. could you please update the tutorial page to show that this option is now available?

The `--verbose` was very much a stopgap, we've got a better plan coming.

> [@nerdynella](#):
>
> I am multiplying the reported seq/sample by 2 (since my seqs were PE) to obtain the reads retained per sample after DADA2. is this approach correct? does the seq/sample reported in the csv file account for reads that were combined into unique sequences during the dereplication stage ?

No, the frequencies provided for `denoise-paired` are what you should use. The PE data is merged into a single sequence which is then counted once each time it is observed.

Just to double check, what sequencing instrument did you use? On Illumina the primers aren't independent, so your forward and reverse reads represent the same "sampling event" on the instrument.

> [@nerdynella](#):
>
> Now, --verbose only shows stats for the first 6 samples, how do you visualize stats for the remaining samples?

That will be coming soon once we've fixed [this issue](https://github.com/qiime2/q2-dada2/issues/79). The plan is to have that information be an artifact which you can visualize as metadata or use otherwise.

---

<div class="post-metadata">

### Author: ![ebolyen](https://forum.qiime2.org/user_avatar/forum.qiime2.org/ebolyen/32/11_2.png) [@ebolyen](https://forum.qiime2.org/u/ebolyen)
#### Post date: [April 16, 2018, 10:30pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735/6 "2018-04-16T22:30:53Z")

</div>



---

<div class="post-metadata">

### Author: ![nerdynella](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nerdynella/32/21728_2.png) [@nerdynella](https://forum.qiime2.org/u/nerdynella)
#### Post date: [April 17, 2018, 3:49pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735/7 "2018-04-17T15:49:22Z")

</div>

> [@ebolyen](#):
>
> No, the frequencies provided for denoise-paired are what you should use. The PE data is merged into a single sequence which is then counted once each time it is observed.

that's what i thought, but since the last column (non-chimeric reads) is what is presented in the freq/sample csv i was hopping that it could somehow be used to calculate the filtered/denoised data

---

<div class="post-metadata">

### Author: ![nerdynella](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nerdynella/32/21728_2.png) [@nerdynella](https://forum.qiime2.org/u/nerdynella)
#### Post date: [April 17, 2018, 3:50pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735/8 "2018-04-17T15:50:08Z")

</div>

> [@ebolyen](#):
>
> Just to double check, what sequencing instrument did you use? On Illumina the primers aren’t independent, so your forward and reverse reads represent the same “sampling event” on the instrument.

yes, we used the Illumina HiSeq platform, and thanks for the heads up

---

<div class="post-metadata">

### Author: ![nerdynella](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nerdynella/32/21728_2.png) [@nerdynella](https://forum.qiime2.org/u/nerdynella)
#### Post date: [April 17, 2018, 3:50pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735/9 "2018-04-17T15:50:38Z")

</div>

> [@ebolyen](#):
>
> ☹ I’m afraid not.

oh well...thanks anyway.

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [May 18, 2018, 9:50pm UTC](https://forum.qiime2.org/t/dada2-filtering-stats/3735/10 "2018-05-18T21:50:42Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
