# loosing reads after dada2 filteration

**URL:** https://forum.qiime2.org/t/loosing-reads-after-dada2-filteration/32373
**Category:** General Discussion
**Created:** [January 23, 2025, 12:28am UTC](https://forum.qiime2.org/t/loosing-reads-after-dada2-filteration/32373 "2025-01-23T00:28:47Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![Sabrin](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/s/59ef9b/32.png) [@Sabrin](https://forum.qiime2.org/u/Sabrin)
#### Post date: [January 23, 2025, 12:28am UTC](https://forum.qiime2.org/t/loosing-reads-after-dada2-filteration/32373/1 "2025-01-23T00:28:47Z")

</div>

Dear forum memebers,

I need your kind help to figure out why I loose almost 90% of reads in 10 samples out of 30 after dada2 **filteration** step.

 ![trimmed-demux](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21907eb3cb3b1ea93b09c69012721adfbb9e7674.png)  
[trimmed-demux.qzv](https://forum.qiime2.org/uploads/short-url/jQrkjkLlKrMBYCNfQwEcZR9VHhY.qzv) (319.9 KB)  
 ![Screenshot 2025-01-23 011936](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/f/b/fb52629f3c6a1bb6f7e7b63bcb3f353ae69869b4.png)  
 ![lost_dada2_filtered](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/7/4/748d3cf10e403763c38540a642b6ce7a96a26717.jpeg)

I did the calculations as follow

# Bacterial 16S amplicon V3/V4 (341\_F/805\_R)

805-341 equals 464 length-of-amplicon  
trunc-len-r + trunc-len-l - length-of-amplicon = overlap  
280 + 200 - 464 = overlap  
16 = overlap  
16 bp gap!! ☹

Please see attached as ref. the quality of my forward and reverse reads.

Many thanks in advance.

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [January 23, 2025, 9:33am UTC](https://forum.qiime2.org/t/loosing-reads-after-dada2-filteration/32373/2 "2025-01-23T09:33:59Z")

</div>

> 805-341 equals 464 length-of-amplicon  
> trunc-len-r + trunc-len-l - length-of-amplicon = overlap  
> 280 + 200 - 464 = overlap  
> 16 = overlap

That's correct! The reads are expected to overlap by 16 basepairs.  
(There is no gap, which is good!)

It looks like many of your reads are merging and there are thousands of reads in most samples, so this may be okay!

If I had data like this, I would continue with analysis and return to this step if I discovered issues later.

---

<div class="post-metadata">

### Author: ![Sabrin](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/s/59ef9b/32.png) [@Sabrin](https://forum.qiime2.org/u/Sabrin)
#### Post date: [January 23, 2025, 3:28pm UTC](https://forum.qiime2.org/t/loosing-reads-after-dada2-filteration/32373/3 "2025-01-23T15:28:05Z")

</div>

@colinbrislawn thank you so much for your quick response!

Unfortunately the 8 samples I loose at the dada2 filteration step is important for my study. They all have over 30k reads but turn into less than 1k after dada2 filteration step.

 ![lost_dada2_filtered](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/4/3/43227bd0e968380e7aeccf35a1a6d647886ff2f4.jpeg)

Any way I could diagnoise why I loose them although I should have enough overlap for merging?

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [January 23, 2025, 7:53pm UTC](https://forum.qiime2.org/t/loosing-reads-after-dada2-filteration/32373/4 "2025-01-23T19:53:15Z")

</div>

> [@Sabrin](#):
>
> Any way I could diagnoise why I loose them although I should have enough overlap for merging?

Yes. Review how the DADA2 algorithm works.

> **[DADA2: High resolution sample inference from Illumina amplicon data](https://pmc.ncbi.nlm.nih.gov/articles/PMC4927377/)**
>
> We present DADA2, a software package that models and corrects Illumina-sequenced amplicon errors. DADA2 infers sample sequences exactly, without coarse-graining into OTUs, and resolves differences of as little as one nucleotide. In several mock ...

Because filtering is the very first step, it's pretty easy to change [DADA2 filtering settings](https://docs.qiime2.org/2024.10/plugins/available/dada2/denoise-paired/) and see how many reads now pass the filter!

> Unfortunately the 8 samples I loose at the dada2 filteration step is important for my study.

Are these samples related in some way?  
Is there any biological reason they may be different from the other samples, like low biomass?

---

<div class="post-metadata">

### Author: ![Sabrin](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/s/59ef9b/32.png) [@Sabrin](https://forum.qiime2.org/u/Sabrin)
#### Post date: [January 23, 2025, 10:10pm UTC](https://forum.qiime2.org/t/loosing-reads-after-dada2-filteration/32373/5 "2025-01-23T22:10:52Z")

</div>

no biological reason, failed samples r random from different treatments while their replicates are ok. But with those failing i end up with one replicate per treatment!

I tried --max\_ee 7 with dada2 but same results, does this means it is not sequencing issue, if yes then what?

Using only forward reads runs ok but then my taxonomic resolution goes bad, right!

---

<div class="post-metadata">

### Author: ![Sabrin](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/s/59ef9b/32.png) [@Sabrin](https://forum.qiime2.org/u/Sabrin)
#### Post date: [January 24, 2025, 1:28am UTC](https://forum.qiime2.org/t/loosing-reads-after-dada2-filteration/32373/6 "2025-01-24T01:28:00Z")

</div>

I tried to use deblur instead of dada2 following this tutorial. It looks like it works to join but eventually i loose like 90% after denoising with deblur but after joining all reads retained.

[deblur-stats\_l\_427.qzv](https://forum.qiime2.org/uploads/short-url/xpwljBd13pyPft91ghHPVrKhM4R.qzv) (214.7 KB)  
[deblur-table\_427.qzv](https://forum.qiime2.org/uploads/short-url/CE6DszmJhUzV83JOP4kCUVxCgG.qzv) (432.0 KB)  
[merged-vsearch-seqs.qzv](https://forum.qiime2.org/uploads/short-url/ot5yMzvktzRLkJlsMrugTolzSbU.qzv) (303.9 KB)  
[filter-stats.qzv](https://forum.qiime2.org/uploads/short-url/8yPjAWNeNRJot1QU9iLV5IpyLu2.qzv) (1.2 MB)

here is steps i followed:

```auto
qiime vsearch merge-pairs \
  --i-demultiplexed-seqs trimmed-demux.qza \
  --o-merged-sequences merged-vsearch-seqs.qza \
  --o-unmerged-sequences unmerged-vsearch-seqs.qza

qiime quality-filter q-score \
  --i-demux merged-vsearch-seqs.qza \
  --o-filtered-sequences filtered-seqs.qza \
  --o-filter-stats filter-stats.qza

  qiime deblur denoise-16S \
      --i-demultiplexed-seqs filtered-seqs.qza \
      --p-trim-length 427 \
      --o-table deblur-table_427.qza \
      --o-representative-sequences deblur-rep-seqs_427.qza \
      --o-stats deblur-stats_l_427.qza

```

however deblur stats file \*qzv looks corrupted , run with this code, **What is wrong here please**?:

```auto
qiime deblur visualize-stats \
--i-deblur-stats deblur-stats_l_427.qza \
--o-visualization deblur-stats_l_427.qzv

```

..  
my main question with this new approach , am I choosing wrong length to trim with deblur?

 ![image](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/f/9/f989f3b8eb18d58a25753f08e42c2d50497b0c1c.png)

when i tried ` --p-trim-length 385` as that was the minimum length at subsampling I got much better results , but still not sure if that is the best choice to move forward with!  
[deblur-table\_385.qzv](https://forum.qiime2.org/uploads/short-url/2zeLcIOwornrOuQIPSE5PYnq4it.qzv) (465.2 KB)

Extremely grateful for your assistance ☺

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [January 25, 2025, 4:08pm UTC](https://forum.qiime2.org/t/loosing-reads-after-dada2-filteration/32373/7 "2025-01-25T16:08:25Z")

</div>

> [@Sabrin](#):
>
> no biological reason, failed samples r random from different treatments while their replicates are ok.

Okay, good to know!

> [@Sabrin](#):
>
> however deblur stats file \*qzv looks corrupted , run with this code, **What is wrong here please**?:

Yeah, I don't see the data columns either. Perhaps you can try to run that command again and see if the rerun fixes the file?

I'm not very familiar with the deblur plugin, so I will not offer any advice here.

I think trying DADA2 denoise-single for just your forward reads is a good idea.

> [@Sabrin](#):
>
> Using only forward reads runs ok but then my taxonomic resolution goes bad, right!

The taxonomic resolution may be reduced, but it's not 'bad'. The first amplicon studies using the Illumina miseq used \<100 bp reads and they got published. Your forward reads alone are more then twice that.

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [January 25, 2025, 4:12pm UTC](https://forum.qiime2.org/t/loosing-reads-after-dada2-filteration/32373/8 "2025-01-25T16:12:40Z")

</div>

Here is one last clue that may be related to your problem:

It's by benjjneb, who is the developer of DADA2, and related to the V3-V4 primers, like you are using:

> [@Questions about V3–V4 primers for 16S rRNA amplicon sequencing, and calculating overlap](https://forum.qiime2.org/t/questions-about-v3-v4-primers-for-16s-rrna-amplicon-sequencing-and-calculating-overlap/20250/2):
>
> Yes, there is variation in lengths of 16S segments. It isn't large, but it exists. In particular, there are two modes of V3V4 length in nature, one at 460 nts and another ~440 nts (using these primers). So you'll want to make sure even the longer natural amplicons will sufficiently overlap after truncation.

If these 8 samples happen to have lots of microbes with the longer ~460 bp amplicon, then they need that extra length to merge.
