Would you be able to track down a log file for one of these runs? I think we should see the number of unique sequences printed to stdout which may be helpful in finding the issue.
I see that you even used a higher max-ee value, will try to use a higher one on mine to see if that helps. Additionally, I even used 'indepedent' pooling method but still came up short.
Perhaps a question for everyone, can one still proceed with downstream analyses despite this significant loss of reads after filtering?
Hi Mukhari_R,
I tried multiple max error, like 2, 5 but it did not make much difference in the reads that passed the final filter.
I am wondering if that might be the sequence quality itself. I am not entirely sure how that could happen.
Low rates of successful denoising, which is where you are losing most of your 16S reads, can happen when there is little duplication in the data, either due to relatively high error rates (more of an issue with long reads) or very high levels of sample diversity (so even real biological sequences are only being observed a handful of times or less). But what is surprising is that you aren't seeing the same issue in your ITS sequences.
Were there any differences in the library preparation/sequencing of your 16S data versus your ITS data? For example, was Kinnex used for one and not the other? How long is the ITS region you are amplifying?