# concerns vsearch clustering speed

**URL:** https://forum.qiime2.org/t/concerns-vsearch-clustering-speed/31032
**Category:** User Support
**Tags:** vsearch
**Created:** [August 2, 2024, 9:36am UTC](https://forum.qiime2.org/t/concerns-vsearch-clustering-speed/31032 "2024-08-02T09:36:36Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![Rob\_DNA](https://forum.qiime2.org/user_avatar/forum.qiime2.org/rob_dna/32/18558_2.png) [@Rob\_DNA](https://forum.qiime2.org/u/Rob_DNA)
#### Post date: [August 2, 2024, 9:36am UTC](https://forum.qiime2.org/t/concerns-vsearch-clustering-speed/31032/1 "2024-08-02T09:36:36Z")

</div>

Hi,

I'm clustering ITS2 reads (denoised using DADA2) using vsearch:

```auto
qiime vsearch cluster-features-de-novo \
--i-sequences rep-seqs-dada2.qza \
--i-table table-dada2.qza \
--p-perc-identity 0.97 \
--p-threads 1 \
--o-clustered-table clustered_table.qza \
--o-clustered-sequences clustered_seq.qza

```

In my experience, clustering is often quite an computational intensive and long process.

However, this clustering step is extremely fast, even for ~200 samples with many 100K+ reads for each sample. See the dada2 output stats of the data before clustering:  
[dada2\_output\_stats\_pre-cluster.tsv](https://forum.qiime2.org/uploads/short-url/pRpISNahS85L6tZut9REbteRWGi.tsv) (11.7 KB)

I timed the vsearch command above using `time` and it took only 6 seconds on a single thread of a AMD® Ryzen 9 5900x CPU.

This seems so unlikely for so many samples and reads. Or does the `.qza` format make it very efficient?

I guess not many people complain about some processing step going to fast, but with this speed I get the feeling that it is actually not clustering properly.

So my question: is this a normal time for de novo clustering with `vsearch` of a quite extensive data set?

---

<div class="post-metadata">

### Author: ![SoilRotifer](https://forum.qiime2.org/user_avatar/forum.qiime2.org/soilrotifer/32/21071_2.png) [@SoilRotifer](https://forum.qiime2.org/u/SoilRotifer)
#### Post date: [August 2, 2024, 5:04pm UTC](https://forum.qiime2.org/t/concerns-vsearch-clustering-speed/31032/2 "2024-08-02T17:04:25Z")

</div>

Hi @Rob_DNA,

I suggest you read up on how userach / vsearch works. Different algorithms default for either speed or accuracy, with parameters to adjust them. For example the best case would be to perform an exhaustive search, but that may be untenable in some cases, and lead to days, or weeks of runtime.

Thus, tools like usearch / vsearch will, by default, not perform exhaustive searches and run very fast. They will operate by certain termination or stop criteria. These are often modified by the `maxaccepts` and `maxrejects` options. Once either of these are satisfied the search will stop, and then the next query search will begin.

Adjusting these values will help with better OTU counts and OTU table construction. That is, one issue with usearch / vsearch clustering, is that a sequence might be placed within an OTU seed just because it fits the match criteria (_i.e._ 97%) . However, there may be better OTU seed match for your query downstream (that the algorithm has not found yet). That may be an OTU seed might match your sequence at 99%, and should be placed within that OTU. You can place reads into incorrect OTUs, if the search criteria are not set properly, as the search will terminate before it finds the 'best' match.

_That being said, this may not be an issue given your data! I am just pointing out some things to consider_.

Anyway, You can read more about this [here](https://drive5.com/usearch/manual/usearch_algo.html) and [here](https://drive5.com/usearch/manual/termination_options.html). The vsearch manual is [here](https://vcru.wisc.edu/simonlab/bioinformatics/programs/vsearch/vsearch_manual.pdf).

Sadly, the `maxaccepts` and `maxrejects` options are not currently available via `qiime vsearch cluster-features-de-novo...`. I believe the defaults for vsearch are `maxaccepts=1 maxrejects=32`.

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [August 2, 2024, 6:09pm UTC](https://forum.qiime2.org/t/concerns-vsearch-clustering-speed/31032/3 "2024-08-02T18:09:02Z")

</div>

> [@Rob\_DNA](#):
>
> I guess not many people complain about some processing step going to fast, but with this speed I get the feeling that it is actually not clustering properly.

Remember that _de novo_ clustering was used to make OTUs right after quality filtering. So it was done on all 10-20 million reads on the Illumina run!

> [@Rob\_DNA](#):
>
> In my experience, clustering is often quite an computational intensive and long process.

Yes, clustering is slow when you run it on millions of raw reads or thousands of dereplicated reads.

It's fast when you run it on a few hundred ~~OTUs~~ DADA2 output sequences.

It really is ['orders of magnitude faster than blast'](https://academic.oup.com/bioinformatics/article/26/19/2460/230188?login=false)

---

<div class="post-metadata">

### Author: ![Rob\_DNA](https://forum.qiime2.org/user_avatar/forum.qiime2.org/rob_dna/32/18558_2.png) [@Rob\_DNA](https://forum.qiime2.org/u/Rob_DNA)
#### Post date: [August 5, 2024, 6:27am UTC](https://forum.qiime2.org/t/concerns-vsearch-clustering-speed/31032/5 "2024-08-05T06:27:51Z")

</div>

thanks @SoilRotifer for the great insight!

---

<div class="post-metadata">

### Author: ![Rob\_DNA](https://forum.qiime2.org/user_avatar/forum.qiime2.org/rob_dna/32/18558_2.png) [@Rob\_DNA](https://forum.qiime2.org/u/Rob_DNA)
#### Post date: [August 5, 2024, 6:29am UTC](https://forum.qiime2.org/t/concerns-vsearch-clustering-speed/31032/6 "2024-08-05T06:29:56Z")

</div>

> [@colinbrislawn](#):
>
> It's fast when you run it on a few hundred ~~OTUs~~ DADA2 output sequences.

Aha yes this makes absolutely sense! Thanks.

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [September 5, 2024, 12:30pm UTC](https://forum.qiime2.org/t/concerns-vsearch-clustering-speed/31032/7 "2024-09-05T12:30:51Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
