# Error while getting plenty of data from NCBI

**URL:** https://forum.qiime2.org/t/error-while-getting-plenty-of-data-from-ncbi/17635
**Category:** Library Support
**Tags:** rescript
**Created:** [November 29, 2020, 1:53pm UTC](https://forum.qiime2.org/t/error-while-getting-plenty-of-data-from-ncbi/17635 "2020-11-29T13:53:59Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![TurboQiimer](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/t/46a35a/32.png) [@TurboQiimer](https://forum.qiime2.org/u/TurboQiimer)
#### Post date: [November 29, 2020, 1:53pm UTC](https://forum.qiime2.org/t/error-while-getting-plenty-of-data-from-ncbi/17635/1 "2020-11-29T13:53:59Z")

</div>

Hi again guys,

It sounds parameter `--p-n-jobs` with digit 5 is unable to work in my case when a lot of data is downloading.

While I wanted to download lots of data with gene name, I got this error (it take around 20 min to get the error):

qiime rescript get-ncbi-data \

> ```
> --p-query dsrB \
> --o-sequences RefsequencesdsrB.qza \
> 
> ```
> 
> --p-n-jobs 5   
> --o-taxonomy 11.qza  
> Plugin error from rescript:

A worker process managed by the executor was unexpectedly terminated. This could be caused by a segmentation fault while calling the function or by an excessive memory usage causing the Operating System to kill the worker.

The exit codes of the workers are {SIGKILL(-9)}

Debug info has been saved to /tmp/qiime2-q2cli-err-3qh5xqo6.log

In the guideline it mentions to digit 5 " **If you are downloading lots of data, you should set the `--p-n-jobs` parameter to a number greater than one. Five works well in most cases**." It failed then!

The only way that I catch all data is gene name, because accession number and GI are exclusively unique, then I used gene name in order to get all data. Would it be the point? How the problem could be solved? What do you think?

Thanks a lot.  
Qiimer

---

<div class="post-metadata">

### Author: ![TurboQiimer](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/t/46a35a/32.png) [@TurboQiimer](https://forum.qiime2.org/u/TurboQiimer)
#### Post date: [November 29, 2020, 4:41pm UTC](https://forum.qiime2.org/t/error-while-getting-plenty-of-data-from-ncbi/17635/2 "2020-11-29T16:41:07Z")

</div>

I used a server (remote machine)to prevent memory-associated limitation. After around two hours, it reported an unknown error.

qiime rescript get-ncbi-data \

> ```
> --p-query dsrB \
> --o-sequences RefsequencesdsrB.qza \
> 
> ```
> 
> --p-n-jobs 5   
> --o-taxonomy 11.qza  
> Plugin error from rescript:

**Download did not finish. Reason unknown.**

Debug info has been saved to /tmp/qiime2-q2cli-err-we4d1npy.log

I attached this result to complete my report.

Best regards,  
Qiimer

---

<div class="post-metadata">

### Author: ![SoilRotifer](https://forum.qiime2.org/user_avatar/forum.qiime2.org/soilrotifer/32/21071_2.png) [@SoilRotifer](https://forum.qiime2.org/u/SoilRotifer)
#### Post date: [November 30, 2020, 8:45pm UTC](https://forum.qiime2.org/t/error-while-getting-plenty-of-data-from-ncbi/17635/3 "2020-11-30T20:45:31Z")

</div>

> [@TurboQiimer](#):
>
> A worker process managed by the executor was unexpectedly terminated. This could be caused by a segmentation fault while calling the function or by an excessive memory usage causing the Operating System to kill the worker.
> 
> The exit codes of the workers are {SIGKILL(-9)}

It looks like you figured this out as per:

> [@TurboQiimer](#):
>
> I used a server (remote machine)to prevent memory-associated limitation

Your next error:

> [@TurboQiimer](#):
>
> Download did not finish. Reason unknown.

can be problematic to figure out. There are many _"out of our hands"_ reasons that fetching data from NCBI can fail:

> [@rescript get-ncbi-data ‘dict’ object has no attribute ‘add’](https://forum.qiime2.org/t/rescript-get-ncbi-data-dict-object-has-no-attribute-add/16560/5):
>
> Hey @the_dummy, Just to follow on @SoilRotifer's comments to clarify — we see various server-side issues with NCBI so sometimes using get-ncbi-data can be a bit bumpy depending on: the time of day the size/content of your query which way the wind is blowing (just kidding wink) Usually trying again later works (e.g., if you are trying to run the job during peak hours) and we have some pending changes to rescript to make these failures more graceful in the future. Good luck!

I'd suggest trying to add more selective terms to your query, or try to download in batches. This is a great place to start:

> [@Building a COI database from NCBI references](https://forum.qiime2.org/t/building-a-coi-database-from-ncbi-references/16500):
>
> construction stop_signconstruction stop_signconstruction stop_signconstruction stop_signconstruction stop_signCitation: If you use the following COI resources or RESCRIPt for COI database preparation, please cite the following: Michael S Robeson II, Devon R O’Rourke, Benjamin D Kaehler, Michal Ziemski, Matthew R Dillon, Jeffrey T Foster, Nicholas A Bokulich. RESCRIPt: Reproducible sequence taxonomy reference database management for the masses. bioRxiv 2020.10.05.326504; d…

Good luck!

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [January 1, 2021, 2:45am UTC](https://forum.qiime2.org/t/error-while-getting-plenty-of-data-from-ncbi/17635/4 "2021-01-01T02:45:37Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
