Hello all,
I'm seeking advice on creating a 16S rDNA database for an avian diet metabarcoding project. I chose the primers 16S1F-degenerate and 16S2R-degenerate (Deagle et al 2007) alongside COI ANML primers to generate 150bp paired end data. I am currently building custom reference databases for each of these primers; my study species is a generalist shorebird so my reasoning was that ANML would target insect prey while 16S would be able to capture other potential prey items like flatworms, crustaceans or fish for example. For the ANML set, I'm following this tutorial using get-ncbi-data. However, I am uncertain about whether I should use NCBI or SILVA when curating sequences for the 16S. My impression is that SILVA is more bacteria focused, and I read this thread where they find that SILVA collapses the taxonomy of some eukaryotes. However, in reading the SILVA manual, it seems more curated than NCBI. Overall, I'm curious what recommendations users who have worked with either datasets would have for creating a metazoan/animal reference dataset. Depending on your recommendations, I also have additional questions:
If I should go the get-ncbi-data route, do you have any pointers on how I can find a list of 16S alternative dictions to use as --p-query search terms? For COI, this tutorial had a nice list. Would something like this be sufficient?: --p-query "txid33208[ORGN] AND (16S OR 16S rDNA OR 16S ribosomal DNA)".
If I should got the get-silva-data route, I understand that the SILVA database is largely in rRNA format, so should I reverse transcribe the results of get-silva-data, then follow this tutorial as normal?
Thank you in advance for your time and help. I am a complete newbie to QIIME and metabarcoding so lengthy explanations are welcome!
Best,
Clare