# Mixed\_oriented\_reads

**URL:** https://forum.qiime2.org/t/mixed-oriented-reads/6242
**Category:** Other Bioinformatics Tools
**Created:** [September 29, 2018, 12:41pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242 "2018-09-29T12:41:57Z")
**Posts on this page:** 15
**Page:** 1

<div class="post-metadata">

### Author: ![amit](https://forum.qiime2.org/user_avatar/forum.qiime2.org/amit/32/20042_2.png) [@amit](https://forum.qiime2.org/u/amit)
#### Post date: [September 29, 2018, 12:41pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/1 "2018-09-29T12:41:57Z")

</div>

Hi,  
I have mixed oriented reads from Illumina and want to put forth this question whether QIIME2 can solve the problem with such data analysis. The main problem is the demultiplexing aand the pre-processing of Miseq reads which depends on the defined orientations.

---

<div class="post-metadata">

### Author: ![colinbrislawn](https://forum.qiime2.org/user_avatar/forum.qiime2.org/colinbrislawn/32/6221_2.png) [@colinbrislawn](https://forum.qiime2.org/u/colinbrislawn)
#### Post date: [September 29, 2018, 5:48pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/2 "2018-09-29T17:48:29Z")

</div>

Unfortunately, Qiime 2 does not have a built-in method for dealing with these. 😞

[https://github.com/qiime2/q2-cutadapt/issues/11](https://github.com/qiime2/q2-cutadapt/issues/11)

You can deal with these outside of QIime then import, but I might be outside the scope of your question. There is some more advice over here:

> [@Problem with demux](https://forum.qiime2.org/t/problem-with-demux/2402):
>
> Dear Qiime2 team, Happy new year! I am trying to learn how to use Qiime2. I have already tried the tutorials and now I am working on my own data. I am dealing with two runs of pair end data. Unfortunately, I did not get a barcode.fastq file from the sequencing company. I could not find a way to transform my data to what is needed for Qiime2 according to the "Atacama tutorial", so I have done the following. I started off with files provided by the company that had already been joined. I then…

May I ask how you came in position of mixed-origination reads? Most Illumina reads are one way.

Colin

---

<div class="post-metadata">

### Author: ![shira](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/s/aeb1de/32.png) [@shira](https://forum.qiime2.org/u/shira)
#### Post date: [September 29, 2018, 6:26pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/3 "2018-09-29T18:26:28Z")

</div>

There is a very straightforward solution to this problem. I came across it after many trials and much reading and asking around.

Basically it involves a single (quick) step in qiime1 using the raw illumina data as input. After that step you can import into qiime2 as paired end sequences, and do your processing there as you would with 'normal' data.

If you are willing to work with qiime1 I am happy to share my code to help you easily overcome this.

---

<div class="post-metadata">

### Author: ![Nicholas\_Bokulich](https://forum.qiime2.org/user_avatar/forum.qiime2.org/nicholas_bokulich/32/19937_2.png) [@Nicholas\_Bokulich](https://forum.qiime2.org/u/Nicholas_Bokulich)
#### Post date: [September 29, 2018, 7:51pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/4 "2018-09-29T19:51:30Z")

</div>

> [@shira](#):
>
> If you are willing to work with qiime1 I am happy to share my code to help you easily overcome this.

Please do share — even if @amit does not want to install qiime1 to accomplish this, other community members may find it useful!

It might also be possible to translate your qiime1 code into a :qiime2: workflow.

Thanks @shira!

---

<div class="post-metadata">

### Author: ![amit](https://forum.qiime2.org/user_avatar/forum.qiime2.org/amit/32/20042_2.png) [@amit](https://forum.qiime2.org/u/amit)
#### Post date: [September 30, 2018, 12:24pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/5 "2018-09-30T12:24:21Z")

</div>

I would be glad if the steps are explained alongwith the script. This will solve all my problems. Thank you so much for your support.

regards,  
Amit.

---

<div class="post-metadata">

### Author: ![shira](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/s/aeb1de/32.png) [@shira](https://forum.qiime2.org/u/shira)
#### Post date: [October 2, 2018, 8:21am UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/6 "2018-10-02T08:21:56Z")

</div>

I am very happy to share the simple solution here, and hope it can help others! I have struggled quite a bit with this until arriving at this solution, and would love to spare the pain from others. I must give credit to @William who gave me the code and patiently answered my many many questions. Obviously it would be fabulous to have this integrated into the qiime2 pipeline 😀, I know there are quite a bit of people struggling with this (@Martin, @emescioglu, @amit - to name a few), and at least one big company producing this kind of mixed-orientation data. Here is the simple code:

# Use qiime1 to sort the sequences into R1 and R2 based on primer sequences, and extract barcodes

extract\_barcodes.py -f SAM1-31\_S2\_L001\_R1\_001.fastq -r SAM1-31\_S2\_L001\_R2\_001.fastq -o ext\_barcodes\_data --bc1\_len 8 --bc2\_len 0 --input\_type barcode\_paired\_end -m 120117SNwhoi341F-mapping.csv --attempt\_read\_reorientation --verbose

#This works like a charm to divide the reads into forward and reverse and extract the barcodes, resulting in true R1 and R2 files, as well as a barcode file for demultiplexing later. The primers are still in the sequence. You will need your R1 and R2 fastq files as well as your mapping file. Adjust the barcode length, in my case it was 8nt and only found on the forward reads.

# Rename, move and Zip files

mkdir data  
mv ext\_barcodes\_data/reads1.fastq data/forward.fastq  
mv ext\_barcodes\_data/reads2.fastq data/reverse.fastq  
mv ext\_barcodes\_data/barcodes.fastq data/  
cd data/  
gzip _._

# Import into QIIME2 and start processing

```auto
qiime tools import \
–type EMPPairedEndSequences \
–input-path data/ \
–output-path emp-paired-end-sequences.qza

qiime demux emp-paired \
–m-barcodes-file sample-metadata.tsv \ 
–m-barcodes-column BarcodeSequence \
–i-seqs emp-paired-end-sequences.qza \
–o-per-sample-sequences demux

qiime demux summarize \
–i-data demux.qza \
–o-visualization demux.qzv

mkdir 277-223

qiime dada2 denoise-paired \
–i-demultiplexed-seqs demux.qza \
–p-trim-left-f 17 \
–p-trim-left-r 21 \
–p-trunc-len-f 277 \
–p-trunc-len-r 223 \
–o-table 277-223/table.qza \
–o-representative-sequences 277-223/rep-seqs.qza \
–o-denoising-stats 277-223/denoising-stats.qza \
–p-n-threads 0 \
–verbose

```

#Note about denoising - as you probably know you will need to try several options to determine how to best trim the reads. I do this by first randomly subsampling the data using seqkit, and it seems crucial to allow enough overlap between the fwd and rev reads. In most cases I see the best results when I leave enough length to have about 50 bp overlap (assuming 450bp contigs), while trimming more from the rev reads that the fwd reads.

I hope this helps!🌈

---

<div class="post-metadata">

### Author: ![amit](https://forum.qiime2.org/user_avatar/forum.qiime2.org/amit/32/20042_2.png) [@amit](https://forum.qiime2.org/u/amit)
#### Post date: [October 3, 2018, 4:42pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/7 "2018-10-03T16:42:10Z")

</div>

Hi @shira is this to be done for Illumina Miseq as well. Kindly help.

regards,  
Amit.

---

<div class="post-metadata">

### Author: ![shira](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/s/aeb1de/32.png) [@shira](https://forum.qiime2.org/u/shira)
#### Post date: [October 3, 2018, 5:07pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/8 "2018-10-03T17:07:48Z")

</div>

Yes, that is the platform my data came from.

---

<div class="post-metadata">

### Author: ![amit](https://forum.qiime2.org/user_avatar/forum.qiime2.org/amit/32/20042_2.png) [@amit](https://forum.qiime2.org/u/amit)
#### Post date: [October 4, 2018, 2:02pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/9 "2018-10-04T14:02:56Z")

</div>

Hi @shira I tried using the extract\_barcode.py command in qiime1. The problem I noticed is that the primers are also removed.  
I gave this command ; extract\_barcodes.py -f Pool4\_S4\_L001\_R1\_001.fastq -r Pool4\_S4\_L001\_R2\_001.fastq -o ext\_barcodes\_data --bc1\_len 7 --bc2\_len 7 --input\_type barcode\_paired\_end -m mapping\_file\_barcodes\_pool4\_new.csv --attempt\_read\_reorientation --verbose.  
It removes the primers. Can you kindly help. I have attached the mapping file information.[mapping\_file\_barcodes\_pool4\_new.csv](https://cdck-file-uploads-global.s3.dualstack.us-west-2.amazonaws.com/flex002/uploads/qiime21/original/2X/c/c6177b9ce7ed92b2570d1b649d80c0700398b7b7.csv) (839 Bytes)

---

<div class="post-metadata">

### Author: ![shira](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/s/aeb1de/32.png) [@shira](https://forum.qiime2.org/u/shira)
#### Post date: [October 5, 2018, 12:38am UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/10 "2018-10-05T00:38:42Z")

</div>

How did you obtain your napping file? Do you have some info on how the libraries were built? Your code is removing 7 nt from both ends of your sequences, are you sure you have barcode there? In general, your mapping file looks odd. Was the same contig amplified in each sample? If so, you primer column should list the same sequence in all rows, at the 5'--3' direction.

---

<div class="post-metadata">

### Author: ![amit](https://forum.qiime2.org/user_avatar/forum.qiime2.org/amit/32/20042_2.png) [@amit](https://forum.qiime2.org/u/amit)
#### Post date: [October 5, 2018, 7:12am UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/11 "2018-10-05T07:12:02Z")

</div>

Hi @shira....The fastq R1 and R2 files were given to me by my collaborators. The mapping file I made to analyze the data are derived out of these sheets (see attachment). The file sample \_library has all the details of the barcode and the data pools. The file Primers are from where the Primers were selected. And I have attached a mapping file for Pool3 which I had used in Qiime1 with the extract\_barcodes.py command. [mapping\_file\_barcodes\_pool3.csv](https://cdck-file-uploads-global.s3.dualstack.us-west-2.amazonaws.com/flex002/uploads/qiime21/original/2X/4/4300cd9edb70900035d5b29379bcbe553265feaa.csv) (743 Bytes)  
[Primers.csv](https://cdck-file-uploads-global.s3.dualstack.us-west-2.amazonaws.com/flex002/uploads/qiime21/original/2X/b/bfe571592b91adcc3a2a264cf87367d6186ac76c.csv) (814 Bytes)  
This amplicon library was from MiSeq. [sample\_library.csv](https://cdck-file-uploads-global.s3.dualstack.us-west-2.amazonaws.com/flex002/uploads/qiime21/original/2X/1/1006afe3728474927d216b5f90465195583fca05.csv) (2.1 KB)  
I KNOW THIS IS TOO MUCH FOR YOU TO LOOK INTO 😒, but kindly help me as I have no clue why the command is removing the primers. Am I making the mapping files correctly?. Kindly respond.

warm regards,  
Amit.

---

<div class="post-metadata">

### Author: ![shira](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/s/aeb1de/32.png) [@shira](https://forum.qiime2.org/u/shira)
#### Post date: [October 5, 2018, 3:12pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/12 "2018-10-05T15:12:56Z")

</div>

Talk to your collaborators to understand how they prepared the sequencing libraries.

There is no mystery in the code. Based on the primer sequences in your mapping file it will sort your reads into forward and reverse, and based on your numeric input it will remove the first x nucleotides from the forward and reverse ends and store threm in a separate barcode file, with reference to the read they are associated with.

It is always useful to take a look at the raw data and see if you understand the structure of the reads. Do they start with the barcode? Does the primer follow? Can you distinguish between  
fwd and rev reads?

---

<div class="post-metadata">

### Author: ![willowblade](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/w/b782af/32.png) [@willowblade](https://forum.qiime2.org/u/willowblade)
#### Post date: [October 10, 2018, 5:50pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/13 "2018-10-10T17:50:57Z")

</div>

Hi,

I've run into this issue as well, and solved it using a tool called sabre. If you want, I can share the code I used with sabre too.

---

<div class="post-metadata">

### Author: ![amit](https://forum.qiime2.org/user_avatar/forum.qiime2.org/amit/32/20042_2.png) [@amit](https://forum.qiime2.org/u/amit)
#### Post date: [October 11, 2018, 9:09am UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/14 "2018-10-11T09:09:09Z")

</div>

Hi @willowblade ....that would be kind on your part....because I am lost 😫

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [November 11, 2018, 3:22pm UTC](https://forum.qiime2.org/t/mixed-oriented-reads/6242/15 "2018-11-11T15:22:44Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
