# Using QIIME2 to view 454 pyrosequencing data

**URL:** https://forum.qiime2.org/t/using-qiime2-to-view-454-pyrosequencing-data/17500
**Category:** General Discussion
**Tags:** import
**Created:** [November 18, 2020, 6:28am UTC](https://forum.qiime2.org/t/using-qiime2-to-view-454-pyrosequencing-data/17500 "2020-11-18T06:28:54Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![perdita](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/p/ebca7d/32.png) [@perdita](https://forum.qiime2.org/u/perdita)
#### Post date: [November 18, 2020, 6:28am UTC](https://forum.qiime2.org/t/using-qiime2-to-view-454-pyrosequencing-data/17500/1 "2020-11-18T06:28:54Z")

</div>

Hi everyone, thanks for taking the time to read this.

I have some 454 pyrosequencing data from years ago, that was processed by Research and Testing. I would really like to take a look at it on QIIME 2 using an EC2 instance, but I am having a hard time. I converted the .fna and .qual files into a .fastq on QIIME(1) using convert\_fastaqual\_fastq.py, with the hope of using that to see what I could see. Before attempting to run anything though, I want to make sure that my methodology makes sense, and I'm hoping that someone more familiar with with these file types could tell me what I'm looking at, here.

 ![Screen Shot 2020-11-17 at 19.22.42](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/3/30d08f9d1f4bdf425d87d2e01333e6b562b89e56.png) This is an overview of RaT processing that was done before data delivery at the time, from the document that arrived with the samples.

When I look at the raw data I see lines like this: A-03-[Primer 1]::M02233:62:000000000-A9GLW:1:1116:21355:24484 in the sequence identifier / description fields. This sample identifier might show up in a few entries but with different numbers at the end. It looks like 8 base pair barcodes are present at the beginning of the sequences (5' end) and in quality data, and the primers were removed? Is there an easy way to see whether this is single end or paired end? I'm trying to figure out whether it's demultiplexed, and whether I can import this kind of data into QIIME 2 and have any hope of getting meaningful interpretations out of it. This isn't EMP or Casava 1.8 data from what I can see, but is the generated fastq file (from fna + qual) best fit to multiplexed fastq data, or by using a manifest file?

Any thoughts or advice on this would be very welcome. I really appreciate your time. 🙂

---

<div class="post-metadata">

### Author: ![andrewsanchez](https://forum.qiime2.org/user_avatar/forum.qiime2.org/andrewsanchez/32/20160_2.png) [@andrewsanchez](https://forum.qiime2.org/u/andrewsanchez)
#### Post date: [November 18, 2020, 11:19pm UTC](https://forum.qiime2.org/t/using-qiime2-to-view-454-pyrosequencing-data/17500/2 "2020-11-18T23:19:00Z")

</div>

Hi, @perdita!

> [@perdita](#):
>
> This is an overview of RaT processing that was done before data delivery at the time, from the document that arrived with the samples.

Please correct me if I am mistaken, but it seems to me that the only relevant portion of this diagram is the upper left corner which shows that the 454 Sequencer generated .fna/.qual files. Given that you have the .fna/.qual files, and no\* the results of the rest of that analysis process (right?), I don't think the rest of the diagram is relevant. Does that make sense or am I missing something?

> [@perdita](#):
>
> Is there an easy way to see whether this is single end or paired end?

It appears that the question is not so straightforward. Please see this [discussion on biostars](https://www.biostars.org/p/111047/).

I need to do a bit more research to answer your other questions confidently. But elsewhere on the forum, people have discussed importing this type of data using a manifest file. For example, see this post: [Importing 454 data to run dada2 denoise-pyro - #3 by Mdavrandi](https://forum.qiime2.org/t/importing-454-data-to-run-dada2-denoise-pyro/7371/3)

You might also look into using q2-cutadapt for demultiplexing if required.

Also, depending on what you need to do, you might have [other options](https://forum.qiime2.org/t/importing-454-raw-reads/2469/3).

---

<div class="post-metadata">

### Author: ![perdita](https://forum.qiime2.org/letter_avatar_proxy/v4/letter/p/ebca7d/32.png) [@perdita](https://forum.qiime2.org/u/perdita)
#### Post date: [November 19, 2020, 12:01am UTC](https://forum.qiime2.org/t/using-qiime2-to-view-454-pyrosequencing-data/17500/3 "2020-11-19T00:01:39Z")

</div>

> [@andrewsanchez](#):
>
> Please correct me if I am mistaken, but it seems to me that the only relevant portion of this diagram is the upper left corner which shows that the 454 Sequencer generated .fna/.qual files. Given that you have the .fna/.qual files, and no\* the results of the rest of that analysis process (right?), I don't think the rest of the diagram is relevant. Does that make sense or am I missing something?

Thank you for the reply! You are correct that I am picking up at the .fna/.qual node of the flowchart, however other versions of this lab's charts include "FASTA / Qual Prepared for QIIME" as an end step from the "Quality Checking / Demultiplexing" node, so I am trying to figure out what I am working with.

 ![Screenshot_20201014_181408_methodology_overview](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/7/733871c2cb53703f6f3ffab19a42e2d103d0fce0.png)

From [Importing and Demultiplexing Sequence Data Quick Reference](https://forum.qiime2.org/t/importing-and-demultiplexing-sequence-data-quick-reference/14002), and the [QIIME2 Import Tutorial](https://docs.qiime2.org/2020.8/tutorials/importing/#sequence-data-with-sequence-quality-information-i-e-fastq), it appears that I need to know if the data are multiplexed or not.

Where I am at:  
✅ Barcodes in Sequence (If they are only at the beginning, does that always imply single-end?)  
❓ Multiplexed? I think it is multiplexed, as several sequences share the same Barcode?  
❓ Need to strip out Barcodes to a metadata file for QIIME2 import?

Thanks for your time and input.

---

<div class="post-metadata">

### Author: ![andrewsanchez](https://forum.qiime2.org/user_avatar/forum.qiime2.org/andrewsanchez/32/20160_2.png) [@andrewsanchez](https://forum.qiime2.org/u/andrewsanchez)
#### Post date: [November 20, 2020, 11:46pm UTC](https://forum.qiime2.org/t/using-qiime2-to-view-454-pyrosequencing-data/17500/4 "2020-11-20T23:46:27Z")

</div>

Sounds like you are on the right track, @perdita!

> [@perdita](#):
>
> ❓ Multiplexed? I think it is multiplexed, as several sequences share the same Barcode?

That sounds right to me.

> [@perdita](#):
>
> ❓ Need to strip out Barcodes to a metadata file for QIIME2 import?

The metadata file can be used to map **barcodes to sample-ids** whereas the _manifest_ file maps **filepaths to sample IDs**.

Hope that helps. Let me know if you still need a hand going forward. Feel free to share the data you're working with (here or in a DM).

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [December 22, 2020, 5:46am UTC](https://forum.qiime2.org/t/using-qiime2-to-view-454-pyrosequencing-data/17500/5 "2020-12-22T05:46:29Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
