# Questions about PCoA biplots

**URL:** https://forum.qiime2.org/t/questions-about-pcoa-biplots/18152
**Category:** User Support
**Created:** [January 20, 2021, 7:56am UTC](https://forum.qiime2.org/t/questions-about-pcoa-biplots/18152 "2021-01-20T07:56:34Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![sbslee](https://forum.qiime2.org/user_avatar/forum.qiime2.org/sbslee/32/9039_2.png) [@sbslee](https://forum.qiime2.org/u/sbslee)
#### Post date: [January 20, 2021, 7:56am UTC](https://forum.qiime2.org/t/questions-about-pcoa-biplots/18152/1 "2021-01-20T07:56:35Z")

</div>

Hello everyone,

I have been creating PCoA biplots, and they have been very useful for identifying features that contribute the most in terms of separating the samples in a PCoA plot. Shown below is an example of PCoA biplot created with QIIME 2 and the `dokdo` [package](https://github.com/sbslee/dokdo/wiki/Dokdo-API#addbiplot-) I wrote:

 ![addbiplot-1-20201217](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/a/a5fbdc50664898128375b2de08559b0b919df9fb.png)

I think I understand enough how to interpret a PCoA biplot. That is, data points represent samples and arrows represent features with their length indicating the amount of loading (importance) with respect to the given axes.

I also know how to generate a PCoA biplot in QIIME 2. This [post](https://forum.qiime2.org/t/how-to-use-pcoa-biplot/7953/2) gives a good overview, but basically you need to use the `pcoa-biplot` [method](https://docs.qiime2.org/2020.11/plugins/available/diversity/pcoa-biplot/?highlight=pcoa%20biplot), which accepts the Artifact types `PCoAResults` and `FeatureTable[RelativeFrequency]` as input.

However, I have some questions about the way a PCoA biplot is generated in QIIME 2 that I could not find answers to in this forum.

Q1. Why does `pcoa-biplot` require `FeatureTable[RelativeFrequency]` instead of `FeatureTable[Frequency]`?

One might ask why we need to provide a feature table to begin with, but I think I got this part: Unlike PCA, PCoA is based on a distance matrix or dissimilarity between the samples, so all the feature data are lost in the way and that's why we need to separately provide a feature table. But why use relative frequency? I did find this in the doc:

> Project features into a principal coordinates matrix. The features used should be the features used to compute the distance matrix. It is recommended that these variables be normalized in cases of dimensionally heterogeneous physical variables.

However, I still couldn't understand why it's relevant. Any clarification (e.g. "dimensionally heterogeneous physical variables") would be appreciated.

Q2. How does `pcoa-biplot` correctly project features into an existing PCoA space?

This may be obvious to many, but I just can't wrap my head around this. By creating a `PCoAResults` artifact you make a N-dimension space with defined axes. Your samples are projected into this space. So far so good. But suddenly, `pcoa-biplot` projects the features (i.e. relative frequency) into the same space using the same axes and then draws arrows to the features from the origin. How does `pcoa-biplot` correctly orient (?) itself when projecting the features into the space of samples and why does it work?! Any guidance regarding this would be deeply appreciated 🙂

Q3. Does `pcoa-biplot` compute a distance matrix between the features before they are projected into a PCoA space?

Q4. What kind of ordination method does `pcoa-biplot` use for projecting the features into a PCoA space?

Please let me know if any of my questions are unclear. Appreciate your help!

---

<div class="post-metadata">

### Author: ![yoshiki](https://forum.qiime2.org/user_avatar/forum.qiime2.org/yoshiki/32/16_2.png) [@yoshiki](https://forum.qiime2.org/u/yoshiki)
#### Post date: [January 26, 2021, 11:06pm UTC](https://forum.qiime2.org/t/questions-about-pcoa-biplots/18152/3 "2021-01-26T23:06:52Z")

</div>

> Q1. Why does `pcoa-biplot` require `FeatureTable[RelativeFrequency]` instead of `FeatureTable[Frequency]` ?

The answer to this question lies on the suggestion that Legendre and Legendre give in their Numerical Ecology book. If you are interested the underlying functionality is implemented [here](http://scikit-bio.org/docs/0.5.6/generated/skbio.stats.ordination.pcoa_biplot.html#skbio.stats.ordination.pcoa_biplot).

> Q2. How does `pcoa-biplot` correctly project features into an existing PCoA space?

The projection step is done by means of computing a covariance matrix between the PCoA matrix (samples by principal coordinates), and the feature table (samples by features).

> Q3. Does `pcoa-biplot` compute a distance matrix between the features before they are projected into a PCoA space?

It does not.

> Q4. What kind of ordination method does `pcoa-biplot` use for projecting the features into a PCoA space?

There's no ordination method used as part of pcoa-biplot. The ordination is as provided by the user (usually PCoA). This method won't rank or re-arrange principal axes.

* * *

Hopefully these answers are helpful. In general, I would say we now have better ways to compute and estimate these biplots (for example DEICODE). `pcoa-biplot` is one of the only methods I found that would work when you want to see the relationship of features and a given distance matrix (for example UniFrac).

---

<div class="post-metadata">

### Author: ![system](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/3X/2/1/21af5fe23cb6f4579467c66a9ed94e55274ca7bd.svg) [@system](https://forum.qiime2.org/u/system)
#### Post date: [February 27, 2021, 5:06am UTC](https://forum.qiime2.org/t/questions-about-pcoa-biplots/18152/4 "2021-02-27T05:06:54Z")

</div>

This topic was automatically closed 31 days after the last reply. New replies are no longer allowed.
