# Help with understanding beta diversity calculation + results

**URL:** https://forum.qiime2.org/t/help-with-understanding-beta-diversity-calculation-results/17273
**Category:** User Support
**Tags:** best-of-the-forum
**Created:** [October 30, 2020, 1:38pm UTC](https://forum.qiime2.org/t/help-with-understanding-beta-diversity-calculation-results/17273 "2020-10-30T13:38:52Z")
**Posts on this page:** 1
**Showing post:** 3

<div class="post-metadata">

### Author: ![ChrisKeefe](https://forum.qiime2.org/user_avatar/forum.qiime2.org/chriskeefe/32/1907_2.png) [@ChrisKeefe](https://forum.qiime2.org/u/ChrisKeefe)
#### Post date: [October 30, 2020, 6:14pm UTC](https://forum.qiime2.org/t/help-with-understanding-beta-diversity-calculation-results/17273/3 "2020-10-30T18:14:13Z")

</div>

Lots of good questions, @fgara. I'll tackle the ones I'm confident in, and try to get you resources for the others.

> [@fgara](#):
>
> These beta diversity values are calculated using PCoA via scikit-bio (e.g. [http://scikit-bio.org/docs/0.5.0/diversity.html](http://scikit-bio.org/docs/0.5.0/diversity.html)), am I right?

You've got the order of operations a little mixed up, but you're headed in the right direction. When making a PCoA plot, distance matrices (distances between samples) are calculated first, and then PCoA results are produced from the distance matrices.

The distance matrices are built using [skbio](http://scikit-bio.org/docs/0.5.0/generated/skbio.diversity.beta_diversity.html#skbio.diversity.beta_diversity) or [unifrac](https://github.com/biocore/unifrac). Most of the `skbio` calculations are actually passed off to `sklearn` or `scipy` (details [here](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.pairwise_distances.html)), but some are [implemented in `q2-diversity-lib`](https://github.com/ChrisKeefe/q2-diversity-lib/blob/6357399c95a2fdd0dd9d9d95ba88f3b366fe8082/q2_diversity_lib/beta.py#L58).

> [@fgara](#):
>
> When I use PCoA for my data set that has hundreds of different taxa, how can I know which taxa contribute most to the PCoA axes?

> [@fgara](#):
>
> What does qiime diversity pcoa-biplot do?

> [@fgara](#):
>
> When I use PCoA for my data set that has hundreds of different taxa, how can I know which taxa contribute most to the PCoA axes?

Biplots are a great way of doing exactly this!  
I made this example with `diversity pcoa_biplot`, and then visualized it with `emperor biplot`.

 ![image](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/1/128f313ee668093d6bc24c9a9608727ba070a5c4.jpeg)

Each arrow describes one feature's contribution to the difference described by the PCoA plot. The red arrow (feature `b3c...965a`) is the biggest; it contributes the most. The purple arrow (`82b...cda`) is the smallest _of the five most-prominent features_; it contributes _the fifth-most_. You can adjust how many arrows are plotted with [`emperor biplot`](https://docs.qiime2.org/2020.8/plugins/available/emperor/biplot/).

If you want to know what features these are, use a classifier to create a `FeatureData[Taxonomy]` artifact, and visualize it with `qiime metadata tabulate`. You'll get a nice list where you can search the feature ids you see at the ends of the arrows.

 ![image](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/8/844e25395d9f1ff45190427bacd6f9253a866f9d.png)

> [@fgara](#):
>
> Is the first principal coordinate axis the line defined by the first eigenvector?

IIRC, the first principal coordinate axis is the vector with the largest eigenvalue - the most "important" vector, if you will. The percentage of variance each axis explains is displayed in parentheses at the end of that axis. Again, IIRC, this is the axis eigenvalue, divided by the sum of all axis eigenvalues. The "axes" tab of the visualization also has a nice little plot explaining percent variance explained by the top 5 axes.

 ![image](https://forum-qiime2-org.s3.dualstack.us-west-2.amazonaws.com/original/2X/e/ea99020de7bf773d9d212aa412d6186349d8b20f.png)

> [@fgara](#):
>
> I read somewhere on this forum, that if there is same distance between 2 points in axis1 and there is same distance between 2 points in axis2, the actual difference is greater between those points in axis1, because axis1 is more important that axis2? Could anyone help explaining this further please?

I think the idea here is that how important one unit of difference is should be scaled by how important the axis is. If PC52 only explains .00001% of the variance in your samples, it doesn't matter much how far apart two points are on that axis. If, on the other hand, PC1 explains 45% of the total variance, differences between samples on PC1 contribute much more relative to the overall variance.

> [@fgara](#):
>
> - I watched a Youtube video that says that PCoA is similar to PCA, in the fact that they both have loadings/loading scores… what are these loading scores exactly?
> - Where can I find the PCoA loading scores in QIIME2?

Unfortunately, @fgara, this is where you lose me. This [ResearchGate post](https://www.researchgate.net/post/Does_Principal_Coordinates_Analysis_PCoA_generate_loadings) has a couple resources that might help you with how to interpret loadings in PCoA, but it sounds like they may not carry the same meaning as PCA loadings. I'm not sure they're available at all in QIIME 2 artifacts, but if they were, you'd probably have to extract them from a `PCoAResults` or `PCoAResults % Properties('biplot')` Artifact. Maybe someone with more experience with the nuts and bolts of PCoA will weigh in on this. 🤞

Best,  
Chris 🦈

---

_[View the full topic](https://forum.qiime2.org/t/help-with-understanding-beta-diversity-calculation-results/17273)._
