Free tools. No account. Nothing stored.

Phylogenetic Distance and Tree Calculator

Most phylogenetic distance calculators hand you the percentage of sites that differ and call it a distance. That number is not comparable between a closely related pair and a distant one. It cannot tell you that a 20% difference over 50 sites means something quite different from a 20% difference over 5000, and it silently assumes every site changes at most once. Kimura’s 2-parameter model fixes both of those, and the ratio of transitions to transversions is usually the first thing worth knowing about any pair of sequences.

Most phylogenetic distance calculators hand you a percentage of differences and call it a distance. That number is not comparable between a closely related pair and a distant one, because sites that differ twice get counted once. Kimura’s 2-parameter model corrects for that, and the ratio of transitions to transversions is usually the first thing worth knowing about a pair of sequences.

p-distance and why it is not enough

p-distance is the proportion of aligned positions that differ. It is easy, it is in every textbook, and it is what most online tools stop at.

The problem is that it treats a site that changed once and a site that changed four times as identical evidence. Once a sequence has been changing for long enough, some sites will have hit the same base again by chance, and each of those is silently undercounted. The effect is saturation, and it is worst exactly where phylogeneticists most want resolution: between distant taxa.

The second problem is that p-distance ignores which bases changed. Genetic code tells us substitutions are not equally likely, and that is not a small effect.

Transitions and transversions

A transition is a change within a chemical class: purine to purine (A↔G) or pyrimidine to pyrimidine (C↔T). A transversion is a change between classes, and there are four possible ones against only two transitions.

Real DNA substitution patterns are not the 1:2 ratio the raw counts suggest, because transitions are generally accepted and repaired more readily than transversions, especially in coding regions where a wrong base is more likely to damage a protein. The Ts/Tv ratio summarises that. In real coding sequence it usually lands somewhere around 2 to 3. A ratio far below 1 in supposedly orthologous coding sequence is a red flag, and it usually means the two sequences are not what you think they are, or the alignment is wrong.

Be careful with the ratio on short sequences. With only a handful of differences, one extra transversion swings it wildly, and a ratio of exactly zero just means no transitions happened, which is not evidence of anything.

Kimura 2-parameter

K80 assumes transitions and transversions each occur at their own constant rate, and corrects the observed distance for two things p-distance ignores: the non-random ratio between the two types, and the chance of multiple hits at the same site.

K80 = −0.5 × ln(1 − 2P − Q) + 0.25 × ln(1 − 2Q)

where P is the transition proportion and Q the transversion proportion. Because of the multiple-hit correction, K80 comes out below p-distance for divergent pairs — some observed differences are the same event counted twice, so the true number of changes is fewer than the raw count.

One practical note: K80 is a maximum-likelihood estimator and can return a very slightly negative value for near-identical sequences, which is not a meaningful distance. This tool clamps it at zero rather than reporting a negative distance or building a negative branch height out of it.

K80 has a limit too. When 1 − 2P − Q goes to zero or below, the correction breaks down because the sequences are so divergent that the model no longer describes them. That is saturation, and this tool flags it, because a saturated pair should be dropped or handled with a different model rather than reported as a large distance.

Alignment

You cannot compute a distance without an alignment, and unequal-length sequences get gapped. This tool does a global Needleman–Wunsch alignment with a simple match, mismatch and gap scheme, which is the right tool for the short sequences a phylogeny exercise hands you.

Be honest about what that is. Real alignment is a much harder problem with its own error bars, and if the sequences you are comparing are divergent enough to need gaps, the alignment choice will affect the distances more than the model will. A global alignment is also the wrong tool for sequences that genuinely differ in length by a lot, where a local or progressive method would be more appropriate.

Reading the tree

The tree is built by UPGMA, Unweighted Pair Group Method using Arithmetic averages. It repeatedly joins the two most similar clusters, and the branch height is half the distance at which they joined.

Two things to hold onto. UPGMA assumes a molecular clock, that all lineages accumulate changes at the same constant rate. When that is false the tree gets distorted, and branch lengths become evolutionary distances rather than actual time. And UPGMA is a distance method, so it cannot use the information in the shape of a sequence, only the differences — which is why maximum-likelihood and parsimony methods can and do disagree with it.

The outgroup, the single sequence that ends up furthest from everything, is the one you most need to have chosen on biological grounds rather than on the data. Rotate the tree around any branch and the relationships do not change, so do not read meaning into which side something is drawn on.

The Newick string is there because a tree has to go into a report somehow, and Newick is what tree-viewing software reads.

Worked example

Four 17-base sequences, one substitution apart for human and chimp, two for chimp and mouse.

Human and chimp come out as the closest pair, and every other distance is larger. UPGMA joins them first, then joins that cluster to mouse, and yeast is left as the outgroup. That is the expected topology for those relationships, and the tool reproduces it from the sequences alone.

Switch the model to K80 and the closest pair’s distance drops to zero, because with a single transversion in 17 sites the corrected estimate is not meaningfully different from identical. The other pairs stay clearly separated. This is the behaviour you want: near-identical sequences should not be assigned a spurious positive distance.

What this tool cannot do is establish which tree is right. With four short sequences you have almost no information, and a large K80 value in particular is a signal that the model is being pushed past what the data can support.

Frequently Asked Questions

What is the difference between p-distance and Kimura 2-parameter?

p-distance counts every differing site once and ignores whether a site changed once or several times. K80 corrects for that multiple-hit saturation and for transitions and transversions not being equally likely, so its numbers are not comparable across pairs of different divergence.

What is the Ts/Tv ratio and what should it be?

Transitions are purine to purine or pyrimidine to pyrimidine changes; transversions go between the classes. The ratio of the two summarises the bias, and real coding sequence usually gives something around 2 to 3. Far below 1 suggests a bad alignment or sequences that are not what you think.

Why can a distance come out as zero or negative?

K80 is a maximum-likelihood estimator and can go slightly negative for near-identical sequences, which is not a meaningful distance. This tool clamps it at zero instead of reporting a negative distance or building a negative branch height out of it.

What does saturation mean?

It means a pair of sequences is so divergent that some sites have changed more than once, and p-distance and K80 both undercount. Once the model is saturated its correction breaks down, and the pair should be dropped or handled with a different model rather than reported as a large distance.

What does the UPGMA tree assume?

A molecular clock, meaning all lineages change at the same constant rate. When that is false the tree is distorted. UPGMA is also a distance method, so it only uses differences and cannot use the information in the shape of a sequence.

Does a UPGMA tree prove which relationships are correct?

No. With a handful of short sequences you have very little information, and maximum likelihood or parsimony can and do disagree. The outgroup also has to be chosen on biological grounds, not because the data made it look isolated.

How reliable is the alignment behind these distances?

Less reliable than the distance model, if the sequences are divergent. A global alignment is the right tool for the short sequences a course exercise gives you, but real alignment is harder and its choice can move the distances more than the model does.

Related tools