phylogenetic
427 papers tagged with this keyword
Hyperbolic distance matrix completion
A completion theory for hyperbolic distance data is developed at the interface of matrix analysis, graph theory, and hyperbolic geometry. Krein's characterization of the metric space embeddability in Lobachevsky space leads to a natural anchoring procedure that transforms the indefinite data into a positive semidefinite kernel. In analogy with positive semidefinite and Euclidean distance matrix completion, chordality of the specification graph is shown to be the necessary and sufficient condition for local Lorentz-Gram data to admit global completion. Existence is complemented by explicit constructions. For trees, we obtain geodesic-rectification and product-distance completions; for chordal graphs, the latter extends to matrix-valued transfers along clique-trees. The resulting canonical completion is characterized by sparsity of its inverse and by a maximum-absolute-determinant principle. Its metric distortion exhibits a sharp dichotomy governed by clique separator size. Applications to exact recovery from sparse hyperbolic measurements and to hierarchical and phylogenetic data are developed.
A Short Combinatorial Proof of the Pons-Batle Identity for Counting Tree-Child Networks
Tree-child networks are a useful class of binary phylogenetic networks. The Pons--Batle identity (Pons and Batle, \textit{Scientific Reports}, 2021) states that the number $a_{n,k}$ of tree-child networks with $k$ reticulations on $n$ taxa satisfies \[ a_{n,k}=(n-k+1)a_{n,k-1} +\frac{n(2n+k-3)}{n-k}a_{n-1,k}. \] In this paper, we present a short combinatorial proof of this identity.
An Affine Semigroup from Orbifold Boundary Conditions: cut, phylogenetic and hierarchical models in the unit-weight sector, and weighted configurations beyond them
The equivalence classes of boundary conditions of a gauge theory on a two-dimensional orbifold are the fibres of a marginal map, indexed by an affine semigroup: one generator per alphabet label, graded by weight, embedded by its local data at the fixed points. This note identifies that semigroup. Without weights the configuration has a name and a literature, whose results about our cases are attributed here: over $\mathbb{Z}_2$ it is the cut configuration of an explicit graph in the sense of Sturmfels-Sullivant --- the four-cycle for $T^2/\mathbb{Z}_2$, the wheel $W_4$ for $S^1/\mathbb{Z}_2\times S^1/\mathbb{Z}_2$ --- verified as an equality of configurations; over $\mathbb{Z}_m$ with equal cone orders, the group-based phylogenetic model on a claw tree; with unequal orders, a mixed-order variant we do not find in the literature; for higher products, the binary hierarchical model of a cross-polytope boundary complex. The product orbifold's ring is a row of a 2008 table --- codimension, degree, minimal generators, normality --- every invariant of which our machinery reproduced without knowing of it. What none of the three covers is the alphabet with weights, which arise from induction to higher-dimensional irreducibles of a non-abelian space group and from conjugate-pair recombination over real or quaternionic ground. That sector is adjacent to, but not identified with, the non-abelian direction Sturmfels and Sullivant raised in 2005, and is where our contributions sit: gluing trees for the weighted alphabets and the orthogonal and symplectic columns, and the group-based model on the tripod, a complete intersection exactly when the finite abelian group has order at most three. The first group beyond $\mathbb{Z}_3$ separates local from global: the $\mathbb{Z}_4$ tripod is a complete intersection on the Zariski-open set the phylogenetics literature works in, and not globally.
Tree Buckets and the Reconstruction of Pairs of Phylogenetic Trees
Phylogenetic trees are used in evolutionary biology to represent the evolutionary history of a collection of taxa. As we have incomplete information about any evolutionary history, recovering trees from partial information is a focus of phylogenetic combinatorics. However, in some cases the available data does not describe a single phylogenetic tree. We consider recovery of pairs of phylogenetic trees from their combined subtrees with $k$ leaves, which we call a $k$-bucket. We establish the exact cases in which these pairs of trees are recoverable from their subtrees with a single leaf removed, both when just considering the structure of the trees, and when additionally considering the set of taxa on the leaves. We also consider recovery of pairs of trees with labelled leaves from their rooted triples, and establish that they are recoverable up to a sequence of subtree swaps.
On the maximum size of 2-weakly compatible split systems
We consider a Turán-type problem arising in phylogenetics: determining the maximum size of a 2-weakly compatible split system. This compatibility condition arises in the reconstruction of phylogenetic networks from quartet weights. It was previously shown that a 2-weakly compatible split system has size at most
\[
3\binom{n}{4}+\binom{n}{2}.
\]
We prove that the maximum size is $O(n^{5/2})$.
A cube-root phase transition in tree-child networks and the enumeration threshold for galled networks
We prove two surprising results about phylogenetic networks. First, we show that the structure of tree-child networks with $n$ leaves and $k$ reticulation nodes undergoes a sharp phase transition at $n^{1/3}$: if $k=o(n^{1/3})$, then a random tree-child network is almost surely a semi-simplex tree-child network, whereas if $k/n^{1/3}\rightarrow\infty$ and $k=o(n^{1/2})$, it is almost surely not. Second, we show that this result implies that the asymptotic counting formula for galled networks with $n$ leaves and a fixed number $k$ of reticulation nodes remains valid in the range $k=o(n^{1/3})$, but not beyond. This is in strong contrast to recently established results for the asymptotic counting formulas for tree-child and normal networks with $n$ leaves and $k$ reticulation nodes, which are valid in the (optimal) range $k=o(n^{1/2})$.
How many cherry-picking sequences are needed to reduce all subtrees of a phylogenetic tree?
Phylogenetic networks are graphs that represent the evolutionary history of species. Recently, the class of orchard phylogenetic networks, which can be reduced by so-called cherry-picking sequences, has gained attention for its computational and biological aspects. In this paper, we study a fundamental question on orchards and their cherry-picking sequences by considering the CoveringNumber problem: given an orchard network $N$, how many cherry-picking sequences are needed to reduce all subnetworks of $N$? We initiate this study by considering the problem for trees. We then show that the covering number can be computed for binary trees recursively using a similar but more fine-grained notion of survival covering number. We also give a recursive formula for the survival covering number of non-binary trees. However, computing the covering number for non-binary trees appears to be considerably more challenging. For this case, we show that the covering number of star trees (whose root is adjacent to all leaves) is equivalent to the so-called SubsetConnectivity problem, which we introduce in this paper. Finally, we show that if there is no restriction on the sequence length, a single sequence of minimum length $\binom{n}{2}$ suffices to reduce all subtrees of a tree on $n$ leaves.
Metropolis-Hastings Sampling of Phylogenetic Networks: Correcting for Symmetries
In phylogenetics, Metropolis-Hastings methods are commonly used to sample phylogenetic trees or networks, for example from Bayesian posteriors. These methods generally use transitions that distinguish all nodes involved, and thus require fully labelled representations of phylogenetic networks. We argue that sampling leaf-labelled phylogenetic networks demands a correction for the number of fully labelled representatives of a leaf-labelled network, or, equivalently, for its internal symmetry. Without correction, there is a danger of undersampling networks with internal symmetries. We show that this correction can be realized by a quotient construction on the Metropolis-Hastings Markov chain, which, in practice, requires the calculation of the size of the network's automorphism group. Using $μ$-vectors, we show that the automorphism group is trivial for orchard networks, and thus also for tree-child networks and trees. This implies that a correction for symmetry is not needed when sampling only from such network classes. More generally, using our Python implementation of the algorithms in this paper, we show that using $μ$-vectors can significantly speed up calculations of automorphism group sizes and thus of Metropolis-Hastings sampling of leaf-labelled networks.
The maximum quartet distance between phylogenetic trees
The quartet distance counts the four-leaf subsets on which two binary phylogenetic trees display different topologies. We prove that its maximum over trees on $n$ leaves is $(2/3+o(1))\binom{n}{4}$, resolving a conjecture of Bandelt and Dress from 1986 and calibrating the scale of a fundamental metric for comparing phylogenetic trees. The proof reduces arbitrary pairs of trees to caterpillars by means of a common-root planarization and an identity on five-leaf trees.
Enumerating monophyletic characters in mathematical phylogenetics
Grouping species according to their phylogenetic relationships often results in different groups than grouping them according to their shared traits. Monophyletic groups play an important role in this regard, as they are groups of species sharing the same trait and being uniquely defined by a joint phylogenetic subtree. This immediately leads to the question of how to identify possible monophyletic groups in characters, which assign each present-day species a certain trait and which are typically used for phylogenetic tree reconstruction.
In our manuscript, we provide a general formula to quantify how many different characters are monophyletic on any given tree and provide simple formulae for binary characters and for certain tree shapes. We also investigate relations between monophyly and the well-known phylogenetic tree reconstruction criterion maximum parsimony by providing a linear-time algorithm which determines the parsimony score together with the monophyly type of a character on a tree.
Classes of phylogenetic networks that are robust to root placement
Standard phylogenetic reconstruction techniques often yield unrooted phylogenetic networks; these are subsequently rooted to infer evolutionary history. A common problem in this process is to determine the structural classes to which the resulting network will belong. In this paper, we investigate unrooted networks in which the choice of any root results in a valid rooted phylogenetic network, a property we define as {\em robustly orientable}. We then establish a strict structural condition for this class, specifically, that an unrooted network is robustly orientable if and only if it contains no sink components. We also show that if an unrooted network is level-$2$ or less, or if it is tree-based, then it is robustly orientable. Furthermore, we define an unrooted network to be {\em robustly class $\mathcal C$} if the choice of any root results in a network belonging to class $\mathcal C$. We demonstrate that an unrooted network is robustly tree-child or robustly stack-free if and only if it is level-$1$ or less. Finally, we show that a phylogenetic network is robustly normal if and only if it is a phylogenetic tree.
On Agreement Subtrees in Multiple Pylogenetic Trees
Snir and Yuster [Discrete Appl. Math. 347 (2026) 160--171] asked for the least number $h(k)$ such that $k$ unrooted binary phylogenetic trees on the same $h(k)$ leaves always share a common quartet. We give a new upper bound for the $k$-tree version of the Maximum Agreement Subtree problem, namely an upper bound for the number of leaves, on which $k$ unrooted binary phylogenetic trees always share a common induced binary subtree on $n$ leaves, which is a four-times iterated exponential function. For $h(k)$, this implies a four-times iterated exponential upper bound. We also set an exponential lower bound for $h(k)$.
Proximity Measures for Classes of Phylogenetic Networks
Phylogenetic networks are used to represent the evolutionary history of species. Due to biological interpretations and computational advantages, researchers have focused on restricted classes of phylogenetic networks, such as tree-child, orchard, and tree-based. These classes capture different notions of tree-likeness: tree-child networks require every internal vertex to have a taxon reachable by a tree path, orchard networks are trees with horizontal arcs (for modelling histories rife with horizontal gene transfers), and tree-based networks are trees with additional (not-necessarily horizontal) arcs. A natural question to ask is ``how far is a given network from belonging to a particular class?'' This motivates the study of proximity measures, which measure the minimum number of graph modifications required to transform a network into one belonging to a particular class. In this paper, we consider three proximity measures based on leaf addition, valid arc deletion, and arc deletion. We study pairwise comparability of the proximity measures, prove complexity results, and derive extremal bounds for the classes of tree, tree-child, orchard, and tree-based networks.
Polynomial encoding of rooted trees with branch lengths
Phylogenetic trees are rooted trees with branch lengths that record genetic divergence or elapsed time, and quantifying differences between them is central to a wide range of evolutionary and epidemiological analyses. Graph-polynomial encodings of rooted trees provide an accurate, interpretable, and computationally efficient way to compare tree shapes, but existing polynomial encodings must be paired with auxiliary structures to study rooted trees with branch lengths. We introduce a bivariate polynomial encoding that incorporates branch lengths directly into a recursive computation from the leaf vertices to the root vertex of a tree. We prove that, for rooted trees with branch lengths and no vertices of degree two, which include all standard phylogenetic trees, two trees have the same polynomial if and only if their underlying unlabeled trees are isomorphic and the branch lengths of corresponding edges are equal. We apply the polynomial encoding to three published HIV-1 phylogenies sampled in different epidemiological settings and show that it accurately separates the three datasets based on their tree topologies and branch lengths, outperforming previous polynomial-based approaches for analyzing rooted trees with branch lengths.
Exact Enumeration of Phylogenetic Networks: The Tree-Child, Reticulation-Visible and Orchard Hierarchy
We develop a unified framework for the exact enumeration and asymptotic analysis of the three most studied classes of phylogenetic networks: tree-child (TC), reticulation-visible (RV) and orchard networks, whose cardinalities satisfy the strict ordering $|\mathrm{TC}_{\ell,k}|<|\mathrm{RV}_{\ell,k}|<|\mathrm{Orch}_{\ell,k}|$ for reticulation number $k\geq2$ (with $\mathrm{TC}\subsetneq\mathrm{RV}$ and $\mathrm{TC}\subsetneq\mathrm{Orch}$, while $\mathrm{RV}$ and $\mathrm{Orch}$ are incomparable as sets). Using the Chang--Fuchs structural theorem, we derive a two-level master functional equation for the RV bivariate generating function and obtain exact closed-form identities for the differences $Δ_k(\ell):=|RV_{\ell,k}|-|TC_{\ell,k}|$ for $k=2,3$, with the asymptotic universality $Δ_k(\ell)/|TC_{\ell,k}|\sim k!/\ell$. For orchard networks, we prove a \emph{universal hypergeometric law} that resolves the exact enumeration problem for all $\ell$: the column generating function $F_\ell(v)$ is rational with denominator $D_\ell(v)=\prod_{j=2}^\ell X_j(v)$, where \[
X_\ell(v) = \sum_{k=0}^{\lfloor\ell/2\rfloor}(-1)^k\,
\frac{\ell!}{(\ell-2k)!\,k!}\,v^k \] is the matching polynomial of the complete graph $K_\ell$ and a rescaled Jacobi polynomial. This immediately resolves the intractable $\ell=9$ case: $D_9$ has degree 20, dominant growth rate $\approx40.73$, and all spectral roots are positive real. A complete enumeration table is provided extending the published data of Cardona, Ribas and Pons.
A parameterized family of balance indices for phylogenetic networks
We introduce a new family of balance indices for phylogenetic networks: the $H_α$ indices, where $α$ is a positive real number. This family includes the $B_2$ index as a special case ($α= 1$) and provides a natural extension of the Sackin index to phylogenetic networks. We show that the $H_α$ indices share many structural properties with the $B_2$ index, most notably a "grafting property" that makes it possible to express the $H_α$ index of a network in terms of the $H_α$ indices of its biconnected components. These properties allow us to identify networks that minimize / maximize $H_α$ for various classes of phylogenetic networks, and to study its distribution for several models of random trees and networks (in particular, Galton-Watson trees and binary Markov branching trees, with a focus on the Yule and PDA models). Finally, we show how local limits can be used to analyze the asymptotic behavior of $H_α$ for large trees and networks, and we obtain general results for the moments of $H_α$ for a broad class of random phylogenetic networks known as blowups of Galton-Watson trees.
Encoding Phylogenetic Networks with Least Common Ancestor Constraints
Encoding phylogenetic networks by suitable substructures is a central problem in phylogenetic combinatorics. We study encodings based on least common ancestor (LCA) constraints. For a directed acyclic graph (DAG) $G$ with leaf set $X$, we consider the relation on pairs of leaves in which $(ab,xy)$ records that the LCAs of $a,b$ and $x,y$ are well-defined and that the former is a descendant of the latter.
We first identify precisely which part of $G$ is determined by this relation. To this end, we compare the canonical DAG constructed from the LCA relation with the 2-regularization of $G$, obtained by removing all vertices that are not LCAs of one or two leaves and then deleting shortcut edges. We prove that these two DAGs are isomorphic. Hence the obstruction to encoding a graph by its LCA relation is exactly the information lost under 2-regularization.
This yields a general reconstruction principle, which we apply to several natural classes of phylogenetic networks. In particular, we show that shortcut-free 2-LCA-relevant DAGs, phylogenetic trees, regular level-1 networks, regular networks with binary clustering systems, regular networks whose clustering systems are closed weak hierarchies, strong-phylogenetic normal networks, separated phylogenetic normal networks, and binary normal networks are encoded by their LCA relations.
We also introduce a sparse triple-like restriction consisting only of comparisons of the form $(ab,ac)$, where $a,b,c\in X$ are pairwise distinct. For graphs with the 2-LCA property, we show that this sparse relation, together with the leaf set, determines the full LCA relation after a natural closure operation. Consequently, several of the above classes can be reconstructed, up to isomorphism, from the sparse relation in polynomial time.
Note on the Maximum Number of Trees Displayed by a Tree-Child Network
In this note, we show that, for all $n\ge 2$, the number of distinct rooted binary phylogenetic $X$-trees displayed by a binary tree-child network $\mathcal{N}$ on $X$ with $n$ leaves is at most $2^{n-1}-1$ and that this upper bound is sharp. Furthermore, if $\mathcal{N}$ displays exactly $2^{n-1}-1$ such trees, then exactly one rooted binary phylogenetic $X$-tree is displayed twice, and this tree can be canonically found by iteratively replacing a reticulated cherry with a cherry.
An Explicit $O(r\log r)$ Threshold for Attaining the Semple--Steel Bound with $r$-State Characters
Let $d_r(n)$ be the maximum, over all binary phylogenetic trees with $n$ leaves, of the minimum number of $r$-state characters required to define the tree. Semple and Steel proved that $d_r(n)\geq\lceil(n-3)/(r-1)\rceil$, and Bordewich and Semple proved that equality holds for each fixed $r$ and all sufficiently large $n$. We study the corresponding threshold $n_r$, the least $N$ for which equality holds for every $n\geq N$. The Bordewich--Semple construction yields an explicit polynomial upper bound of order $O(r^5)$ for this threshold. We prove the near-linear estimate \[
3r+1\leq n_r\leq \ceil{64(r-1)\log_2(r+1)}+3\qquad(r\geq4). \] The proof constructs, for every binary phylogenetic tree with $m=n-3$ internal edges, a linked quartet certificate whose conflict graph has maximum degree at most $16\lceil\log_2(m+2)\rceil+4$. Equitable coloring then packs the certificate into exactly $\lceil m/(r-1)\rceil$ $r$-state characters once $m\geq\lceil64(r-1)\log_2(r+1)\rceil$. We also include the lower bound $n_r\geq3r+1$, obtained from the snowflake obstruction, and state the natural conjecture that this lower bound is the exact threshold for all $r\geq4$. The conjectural endpoint is consistent with the known small-state thresholds: $n_4=13$ and $n_5=16$, while the cases $r=2,3$ are also explicitly classified.
Regularizing and Normalizing DAGs and Phylogenetic Networks
Phylogenetic networks and, more generally, directed acyclic graphs (DAGs) represent hierarchical structure beyond trees, for instance in the presence of reticulate evolutionary events such as hybridization or horizontal gene transfer. A central question is which parts of such graphs are essential with respect to leaf-observable information, and which parts can be removed without changing this information. Resolving this question can lead to principled simplification methods for phylogenetic networks, such as the recent normalization approach of Francis et al.
In this paper, we study this question from three related perspectives: clusters displayed by a DAG $G$, least common ancestors (LCAs) of subsets of its leaf set, and visibility, a path-based property of vertices. We first introduce an LCA-based simplification procedure called $i$-regularization. For a DAG $G$ and $i\geq 1$, the DAG $\reg_i(G)$ retains precisely those vertices that occur as unique LCAs of leaf subsets of size at most $i$, removes the remaining non-leaf vertices by a graph-editing operation $\ominus$, and then deletes shortcuts. We show that $\reg_i(G)$ preserves all such LCAs, is $i$-lca-relevant, and admits a cluster-level description: it is regular, i.e., isomorphic to the Hasse diagram of the corresponding lca-clusters.
We then compare LCA-based regularization with normalization. Using the same $\ominus$-operator, we describe the cover construction underlying normalization, identify visible vertices that are nevertheless removed, and characterize when regularization and normalization coincide. Together, these results provide a unified framework for cluster-based, LCA-based, and visibility-based simplifications of DAGs and phylogenetic networks.