species
265 papers tagged with this keyword
Bounding the softwired parsimony score of a phylogenetic network
In comparison to phylogenetic trees, phylogenetic networks are more suitable to represent complex evolutionary histories of species whose past includes reticulation such as hybridisation or lateral gene transfer. However, the reconstruction of phylogenetic networks remains challenging and computationally expensive due to their intricate structural properties. For example, the small parsimony problem that is solvable in polynomial time for phylogenetic trees, becomes NP-hard on phylogenetic networks under softwired and parental parsimony, even for a single binary character and structurally constrained networks. To calculate the parsimony score of a phylogenetic network $N$, these two parsimony notions consider different exponential-size sets of phylogenetic trees that can be extracted from $N$ and infer the minimum parsimony score over all trees in the set. In this paper, we ask: What is the maximum difference between the parsimony score of any phylogenetic tree that is contained in the set of considered trees and a phylogenetic tree whose parsimony score equates to the parsimony score of $N$? Given a gap-free sequence alignment of multi-state characters and a rooted binary level-$k$ phylogenetic network, we use the novel concept of an informative blob to show that this difference is bounded by $k+1$ times the softwired parsimony score of $N$. In particular, the difference is independent of the alignment length and the number of character states. We show that an analogous bound can be obtained for the softwired parsimony score of semi-directed networks, while under parental parsimony on the other hand, such a bound does not hold.
A Vector Representation for Phylogenetic Trees
Good representations for phylogenetic trees and networks are important for optimizing storage efficiency and implementation of scalable methods for the inference and analysis of evolutionary trees for genes, genomes and species. We introduce a new representation for rooted phylogenetic trees that encodes a binary tree on n taxa as a vector of length 2n in which each taxon appears exactly twice. Using this new tree representation, we introduce a novel tree rearrangement operator, called a HOP, that results in a tree space of diameter n and a quadratic neighbourhood size. We also introduce a novel metric, the HOP distance, which is the minimum number of HOPs to transform a tree into another tree. The HOP distance can be computed in near-linear time, a rare instance of a tree rearrangement distance that is tractable. Our experiments show that the HOP distance is better correlated to the Subtree-Prune-and-Regraft distance than the widely used Robinson-Foulds distance. We also describe how the novel tree representation we introduce can be further generalized to tree-child networks.
Operad structures on the species composition of two operads
We give an explicit description of three operad structures on the species composition $p \circ q$, where $q$ is any given positive operad, and where $p$ is the NAP operad, or a shuffle version of the magmatic operad Mag. No distributive law between $p$ and $q$ is assumed.
Phylogenetic diversity indices from an affine and projective viewpoint
Phylogenetic diversity indices are commonly used to rank the elements in a collection of species or populations for conservation purposes. The derivation of these indices is typically based on some quantitative description of the evolutionary history of the species in question, which is often given in terms of a phylogenetic tree. Both rooted and unrooted phylogenetic trees can be employed, and there are close connections between the indices that are derived in these two different ways. In this paper, we introduce more general phylogenetic diversity indices that can be derived from collections of subsets (clusters) and collections of bipartitions (splits) of the given set of species. Such indices could be useful, for example, in case there is some uncertainty in the topology of the tree being used to derive a phylogenetic diversity index. As well as characterizing some of the indices that we introduce in terms of their special properties, we provide a link between cluster-based and split-based phylogenetic diversity indices that uses a discrete analogue of the classical link between affine and projective geometry. This provides a unified framework for many of the various phylogenetic diversity indices used in the literature based on rooted and unrooted phylogenetic trees, generalizations and new proofs for previous results concerning tree-based indices, and a way to define some new phylogenetic diversity indices that naturally arise as affine or projective variants of each other.
Exact and Heuristic Computation of the Scanwidth of Directed Acyclic Graphs
To measure the tree-likeness of a directed acyclic graph (DAG), a new width parameter that considers the directions of the arcs was recently introduced: scanwidth. We present the first algorithm that efficiently computes the exact scanwidth of general DAGs. For DAGs with one root and scanwidth $k$ it runs in $O(k \cdot n^k \cdot m)$ time. The algorithm also functions as an FPT algorithm with complexity $O(2^{4 \ell - 1} \cdot \ell \cdot n + n^2)$ for phylogenetic networks of level-$\ell$, a type of DAG used to depict evolutionary relationships among species. Our algorithm performs well in practice, being able to compute the scanwidth of synthetic networks up to 30 reticulations and 100 leaves within 500 seconds. Furthermore, we propose a heuristic that obtains an average practical approximation ratio of 1.5 on these networks. While we prove that the scanwidth is bounded from below by the treewidth of the underlying undirected graph, experiments suggest that for networks the parameters are close in practice.
The inhomogeneous $t$-PushTASEP and Macdonald polynomials
We study a multispecies $t$-PushTASEP system on a finite ring of $n$ sites with site-dependent rates $x_1,\dots,x_n$. Let $λ=(λ_1,\dots,λ_n)$ be a partition whose parts represent the species of the $n$ particles on the ring. We show that for each composition $η$ obtained by permuting the parts of $λ$, the stationary probability of being in state $η$ is proportional to the ASEP polynomial $F_η(x_1,\dots,x_n; q,t)$ at $q=1$; the normalizing constant (or partition function) is the Macdonald polynomial $P_λ(x_1,\dots,x_n;q,t)$ at $q=1$. Our approach involves new relations between the families of ASEP polynomials and of non-symmetric Macdonald polynomials at $q=1$. We also use multiline diagrams, showing that a single jump of the PushTASEP system is closely related to the operation of moving from one line to the next in a multiline diagram. We derive symmetry properties for the system under permutation of its jump rates, as well as a formula for the current of a single-species system.
On the correctness of Maximum Parsimony for data with few substitutions in the NNI neighborhood of phylogenetic trees
Estimating phylogenetic trees, which depict the relationships between different species, from aligned sequence data (such as DNA, RNA, or proteins) is one of the main aims of evolutionary biology. However, tree reconstruction criteria like maximum parsimony do not necessarily lead to unique trees and in some cases even fail to recognize the \enquote{correct} tree (i.e., the tree on which the data was generated). On the other hand, a recent study has shown that for an alignment containing precisely those binary characters (sites) which require up to two substitutions on a given tree, this tree will be the unique maximum parsimony tree.
It is the aim of the present paper to generalize this recent result in the following sense: We show that for a tree $T$ with $n$ leaves, as long as $k<\frac{n}{8}+\frac{11}{9}-\frac{1}{18}\sqrt{9\cdot \left(\frac{n}{4}\right)^2+16}$ (or, equivalently, $n>9 k-11+\sqrt{9k^2-22 k+17} $, which in particular holds for all $n\geq 12k$), the maximum parsimony tree for the alignment containing all binary characters which require (up to or precisely) $k$ substitutions on $T$ will be unique in the NNI neighborhood of $T$ and it will coincide with $T$, too. In other words, within the NNI neighborhood of $T$, $T$ is the unique most parsimonious tree for the said alignment. This partially answers a recently published conjecture affirmatively. Additionally, we show that for $n\geq 8$ and for $k$ being in the order of $\frac{n}{2}$, there is always a pair of phylogenetic trees $T$ and $T'$ which are NNI neighbors, but for which the alignment of characters requiring precisely $k$ substitutions each on $T$ in total requires fewer substitutions on $T'$.
Identifying circular orders for blobs in phylogenetic networks
Interest in the inference of evolutionary networks relating species or populations has grown with the increasing recognition of the importance of hybridization, gene flow and admixture, and the availability of large-scale genomic data. However, what network features may be validly inferred from various data types under different models remains poorly understood. Previous work has largely focused on level-1 networks, in which reticulation events are well separated, and on a general network's tree of blobs, the tree obtained by contracting every blob to a node. An open question is the identifiability of the topology of a blob of unknown level. We consider the identifiability of the circular order in which subnetworks attach to a blob, first proving that this order is well-defined for outer-labeled planar blobs. For this class of blobs, we show that the circular order information from 4-taxon subnetworks identifies the full circular order of the blob. Similarly, the circular order from 3-taxon rooted subnetworks identifies the full circular order of a rooted blob. We then show that subnetwork circular information is identifiable from certain data types and evolutionary models. This provides a general positive result for high-level networks, on the identifiability of the ordering in which taxon blocks attach to blobs in outer-labeled planar networks. Finally, we give examples of blobs with different internal structures which cannot be distinguished under many models and data types.
Mallows Product Measure
Published in Electron. J. Probab. 29: 1-33 (2024)
• View Publication
• BIB
Q-exchangeable ergodic distributions on the infinite symmetric group were classified by Gnedin-Olshanski (2012). In this paper, we study a specific linear combination of the ergodic measures and call it the Mallows product measure. From a particle system perspective, the Mallows product measure is a reversible stationary blocking measure of the infinite-species ASEP and it is a natural multi-species extension of the Bernoulli product blocking measures of the one-species ASEP. Moreover, the Mallows product measure can be viewed as the universal product blocking measure of interacting particle systems coming from random walks on Hecke algebras.
For the random infinite permutation distributed according to the Mallows product measure we have computed the joint distribution of its neighboring displacements, as well as several other observables. The key feature of the obtained formulas is their remarkably simple product structure. We project these formulas to ASEP with finitely many species, which in particular recovers a recent result of Adams-Balazs-Jay, and also to ASEP(q,M).
Our main tools are results of Gnedin-Olshanski about ergodic Mallows measures and shift-invariance symmetries of the stochastic colored six vertex model discovered by Borodin-Gorin-Wheeler and Galashin.
Musical Systems with $\mathbb{Z}_n$ -- Cayley Graphs
We apply geometric group theory to study and interpret known concepts from Western music. We show that chords, the circle of fifths, scales and certain aspects of the first species of counterpoint are encoded in the Cayley graph of the group $\mathbb{Z}_{12}$, generated by $3$ and $4$. Using $\mathbb{Z}_{12}$ as a model, we extend the above music concepts to a particular class of groups $\mathbb{Z}_{n}$, which displays geometric and algebraic features similar to $\mathbb{Z}_{12}$. We identify a weaker form of counterpoint which, in particular leads to Fux's dichotomy in $\mathbb{Z}_{12}$, and to consonant sets in $\mathbb{Z}_n$. Using Maple software, we implement these new constructions and show how to experiment with them musically.
Operad Structure of Poset Matrices
This paper examines operad structures derived from poset matrices by formulating a set of new construction rules for poset matrices. In this direction, eleven different partial composition operations will be introduced as the basis for the construction of poset matrices of any given size by extending the combinatorial setting of species of structures to poset matrices. Three of these partial composition operations are shown to define an operad structure for poset matrices. The structural properties of poset matrices and their duals are then studied based on their associated operad constructions.
The doubly asymmetric simple exclusion process, the colored Boolean process, and the restricted random growth model
The multispecies asymmetric simple exclusion process (mASEP) is a Markov chain in which particles of different species hop along a one-dimensional lattice. This paper studies the doubly asymmetric simple exclusion process $\mathrm{DASEP}(n,p,q)$ in which $q$ particles with species $1, \dots, p$ hop along a circular lattice with $n$ sites, but also the particles are allowed to spontaneously change from one species to another. In this paper, we introduce two related Markov chains called the colored Boolean process and the restricted random growth model, and we show that the DASEP lumps to the colored Boolean process, and the colored Boolean process lumps to the restricted random growth model. This allows us to generalize a theorem of David Ash on the relations between sums of steady state probabilities. We also give explicit formulas for the stationary distribution of $\mathrm{DASEP}(n,2,2)$.
Predicting Horizontal Gene Transfers with Perfect Transfer Networks
Horizontal gene transfer inference approaches are usually based on gene sequences: parametric methods search for patterns that deviate from a particular genomic signature, while phylogenetic methods use sequences to reconstruct the gene and species trees. However, it is well-known that sequences have difficulty identifying ancient transfers since mutations have enough time to erase all evidence of such events. In this work, we ask whether character-based methods can predict gene transfers. Their advantage over sequences is that homologous genes can have low DNA similarity, but still have retained enough important common motifs that allow them to have common character traits, for instance the same functional or expression profile. A phylogeny that has two separate clades that acquired the same character independently might indicate the presence of a transfer even in the absence of sequence similarity. We introduce perfect transfer networks, which are phylogenetic networks that can explain the character diversity of a set of taxa under the assumption that characters have unique births, and that once a character is gained it is rarely lost. Examples of such traits include transposable elements, biochemical markers and emergence of organelles, just to name a few. We study the differences between our model and two similar models: perfect phylogenetic networks and ancestral recombination networks. Our goals are to initiate a study on the structural and algorithmic properties of perfect transfer networks. We then show that in polynomial time, one can decide whether a given network is a valid explanation for a set of taxa, and show how, for a given tree, one can add transfer edges to it so that it explains a set of taxa. We finally provide lower and upper bounds on the number of transfers required to explain a set of taxa, in the worst case.
Phylogenetic trees defined by at most three characters
In evolutionary biology, phylogenetic trees are commonly inferred from a set of characters (partitions) of a collection of biological entities (e.g., species or individuals in a population). Such characters naturally arise from molecular sequences or morphological data. Interestingly, it has been known for some time that any binary phylogenetic tree can be (convexly) defined by a set of at most four characters, and that there are binary phylogenetic trees for which three characters are not enough. Thus, it is of interest to characterise those phylogenetic trees that are defined by a set of at most three characters. In this paper, we provide such a characterisation, in particular proving that a binary phylogenetic tree $T$ is defined by a set of at most three characters precisely if $T$ has no internal subtree isomorphic to a certain tree.
On the free commutative monoid over a positive operad
We study algebraic structures on the free commutative twisted algebra generated by a positive operad $\mathbf q$, in the framework of vector species. Given a nonunital commutative twisted algebra structure $μ$ on $\mathbf q$, we introduce the notion of $μ$-compatible operad structure, leading to a nonunital operad structure on $\mathbf E \circ \mathbf q$, where $\mathbf E$ stands for the exponential species. Next, we define nested pre-Lie operads (NPL-operads), a weak form of the notion of operad, in which the nested associativity axiom is weakened down to a nested pre-Lie condition. This structure is new up to our knowledge. Several constructions of NPL-operads are presented. Finally, we define algebras over a NPL-operad, based on the notion of polynomial functions.
Inferring Long-term Dynamics of Ecological Communities Using Combinatorics
In an increasingly changing world, predicting the fate of species across the globe has become a major concern. Understanding how the population dynamics of various species and communities will unfold requires predictive tools that experimental data alone can not capture. Here, we introduce our combinatorial framework, Widespread Ecological Networks and their Dynamical Signatures (WENDyS) which, using data on the relative strengths of interactions and growth rates within a community of species predicts all possible long-term outcomes of the community. To this end, WENDyS partitions the multidimensional parameter space (formed by the strengths of interactions and growth rates) into a finite number of regions, each corresponding to a unique set of coarse population dynamics. Thus, WENDyS ultimately creates a library of all possible outcomes for the community. On the one hand, our framework avoids the typical ``parameter sweeps'' that have become ubiquitous across other forms of mathematical modeling, which can be computationally expensive for ecologically realistic models and examples. On the other hand, WENDyS opens the opportunity for interdisciplinary teams to use standard experimental data (i.e., strengths of interactions and growth rates) to filter down the possible end states of a community. To demonstrate the latter, here we present a case study from the Indonesian Coral Reef. We analyze how different interactions between anemone and anemonefish species lead to alternative stable states for the coral reef community, and how competition can increase the chance of exclusion for one or more species. WENDyS, thus, can be used to anticipate ecological outcomes and test the effectiveness of management (e.g., conservation) strategies.
On the combinatorics of Lotka-Volterra equations
Published in Physica A, Volume 670, 15 July 2025, 130484
• Search Publication
We study an approach to obtaining the exact formal solution of the 2-species Lotka-Volterra equation based on combinatorics and generating functions. By employing a combination of Carleman linearization and Mori-Zwanzig reduction techniques, we transform the nonlinear equations into a linear system, allowing for the derivation of a formal solution. The Mori-Zwanzig reduction reduces to an expansion which we show can be interpreted as a directed and weighted lattice path walk, which we use to obtain a representation of the system dynamics as walks of fixed length. The exact solution is then shown to be dependent on the generator of weighted walks. We show that the generator can be obtained by the solution of PDE which in turn is equivalent to a particular Koopman evolution of nonlinear observables.
Shared ancestry graphs and symbolic arboreal maps
A network $N$ on a finite set $X$, $|X|\geq 2$, is a connected directed acyclic graph with leaf set $X$ in which every root in $N$ has outdegree at least 2 and no vertex in $N$ has indegree and outdegree equal to 1; $N$ is arboreal if the underlying unrooted, undirected graph of $N$ is a tree. Networks are of interest in evolutionary biology since they are used, for example, to represent the evolutionary history of a set $X$ of species whose ancestors have exchanged genes in the past. For $M$ some arbitrary set of symbols, $d:{X \choose 2} \to M \cup \{\odot\}$ is a symbolic arboreal map if there exists some arboreal network $N$ whose vertices with outdegree two or more are labelled by elements in $M$ and so that $d(\{x,y\})$, $\{x,y\} \in {X \choose 2}$, is equal to the label of the least common ancestor of $x$ and $y$ in $N$ if this exists and $\odot$ else. Important examples of symbolic arboreal maps include the symbolic ultrametrics, which arise in areas such as game theory, phylogenetics and cograph theory. In this paper we show that a map $d:{X \choose 2} \to M \cup \{\odot\}$ is a symbolic arboreal map if and only if $d$ satisfies certain 3- and 4-point conditions and the graph with vertex set $X$ and edge set consisting of those pairs $\{x,y\} \in {X \choose 2}$ with $d(\{x,y\}) \neq \odot$ is Ptolemaic. To do this, we introduce and prove a key theorem concerning the shared ancestry graph for a network $N$ on $X$, where this is the graph with vertex set $X$ and edge set consisting of those $\{x,y\} \in {X \choose 2}$ such that $x$ and $y$ share a common ancestor in $N$. In particular, we show that for any connected graph $G$ with vertex set $X$ and edge clique cover $K$ in which there are no two distinct sets in $K$ with one a subset of the other, there is some network with $|K|$ roots and leaf set $X$ whose shared ancestry graph is $G$.
Agreement forests of caterpillar trees: complexity, kernelization and branching
Given a set $X$ of species, a phylogenetic tree is an unrooted binary tree whose leaves are bijectively labelled by $X$. Such trees can be used to show the way species evolve over time. One way of understanding how topologically different two phylogenetic trees are, is to construct a minimum-size agreement forest: a partition of $X$ into the smallest number of blocks, such that the blocks induce homeomorphic, non-overlapping subtrees in both trees. This comparison yields insight into commonalities and differences in the evolution of $X$ across the two trees. Computing a smallest agreement forest is NP-hard (Hein, Jiang, Wang and Zhang, Discrete Applied Mathematics 71(1-3), 1996). In this work we study the problem on caterpillars, which are path-like phylogenetic trees. We will demonstrate that, even if we restrict the input to this highly restricted subclass, the problem remains NP-hard and is in fact APX-hard. Furthermore we show that for caterpillars two standard reductions rules well known in the literature yield a tight kernel of size at most $7k$, compared to $15k$ for general trees (Kelk and Simone, SIAM Journal on Discrete Mathematics 33(3), 2019). Finally we demonstrate that we can determine if two caterpillars have an agreement forest with at most $k$ blocks in $O^*(2.49^k)$ time, compared to $O^*(3^k)$ for general trees (Chen, Fan and Sze, Theoretical Computater Science 562, 2015), where $O^*(.)$ suppresses polynomial factors.
Evaluating The Impact Of Species Specialisation On Ecological Network Robustness Using Analytic Methods
Ecological networks describe the interactions between different species, informing us of how they rely on one another for food, pollination and survival. If a species in an ecosystem is under threat of extinction, it can affect other species in the system and possibly result in their secondary extinction as well. Understanding how (primary) extinctions cause secondary extinctions on ecological networks has been considered previously using computational methods. However, these methods do not provide an explanation for the properties which make ecological networks robust, and can be computationally expensive. We develop a new analytic model for predicting secondary extinctions which requires no non-deterministic computational simulation. Our model can predict secondary extinctions when primary extinctions occur at random or due to some targeting based on the number of links per species or risk of extinction, and can be applied to an ecological network of any number of layers. Using our model, we consider how false positives and negatives in network data affect predictions for network robustness. We have also extended the model to predict scenarios in which secondary extinctions occur once species lose a certain percentage of interaction strength, and to model the loss of interactions as opposed to just species extinction. From our model, it is possible to derive new analytic results such as how ecological networks are most robust when secondary species degree variance is minimised. Additionally, we show that both specialisation and generalisation in distribution of interaction strength can be advantageous for network robustness, depending upon the extinction scenario being considered.