species
265 papers tagged with this keyword
Free pre-lie algebras of finite posets
Published
• View Publication
• BIB
In this paper, we first recall the construction of a twisted pre-Lie algebra structure on the species of finite connected topological spaces. Then we construct the corresponding nonassociative permutative coproduct, and we prove that the vector space generated by isomorphism classes of finite posets is a free pre-Lie algebra and is a co-free non-associative permutative coalgebra. In the end, we give an explicit duality between the non-associative permutative product and the proposed non-associative permutative coproduct. Finally, we prove that the results in this paper remain true for the finite connected topological spaces.
A Model for Birdwatching and other Chronological Sampling Activities
Published
• View Publication
• BIB
In many real life situations one has $m$ types of random events happening in chronological order within a time interval and one wishes to predict various milestones about these events or their subsets. An example is birdwatching. Suppose we can observe up to $m$ different types of birds during a season. At any moment a bird of type $i$ is observed with some probability. There are many natural questions a birdwatcher may have: how many observations should one expect to perform before recording all types of birds? Is there a time interval where the researcher is most likely to observe all species? Or, what is the likelihood that several species of birds will be observed at overlapping time intervals? Our paper answers these questions using a new model based on random interval graphs. This model is a natural follow up to the famous coupon collector's problem.
Hopf monoids of set families
A \textit{grounded set family} on $I$ is a subset $F\subseteq2^I$ such that $\emptyset\in F$. We study a linearized Hopf monoid \textbf{SF} on grounded set families, with restriction and contraction inspired by the corresponding operations for antimatroids. Many known combinatorial species, including simplicial complexes and matroids, form Hopf submonoids of \textbf{SF}, although not always with the "standard" Hopf structure (for example, our contraction operation is not the usual contraction of matroids). We use the topological methods of Aguiar and Ardila to obtain a cancellation-free antipode formula for the Hopf submonoid of lattices of order ideals of finite posets. Furthermore, we prove that the Hopf algebra of lattices of order ideals of chain gangs extends the Hopf algebra of symmetric functions, and that its character group extends the group of formal power series in one variable with constant term 1 under multiplication.
Planar Rooted Phylogenetic Networks
Published
• View Publication
• BIB
A rooted phylogenetic network is a directed acyclic graph with a single root, whose sinks correspond to a set of species. As such networks are useful for representing the evolution of species that have undergone reticulate evolution, there has been great interest in developing the theory behind and algorithms for constructing them. However, unlike evolutionary trees, these networks can be highly non-planar, which can make them difficult to visualise and interpret. Here we investigate properties of planar rooted phylogenetic networks and algorithms for deciding whether or not rooted networks have certain special planarity properties. In particular, we introduce three natural subclasses of planar rooted phylogenetic networks and show that they form a hierarchy. In addition, for the well-known level-k networks, we show that level-1, -2, -3 networks are always outer, terminal, and upward planar, respectively, and that level-4 networks are not necessarily planar. Finally, we show that a regular network is terminal planar if and only if it is pyramidal. Our results make use of the highly developed field of planar digraphs, and we believe that the link between phylogenetic networks and planar graphs should prove useful in future for developing new approaches to both construct and visualise phylogenetic networks.
Joint $q$-moments and shift invariance for the multi-species $q$-TAZRP on the infinite line
This paper presents a novel method for computing certain particle locations in the multi-species $q$-TAZRP (totally asymmetric zero range process). The method is based on a decomposition of the process into its discrete-time embedded Markov chain, which is described more generally as a monotone process on a graded partially ordered set; and an independent family of exponential random variables. A further ingredient is explicit contour integral formulas for the transition probabilities of the $q$-TAZRP. The main result of this method is a shift invariance for the multi-species $q$-TAZRP on the infinite line.
By a previously known Markov duality result, these particle locations are the same as joint $q$-moments. One particular special case is that for step initial conditions, ordered multi-point joint $q$-moments of the $n$-species $q$-TAZRP match the $n$-point joint $q$-moments of the single-species $q$-TAZRP. Thus, we conjecture that the Airy$_2$ process describes the joint multi-point fluctuations of multi-species $q$-TAZRP.
As a probabilistic application of this result, we find explicit contour integral formulas for the joint $q$-moments of the multi-species $q$-TAZRP in the diffusive scaling regime.
A Linear Time, and Constant Space, Algorithm to Compute the Mixed Moments of the Multivariate Normal Distributions
Using recurrences gotten from the Apagodu-Zeilberger Multivariate Almkvist-Zeilberger algorithm we present a linear-time, and constant-space, algorithm to compute the general mixed moments of the k-variate general normal distribution, with any covariance matrix, for any specific k. Besides their obvious importance in statistics, these numbers are also very significant in enumerative combinatorics, since they count in how many ways, in a species with k different genders, a bunch of individuals can all get married, keeping track of the different kinds of heterosexual marriages. We completely implement our algorithm (with an accompanying Maple package, MVNM.txt) for the bivariate and trivariate cases (and hence taking care of our own 2-sex society and a putative 3-sex society), but alas, the actual recurrences for larger k took too long for us to compute. We leave them as computational challenges.
Forest-based networks
Published
• View Publication
• BIB
In evolutionary studies it is common to use phylogenetic trees to represent the evolutionary history of a set of species. However, in case the transfer of genes or other genetic information between the species or their ancestors has occurred in the past, a tree may not provide a complete picture of their history. In such cases,tree-based phylogenetic networks can provide a useful, more refined representation of the species evolution. Such a network is essentially a phylogenetic tree with some arcs added between the tree edges so as to represent reticulate events such as gene transfer. Even so, this model does not permit the representation of evolutionary scenarios where reticulate events have taken place between different subfamilies or lineages of species. To represent such scenarios, in this paper we introduce the notion of a forest-based phylogenetic network, that is, a collection of leaf-disjoint phylogenetic trees on a set of species with arcs added between the edges of distinct trees within the collection. Forest-based networks include the recently introduced class of overlaid species forests which are used to model introgression. As we shall see, even though the definition of forest-based networks is closely related to that of tree-based networks, they lead to new mathematical theory which complements that of tree-based networks. As well as studying the relationship of forest-based networks with other classes of phylogenetic networks, such as tree-child networks and universal tree-based networks, we present some characterizations of some special classes of forest-based networks. We expect that our results will be useful for developing new models and algorithms to understand reticulate evolution, such as gene transfer between collections of bacteria that live in different environments.
What makes a reaction network "chemical"?
Reaction networks (RNs) comprise a set $X$ of species and a set $\mathscr{R}$ of reactions $Y\to Y'$, each converting a multiset of educts $Y\subseteq X$ into a multiset $Y'\subseteq X$ of products. RNs are equivalent to directed hypergraphs. However, not all RNs necessarily admit a chemical interpretation. Instead, they might contradict fundamental principles of physics such as the conservation of energy and mass or the reversibility of chemical reactions. The consequences of these necessary conditions for the stoichiometric matrix $\mathbf{S} \in \mathbb{R}^{X\times\mathscr{R}}$ have been discussed extensively in the literature. Here, we provide sufficient conditions for $\mathbf{S}$ that guarantee the interpretation of RNs in terms of balanced sum formulas and structural formulas, respectively.
Chemically plausible RNs allow neither a perpetuum mobile, i.e., a "futile cycle" of reactions with non-vanishing energy production, nor the creation or annihilation of mass. Such RNs are said to be thermodynamically sound and conservative. For finite RNs, both conditions can be expressed equivalently as properties of $\mathbf{S}$. The first condition is vacuous for reversible networks, but it excludes irreversible futile cycles and - in a stricter sense - futile cycles that even contain an irreversible reaction. The second condition is equivalent to the existence of a strictly positive reaction invariant. Furthermore, it is sufficient for the existence of a realization in terms of sum formulas, obeying conservation of "atoms". In particular, these realizations can be chosen such that any two species have distinct sum formulas, unless $\mathbf{S}$ implies that they are "obligatory isomers". In terms of structural formulas, every compound is a labeled multigraph, in essence a Lewis formula, and reactions comprise only a rearrangement of bonds such that the total bond order is preserved.
Introducing DASEP: the doubly asymmetric simple exclusion process
Published in "Séminaire Lotharingien de Combinatoire" 87B (2023), pp. 81-91
• Search Publication
Research in combinatorics has often explored the asymmetric simple exclusion process (ASEP). The ASEP, inspired by examples from statistical mechanics, involves particles of various species moving around a lattice. With the traditional ASEP particles of a given species can move but do not change species. In this paper a new combinatorial formalism, the DASEP (doubly asymmetric simple exclusion process), is explored. The DASEP is inspired by biological processes where, unlike the ASEP, the particles can change from one species to another. The combinatorics of the DASEP on a one dimensional lattice are explored, including the associated generating function. The stationary probabilities of the DASEP are explored, and results are proven relating these stationary probabilities to those of the simpler ASEP.
On the quartet distance given partial information
Published
• View Publication
• BIB
Let $T$ be an arbitrary phylogenetic tree with $n$ leaves. It is well-known that the average quartet distance between two assignments of taxa to the leaves of $T$ is $\frac 23 \binom{n}{4}$. However, a longstanding conjecture of Bandelt and Dress asserts that $(\frac 23 +o(1))\binom{n}{4}$ is also the {\em maximum} quartet distance between two assignments. While Alon, Naves, and Sudakov have shown this indeed holds for caterpillar trees, the general case of the conjecture is still unresolved. A natural extension is when partial information is given: the two assignments are known to coincide on a given subset of taxa. The partial information setting is biologically relevant as the location of some taxa (species) in the phylogenetic tree may be known, and for other taxa it might not be known. What can we then say about the average and maximum quartet distance in this more general setting? Surprisingly, even determining the {\em average} quartet distance becomes a nontrivial task in the partial information setting and determining the maximum quartet distance is even more challenging, as these turn out to be dependent of the structure of $T$. In this paper we prove nontrivial asymptotic bounds that are sometimes tight for the average quartet distance in the partial information setting. We also show that the Bandelt and Dress conjecture does not generally hold under the partial information setting. Specifically, we prove that there are cases where the average and maximum quartet distance substantially differ.
Convex characters, algorithms and matchings
Published
• View Publication
• BIB
Phylogenetic trees are used to model evolution: leaves are labelled to represent contemporary species ("taxa") and interior vertices represent extinct ancestors. Informally, convex characters are measurements on the contemporary species in which the subset of species (both contemporary and extinct) that share a given state, form a connected subtree. In \cite{KelkS17} it was shown how to efficiently count, list and sample certain restricted subfamilies of convex characters, and algorithmic applications were given. We continue this work in a number of directions. First, we show how combining the enumeration of convex characters with existing parameterised algorithms can be used to speed up exponential-time algorithms for the \emph{maximum agreement forest problem} in phylogenetics. Second, we re-visit the quantity $g_2(T)$, defined as the number of convex characters on $T$ in which each state appears on at least 2 taxa. We use this to give an algorithm with running time $O( φ^{n} \cdot \text{poly}(n) )$, where $φ\approx 1.6181$ is the golden ratio and $n$ is the number of taxa in the input trees, for computation of \emph{maximum parsimony distance on two state characters}. By further restricting the characters counted by $g_2(T)$ we open an interesting bridge to the literature on enumeration of matchings. By crossing this bridge we improve the running time of the aforementioned parsimony distance algorithm to $O( 1.5895^{n} \cdot \text{poly}(n) )$, and obtain a number of new results in themselves relevant to enumeration of matchings on at-most binary trees.
A lattice structure for ancestral configurations arising from the relationship between gene trees and species trees
Published
• View Publication
• BIB
To a given gene tree topology $G$ and species tree topology $S$ with leaves labeled bijectively from a fixed set $X$, one can associate a set of ancestral configurations, each of which encodes a set of gene lineages that can be found at a given node of the species tree. We introduce a lattice structure on ancestral configurations, studying the directed graphs that provide graphical representations of lattices of ancestral configurations. For a matching gene tree topology and species tree topology $G=S$, we present a method for defining the digraph of ancestral configurations from the tree topology by using iterated cartesian products of graphs. We show that a specific set of paths on the digraph of ancestral configurations is in bijection with the set of labeled histories -- a well-known phylogenetic object that enumerates possible temporal orderings of the coalescences of a tree. For each of a series of tree families, we obtain closed-form expressions for the number of labeled histories by using this bijection to count paths on associated digraphs. Finally, we prove that our lattice construction extends to nonmatching tree pairs, and we use it to characterize pairs $(G,S)$ having the maximal number of ancestral configurations for a fixed $G$. We discuss how the construction provides new methods for performing enumerations of combinatorial aspects of gene and species trees.
Quasi-Best Match Graphs
Quasi-best match graphs (qBMGs) are a hereditary class of directed, properly vertex-colored graphs. They arise naturally in mathematical phylogenetics as a generalization of best match graphs, which formalize the notion of evolutionary closest relatedness of genes (vertices) in multiple species (vertex colors). They are explained by rooted trees whose leaves correspond to vertices. In contrast to BMGs, qBMGs represent only best matches at a restricted phylogenetic distance. We provide characterizations of qBMGs that give rise to polynomial-time recognition algorithms and identify the BMGs as the qBMGs that are color-sink-free. Furthermore, two-colored qBMGs are characterized as directed graphs satisfying three simple local conditions, two of which have appeared previously, namely bi-transitivity in the sense of Das et al. (2021) and a hierarchy-like structure of out-neighborhoods, i.e., $N(x)\cap N(y)\in\{N(x),N(y),\emptyset\}$ for any two vertices $x$ and $y$. Further results characterize qBMGs that can be explained by binary phylogenetic trees.
Encoding and ordering X-cactuses
Published
• View Publication
• BIB
Phylogenetic networks are a generalization of evolutionary or phylogenetic trees that are commonly used to represent the evolution of species which cross with one another. A special type of phylogenetic network is an {\em $X$-cactus}, which is essentially a cactus graph in which all vertices with degree less than three are labelled by at least one element from a set $X$ of species. In this paper, we present a way to {\em encode} $X$-cactuses in terms of certain collections of partitions of $X$ that naturally arise from $X$-cactuses. Using this encoding, we also introduce a partial order on the set of $X$-cactuses (up to isomorphism), and derive some structural properties of the resulting partially ordered set. This includes an analysis of some properties of its least upper and greatest lower bounds. Our results not only extend some fundamental properties of phylogenetic trees to $X$-cactuses, but also provides a new approach to solving topical problems in phylogenetic network theory such as deriving consensus networks.
"Normal" phylogenetic networks may be emerging as the leading class
Published
• View Publication
• BIB
The rich and varied ways that genetic material can be passed between species has motivated extensive research into the theory of phylogenetic networks. Features that align with biological processes, or with desirable mathematical properties, have been used to define classes and prove results, with the goal of developing the theoretical foundations for network reconstruction methods. We may have now reached the point where a collection of recent results can be drawn together to make one class of network, the \emph{normal} networks, a leading contender, sitting in the sweet spot between biological relevance and mathematical tractability.
On the Complexity of Optimising Variants of Phylogenetic Diversity on Phylogenetic Networks
Published
• View Publication
• BIB
Phylogenetic Diversity (PD) is a prominent quantitative measure of the biodiversity of a collection of present-day species (taxa). This measure is based on the evolutionary distance among the species in the collection. Loosely speaking, if $\mathcal{T}$ is a rooted phylogenetic tree whose leaf set $X$ represents a set of species and whose edges have real-valued lengths (weights), then the PD score of a subset $S$ of $X$ is the sum of the weights of the edges of the minimal subtree of $\mathcal{T}$ connecting the species in $S$. In this paper, we define several natural variants of the PD score for a subset of taxa which are related by a known rooted phylogenetic network. Under these variants, we explore, for a positive integer $k$, the computational complexity of determining the maximum PD score over all subsets of taxa of size $k$ when the input is restricted to different classes of rooted phylogenetic networks
Phylogenetic Diversity Rankings in the Face of Extinctions: the Robustness of the Fair Proportion Index
Published
• View Publication
• BIB
Planning for the protection of species often involves difficult choices about which species to prioritize, given constrained resources. One way of prioritizing species is to consider their "evolutionary distinctiveness", i.e. their relative evolutionary isolation on a phylogenetic tree. Several evolutionary isolation metrics or phylogenetic diversity indices have been introduced in the literature, among them the so-called Fair Proportion index (also known as the "evolutionary distinctiveness" score). This index apportions the total diversity of a tree among all leaves, thereby providing a simple prioritization criterion for conservation.
Here, we focus on the prioritization order obtained from the Fair Proportion index and analyze the effects of species extinction on this ranking. More precisely, we analyze the extent to which the ranking order may change when some species go extinct and the Fair Proportion index is re-computed for the remaining taxa. We show that for each phylogenetic tree, there are edge lengths such that the extinction of one leaf per cherry completely reverses the ranking. Moreover, we show that even if only the lowest ranked species goes extinct, the ranking order may drastically change. We end by analyzing the effects of these two extinction scenarios (extinction of the lowest ranked species and extinction of one leaf per cherry) for a collection of empirical and simulated trees. In both cases, we can observe significant changes in the prioritization orders, highlighting the empirical relevance of our theoretical findings.
Comparing the topology of phylogenetic network generators
Published
• View Publication
• BIB
Phylogenetic networks represent evolutionary history of species and can record natural reticulate evolutionary processes such as horizontal gene transfer and gene recombination. This makes phylogenetic networks a more comprehensive representation of evolutionary history compared to phylogenetic trees. Stochastic processes for generating random trees or networks are important tools in evolutionary analysis, especially in phylogeny reconstruction where they can be utilized for validation or serve as priors for Bayesian methods. However, as more network generators are developed, there is a lack of discussion or comparison for different generators. To bridge this gap, we compare a set of phylogenetic network generators by profiling topological summary statistics of the generated networks over the number of reticulations and comparing the topological profiles.
Twisted pre-lie algebras of finite topological spaces
Published
• View Publication
• BIB
In this paper, we first study the species of finite topological spaces recently considered by F. Fauvet, L. Foissy, and D. Manchon. Then, we construct a twisted pre-Lie structure on the species of connected finite topological spaces. The underlying pre-Lie structure defines a coproduct on the species of finite topological spaces different from those already defined by the Authors above. In the end, we illustrate the link between the Grossman-Larson product and the proposed coproduct.
Distinguishing Level-2 Phylogenetic Networks Using Phylogenetic Invariants
In phylogenetics, it is important for the phylogenetic network model parameters to be identifiable so that the evolutionary histories of a group of species can be consistently inferred. However, as the complexity of the phylogenetic network models grows, the identifiability of network models becomes increasingly difficult to analyze. As an attempt to analyze the identifiability of network models, we check whether two networks are distinguishable. In this paper, we specifically study the distinguishability of phylogenetic network models associated with level-2 networks. Using an algebraic approach, namely using discrete Fourier transformation, we present some results on the distinguishability of some level-2 networks, which generalize earlier work on the distinguishability of level-1 networks. In particular, we study simple and semisimple level-2 networks. Simple and semisimple level-2 networks can be thought as generalizations of level-1 sunlet and cycle networks, respectively. Moreover, we also compare the varieties associated with semisimple level-2 and cycle networks.