Papers by François Bienvenu
9 paper(s) by this author
· All BibTeX
A parameterized family of balance indices for phylogenetic networks
We introduce a new family of balance indices for phylogenetic networks: the $H_α$ indices, where $α$ is a positive real number. This family includes the $B_2$ index as a special case ($α= 1$) and provides a natural extension of the Sackin index to phylogenetic networks. We show that the $H_α$ indices share many structural properties with the $B_2$ index, most notably a "grafting property" that makes it possible to express the $H_α$ index of a network in terms of the $H_α$ indices of its biconnected components. These properties allow us to identify networks that minimize / maximize $H_α$ for various classes of phylogenetic networks, and to study its distribution for several models of random trees and networks (in particular, Galton-Watson trees and binary Markov branching trees, with a focus on the Yule and PDA models). Finally, we show how local limits can be used to analyze the asymptotic behavior of $H_α$ for large trees and networks, and we obtain general results for the moments of $H_α$ for a broad class of random phylogenetic networks known as blowups of Galton-Watson trees.
Mathematically tractable models of random phylogenetic networks: an overview of some recent developments
Models of random phylogenetic networks have been used since the inception of the field, but the introduction and rigorous study of mathematically tractable models is a much more recent topic that has gained momentum in the last 5 years. This manuscript discusses some recent developments in the field through a selection of examples. The emphasis is on the techniques rather than on the results themselves, and on probabilistic tools rather than on combinatorial ones.
The $B_2$ index of galled trees
In recent years, there has been an effort to extend the classical notion of phylogenetic balance, originally defined in the context of trees, to networks. One of the most natural ways to do this is with the so-called $B_2$ index. In this paper, we study the $B_2$ index for a prominent class of phylogenetic networks: galled trees. We show that the $B_2$ index of a uniform leaf-labeled galled tree converges in distribution as the network becomes large. We characterize the corresponding limiting distribution, and show that its expected value is 2.707911858984... This is the first time that a balance index has been studied to this level of detail for a random phylogenetic network.
One specificity of this work is that we use two different and independent approaches, each with its advantages: analytic combinatorics, and local limits. The analytic combinatorics approach is more direct, as it relies on standard tools; but it involves slightly more complex calculations. Because it has not previously been used to study such questions, the local limit approach requires developing an extensive framework beforehand; however, this framework is interesting in itself and can be used to tackle other similar problems.
0-1 laws for pattern occurrences in phylogenetic trees and networks
Published in Bull. Math. Biol. 86, 94 (2024)
• View Publication
• BIB
In a recent paper, the question of determining the fraction of binary trees that contain a fixed pattern known as the snowflake was posed. We show that this fraction goes to 1, providing two very different proofs: a purely combinatorial one that is quantitative and specific to this problem; and a proof using branching process techniques that is less explicit, but also much more general, as it applies to any fixed patterns and can be extended to other trees and networks. In particular, it follows immediately from our second proof that the fraction of $d$-ary trees (resp. level-$k$ networks) that contain a fixed $d$-ary tree (resp. level-$k$ network) tends to $1$ as the number of leaves grows.
Revisiting Shao and Sokal's $B_2$ index of phylogenetic balance
Published in Journal of Mathematical Biology 83:52 (2021)
• View Publication
• BIB
Measures of phylogenetic balance, such as the Colless and Sackin indices, play an important role in phylogenetics. Unfortunately, these indices are specifically designed for phylogenetic trees, and do not extend naturally to phylogenetic networks (which are increasingly used to describe reticulate evolution). This led us to consider a lesser-known balance index, whose definition is based on a probabilistic interpretation that is equally applicable to trees and to networks. This index, known as the $B_2$ index, was first proposed by Shao and Sokal in 1990. Surprisingly, it does not seem to have been studied mathematically since. Likewise, it is used only sporadically in the biological literature, where it tends to be viewed as arcane. In this paper, we study mathematical properties of $B_2$ such as its expectation and variance under the most common models of random trees and its extremal values over various classes of phylogenetic networks. We also assess its relevance in biological applications, and find it to be comparable to that of the Colless and Sackin indices. Altogether, our results call for a reevaluation of the status of this somewhat forgotten measure of phylogenetic balance.
Combinatorial and stochastic properties of ranked tree-child networks
Published in Random Structures & Algorithms, 60(4):653-689 (2022)
• View Publication
• BIB
Tree-child networks are a recently-described class of directed acyclic graphs that have risen to prominence in phylogenetics (the study of evolutionary trees and networks). Although these networks have a number of attractive mathematical properties, many combinatorial questions concerning them remain intractable. In this paper, we show that endowing these networks with a biologically relevant ranking structure yields mathematically tractable objects, which we term ranked tree-child networks (RTCNs). We explain how to derive exact and explicit combinatorial results concerning the enumeration and generation of these networks. We also explore probabilistic questions concerning the properties of RTCNs when they are sampled uniformly at random. These questions include the lengths of random walks between the root and leaves (both from the root to the leaves and from a leaf to the root); the distribution of the number of cherries in the network; and sampling RTCNs conditional on displaying a given tree. We also formulate a conjecture regarding the scaling limit of the process that counts the number of lineages in the ancestry of a leaf. The main idea in this paper, namely using ranking as a way to achieve combinatorial tractability, may also extend to other classes of networks.
The Moran forest
Published in Random Structures & Algorithms, 59(2):155-188 (2021)
• View Publication
• BIB
Starting from any graph on $\{1, \ldots, n\}$, consider the Markov chain where at each time-step a uniformly chosen vertex is disconnected from all of its neighbors and reconnected to another uniformly chosen vertex. This Markov chain has a stationary distribution whose support is the set of non-empty forests on $\{1, \ldots, n\}$. The random forest corresponding to this stationary distribution has interesting connections with the uniform rooted labeled tree and the uniform attachment tree. We fully characterize its degree distribution, the distribution of its number of trees, and the limit distribution of the size of a tree sampled uniformly. We also show that the size of the largest tree is asymptotically $α\log n$, where $α= (1 - \log(e - 1))^{-1} \approx 2.18$, and that the degree of the most connected vertex is asymptotically $\log n / \log\log n$.
Positive association of the oriented percolation cluster in randomly oriented graphs
Published in Combinatorics, Probability and Computing, 28(6):811-815 (2019)
• View Publication
• BIB
Consider any fixed graph whose edges have been randomly and independently oriented, and write $\{S \leadsto i\}$ to indicate that there is an oriented path going from a vertex $s \in S$ to vertex $i$. Narayanan (2016) proved that for any set $S$ and any two vertices $i$ and $j$, $\{S \leadsto i\}$ and $\{S \leadsto j\}$ are positively correlated. His proof relies on the Ahlswede-Daykin inequality, a rather advanced tool of probabilistic combinatorics.
In this short note, I give an elementary proof of the following, stronger result: writing $V$ for the vertex set of the graph, for any source set $S$, the events $\{S \leadsto i\}$, $i \in V$, are positively associated -- meaning that the expectation of the product of increasing functionals of the family $\{S \leadsto i\}$ for $i \in V$ is greater than the product of their expectations.
The split-and-drift random graph, a null model for speciation
Published in Stochastic Processes and their Applications 129(6):2010-2048 (2019)
• View Publication
• BIB
We introduce a new random graph model motivated by biological questions relating to speciation. This random graph is defined as the stationary distribution of a Markov chain on the space of graphs on $\{1, \ldots, n\}$. The dynamics of this Markov chain is governed by two types of events: vertex duplication, where at constant rate a pair of vertices is sampled uniformly and one of these vertices loses its incident edges and is rewired to the other vertex and its neighbors; and edge removal, where each edge disappears at constant rate. Besides the number of vertices $n$, the model has a single parameter $r_n$.
Using a coalescent approach, we obtain explicit formulas for the first moments of several graph invariants such as the number of edges or the number of complete subgraphs of order $k$. These are then used to identify five non-trivial regimes depending on the asymptotics of the parameter $r_n$. We derive an explicit expression for the degree distribution, and show that under appropriate rescaling it converges to classical distributions when the number of vertices goes to infinity. Finally, we give asymptotic bounds for the number of connected components, and show that in the sparse regime the number of edges is Poissonian.