math.PR ↗ arXiv
440 papers in this category
Universal dualities for Wilson loops in lattice Yang-Mills
We identify a universal finite-$N$ structure underlying Wilson loop expectations in lattice Yang-Mills, in any dimension $d\geq 2$, for gauge group $\mathrm{U}(N)$, and for arbitrary smooth central plaquette actions. The starting point is a state-sum expansion in plaquette labels by irreducible representations, in which each term factorizes into an action-dependent spectral weight and an action-independent topological coefficient. We then analyze these coefficients in three exact ways: as a gauge/string expansion over decorated spanning surfaces, as a local spin-foam/channel model on the dual incidence graph, and as a universal finite-$N$ master loop equation that closes on the coefficient side. As a consequence, several recent Wilson-action results are recovered as specializations of our broader action-agnostic framework.
A counter-example to persistence in generalised preferential attachment trees
Consider a generalised preferential attachment tree with attachment function $f$, that is a random tree, where at each time-step a node connects to an existing node $v$ with probability proportional to $f(\mathrm{deg}(v))$, where $\mathrm{deg}(v)$ denotes the degree of the node in the existing tree. We provide a counter-example to a conjecture of the author asserting that under the assumption $\sum_{j=1}^{\infty} \frac{1}{f(j)^2} < \infty$ there is a persistent hub in the model, that is, a single node that has the maximal degree for all but finitely many time-steps. The counter-example is a minor modification of a related counter-example due to Galganov and Ilienko.
Convolution, cumulants and infinitesimal generators in the formal power series ring
We extend the notions of finite free convolution and finite free cumulants to the setting of formal power series by introducing their natural analogues, namely $t$-deformed convolution and $t$-deformed cumulants. In this framework, we establish $t$-deformed analogues of the law of large numbers and the central limit theorem, revealing structural parallels with classical, free, and finite free probability theories. We show that the case $t=-1$ recovers classical convolution at the level of moment generating functions, thereby connecting the theory directly to classical probability. We further investigate the infinitesimal generators associated with $\boxplus^t$-continuous semigroups, deriving explicit representation formulas that clarify how these generators describe the infinitesimal evolution of the semigroup. In the case $t = d$, our results yield explicit formulas for finite free infinitesimal generators. In the case $t = -1$, we relate these generators to those of one-dimensional Lévy processes by identifying the corresponding terms in their representations. This establishes a direct connection between $\boxplus^t$-convolution semigroups and classical Lévy-Khintchine-type generators.
Sweet Trims are made of Threes: A càdlàg erasure of the Brownian tree
We present a simple trimming algorithm that generates nested uniform binary plane trees by removing leaves one-by-one using a best-of-three-match procedure. While its one-step transition specializes to the Luczak-Winkler & Caraceni-Stauffer coupling, its scaling limit provides a suprising càdlàg erasure of Brownian trees, reminiscent of SLE theory.
On additive averaging kernels for finite Markov chains
We study additive mixtures of Markov kernels of the form $A_α= αP + (1-α)G$, where $α\in [0,1]$, $P$ is a baseline sampler and $G$ is a Gibbs kernel induced by a partition of the state space. We first motivate the study of $A_α$, which can be interpreted as the projection of a lifted Markov chain. We then consider the minimisation of distance to stationarity under two objectives: the squared Frobenius norm and the Kullback-Leibler (KL) divergence. For the Frobenius objective, we derive explicit trace formulas and identify a Cheeger-type functional that characterises optimal two-block partitions. This yields a structured combinatorial optimisation problem admitting a difference-of-submodular decomposition, enabling efficient approximation via majorisation-minimisation. We also obtain geometric decay rates governed by the absolute spectral gap of $P$. For the KL divergence, we establish convexity-based bounds showing that the divergence of $A_α$ is controlled by those of both $P$ and $G$, thereby reducing partition selection to the Gibbs component. Numerical experiments on the Curie-Weiss model demonstrate that suitable choice of both the partition and the parameter $α$ can significantly accelerate convergence in total variation distance. We observe a consistent trade-off between local exploration and global averaging, with intermediate values of $α$ achieving the best performance across regimes.
FlowBoost Reveals Phase Transitions and Spectral Structure in Finite Free Information Inequalities
Using FlowBoost, a closed-loop deep generative optimization framework for extremal structure discovery, we investigate $\ell^p$-generalizations of the finite free Stam inequality for real-rooted polynomials under finite free additive convolution $\boxplus_n$. At $p=2$, FlowBoost finds the Hermite pair as the unique equality case and reveals the spectral structure of the linearized convolution map at this extremal point. As a result, we conjecture that the singular values of the doubly stochastic coupling matrix $E_n$ on the mean-zero subspace are ${2^{-k/2}:k=1,\ldots,n-1}$, independent of $n$. Conditional on this conjecture, we obtain a sharp local stability constant and the finite free CLT convergence rate, both uniform in $n$. We introduce a one-parameter family of $p$-Stam inequalities using $\ell^p$-Fisher information and prove that the Hermite pair itself violates the inequality for every $p>2$, with the sign of the deficit governed by the $\ell^p$-contraction ratio of $E_n$. Systematic computation via FlowBoost supports the conjecture that $p^*\!=2$ is the sharp critical exponent. For $p<2$, the extremal configurations undergo a bifurcation, meaning that they become non-matching pairs with bimodal root structure, converging back to the Hermite diagonal only as $p\to 2^-$. Our findings demonstrate that FlowBoost, can be an effective tool of mathematical discovery in infinite-dimensional extremal problems.
The smallest singular value of signed random combinatorial matrices
Let $M_n$ be an $n\times n$ signed random combinatorial matrix whose rows are independent and uniformly distributed over the set of $\{-1,0,1\}$-vectors with exactly $n/2$ zero coordinates. Despite the dependence induced by the row constraints, we prove that there exist constants $C,c > 0$ such that for any $\varepsilon\ge0$, \begin{align*} \textbf{P}\left(s_{n}(M_n)\le {\varepsilon}{n^{-1/2}}\right)\le C\varepsilon+e^{-cn}. \end{align*} In particular, the probability that $M_n$ is singular is exponentially small. Our approach builds on the Combinatorial Least Common Denominator (CLCD) introduced by Tran and develops the method in the present constrained setting.
A Strict Gap Between Relaxed and Partition-Constrained Spectral Compression in a Six-State Lumpable Markov Chain
This paper studies a finite reversible lumpable Markov chain for which relaxed spectral compression yields a larger determinant than partition-constrained compression. For a symmetric six-state lumpable chain and the positive operator $T=P^2$, I compare the relaxed benchmark \begin{equation*} \mathfrak D^{\mathrm{rel}}_3(T):=\sup_{U^*U=I_3}\det(U^*TU) \end{equation*} and the partition-constrained benchmark \begin{equation*} \sup_{\mathcal A\,\mathrm{3\text{-}partition}}\det Q_{\mathcal A}(T), \qquad Q_{\mathcal A}(T)=H_{\mathcal A}^*TH_{\mathcal A}. \end{equation*} Here the partition-constrained benchmark is the compression induced by normalized indicator vectors of genuine partitions of the state space. I derive closed formulas for the two analytically central partition families, prove strict upper bounds for both in a local-mode-dominated regime, and combine these bounds with an exhaustive enumeration of all $90$ partitions into three nonempty cells in an explicit six-state model. For this model, one obtains a strict global gap: \begin{equation*} \sup_{\mathcal A}\det Q_{\mathcal A}(T)<\mathfrak D^{\mathrm{rel}}_3(T). \end{equation*} Thus, in this model, indicator-based partition frames are strictly weaker than relaxed orthonormal frames even after global partition-constrained optimization.
Asymptotic enumeration of admixed arrays and a different independence heuristic
We introduce a class of paired binary matrices called admixed arrays, which arise in analyses of large-scale genetic data and can be viewed as weighted edge colorings of complete bipartite graphs. This combinatorial structure gives rise to two natural families of marginal constraints: a row-sum constraint and a paired column-sum constraint, the latter inducing an inequality among entries of the matrix pair. We study the enumeration of admixed arrays under these constraints in dense regimes. First, we obtain exact formulas for the sizes of the families defined by each constraint in isolation and derive a finite-size criterion characterizing when one constraint is more restrictive than the other. In the large-dimension limit, this comparison simplifies to an entropy inequality, yielding an information-theoretic interpretation and a quantifiable error bound in the semi-regular case. We then analyze the asymptotic enumeration of the doubly constrained family in a semi-regular setting. Using saddle-point approximation and probabilistic techniques, we derive a detailed asymptotic expansion for the logarithm of the count, isolating an explicit fourth-moment contribution and establishing quantitative control of the higher-order remainder. A consequence of this analysis is a phenomenon absent from classical binary and integer matrix models: in the regime $N=Θ(P)$ with uniform margins and density bounded away from zero, the two constraint families obey the independence heuristic with a correction factor $1/\sqrt[4]{e}$ rather than the familiar $e^{\pm1/2}$. Numerical experiments corroborate the analytical approximations, and we implement and extend an algorithm of Miller and Harrison (2013) as open-source software to enumerate constrained admixed arrays.
Limit laws for longest edges in empty region graphs
Empty region graphs are graphs whose vertices are points in $\mathbb{R}^d$ and where two vertices are connected by an edge whenever some associated region does not contain any other vertices. We investigate the asymptotic behaviour of long edges in empty region graphs generated by a stationary Poisson process in $\mathbb{R}^d$. {Letting} the intensity of the underlying Poisson process tend to infinity, we consider the associated point process of edge midpoints, suitably transformed edge lengths, and directions of the edges. We prove that it converges in distribution to a Poisson process on $\mathbb{R}^d \times \mathbb{R}\times\mathbb{L}^d$, where $\mathbb{L}^d$ is the space of lines in $\mathbb{R}^d$ through the origin, and that the suitably transformed length of the longest edge with midpoint in an observation window converges in distribution to a Gumbel distributed random variable. Our approach yields explicit error bounds in Kantorovich--Rubinstein distance for the point process convergence {when restricting to an observation window} and in Kolmogorov distance for the maximal edge length. The results apply uniformly to a broad class of empty region graphs, including the Gabriel graph, the relative neighbourhood graph, the beta-skeleton graph, the Mastercard graph, and the Pacman graph.
Random 0/1-polytopes expand rapidly
A 0/1-polytope is the convex hull of a subset $V\subseteq \{0,1\}^n$. A celebrated conjecture of Mihail and Vazirani asserts that the graph of every 0/1-polytope has edge-expansion at least 1. In this paper, we show that typical 0/1-polytopes have significantly stronger expansion. Specifically, if $V$ is formed by sampling each vertex of $\{0,1\}^n$ independently with constant probability $p$, then with high probability the edge-expansion is $Θ(n)$ for $p \in (1/2, 1)$, and $n^{Θ(\log \log n)}$ for $p \in (0, 1/2)$. This improves the previously best known bound $Ω(1)$ due to Ferber, Krivelevich, Sales and Samotij.
Random permutations from $q$-Demazure products
We study the $q$-deformation of the Demazure product model from arXiv:2407.21653. Consider the longest element $w_0$ in $S_n$ written as a reduced word in simple transpositions. Independently delete each transposition with probability $1-p$ and apply the $q$-Demazure product to the remaining ones. We show that the law of the resulting permutation converges as $n \to \infty$ to a deterministic permuton, which coincides with the $q=0$ case studied in arXiv:2407.21653 for adjusted probability $p'=p(1-q)/(1-qp)$. This resolves Conjecture 1.13 from arXiv:2407.21653 and identifies the limiting permuton explicitly.
Short proofs in combinatorics, probability and number theory II
We give a quintet of proofs resulting from questions posed by Erdős. These questions concern ordinary lines in planar point sets, sequences with uniformly small exponential sums, $K_4$-free $4$-critical graphs with few chords in any cycle, a counterexample to a "fewnomial" version of the Erdős--Turán discrepancy bound, and a finiteness theorem for integers $n$ such that $n-a k^2$ is prime for all $k\leq \sqrt{n/a}$ coprime to $n$ (for fixed $a\in\mathbb Z_+$). Each proof is due to an internal model at OpenAI.
The Random Subsequence Model and Uniform Codes for the Deletion Channel
We introduce the Random Subsequence Model, a spin glass model on pairs of random strings $(X,Y) \in \{0,1\}^N \times \{0,1\}^M$ whose partition function counts subsequence embeddings of $Y$ into $X$. We study two variants: the null model, where $X$ and $Y$ are independent and uniform, and the planted model, where $X$ is uniform and $Y$ is a uniformly-random length-$M$ subsequence of $X$. We connect the Random Subsequence Model to longstanding problems in various fields, including the best rate achievable by uniformly-random codes in the deletion channel, the longest common subsequence problem between two random strings, and models of directed polymers in statistical physics.
In the regime where $N,M\to\infty$ at a fixed ratio $α= M/N \in (0,1)$, we exhibit strict asymptotic separations between the null annealed free energy and the quenched free energies of the null and planted models at all values of the density parameter $α$. This suggests that these models are in a spin glass phase at zero temperature throughout the entire dense regime. As a consequence, we show that uniformly-random codes achieve a positive rate in the deletion channel for all deletion probabilities $p\in [0,1),$ settling multiple conjectures of the second author, Isik and Weissman (2024) and proving the first such positive rate result for the regime $p \geq 1/2$.
We also give an exact analytic formula for the annealed free energy of the planted model for all values of the density parameter. This implies a corresponding analytic upper bound on the best rate achievable by uniformly-random codes in the deletion channel, complementing the lower bound from our first result. Our upper and lower bounds for the capacity of the deletion channel under uniform codes are far closer to each other than the best known upper and lower bounds for the capacity of the deletion channel.
Simplicity of random hypergraphs
Random hypergraphs extend the classical notion of random graphs by allowing hyperedges to join more than two vertices, making them well-suited for modeling higher-order interactions in complex systems. Despite their broad applicability, many structural properties of random hypergraphs remain less understood than in the graph setting. One such property is simplicity: the absence of self-loops, multi-hyperedges, and, in the hypergraph context, degenerate hyperedges where hyperedges contain a copy of the same vertex at least twice. While the behaviour of the number of such self-loops and multi-hyperedges is well understood for random graphs through the configuration model, analogous results for hypergraphs are comparatively sparse. In this work, we study both undirected and directed hypergraphs generated by the configuration model with prescribed vertex and hyperedge degrees. We derive exact, explicit expressions for the expected number of self-loops, multi-hyperedges and degenerate hyperedges, extending classical results from the graph setting. In addition, an asymptotical analysis shows that, under mild moment conditions on the degree distribution, the expected fraction of self-loops, multi-hyperedges and degenerate hyperedges vanishes as the number of vertices grows. Our results provide a systematic understanding of simplicity in directed and undirected hypergraph models.
Large fringe trees for random trees with given vertex degrees
Published
• View Publication
• BIB
This paper extends the study of fringe trees in random plane trees with a given degree statistic. While previous work established the asymptotic normality of the count of fringe trees isomorphic to a fixed tree, we investigate the case where the target tree grows with the size of the random tree.
We consider three primary subtree counts: the number of fringe trees isomorphic to a specific growing tree, the number of fringe trees sharing a given growing degree statistic, and the number of fringe trees of a specific growing size. To establish our results, we employ and compare four distinct probabilistic frameworks: the method of moments with the Gao-Wormald theorem, Stein's method with coupling (to provide explicit error bounds in total variation distance), the Cai-Devroye method, and Stein's method with exchangeable pairs. Our findings provide conditions for Poisson and normal convergence for these subtree counts.
Additionally, we provide a local limit theorem for sums of values obtained via sampling without replacement that may be of independent interest. Finally, our results and methods are also applied to conditioned critical Galton-Watson trees.
Non-existence probabilities and lower tails in the critical regime via Belief Propagation
We compute the logarithmic asymptotics of the non-existence probability (and more generally the lower-tail probability) for a wide variety of combinatorial problems for a range of parameters in the `critical regime' between the regime amenable to hypergraph container methods and that amenable to Janson's inequality. Examples include lower tails and non-existence probabilities for subgraphs of random graphs and for $k$-term arithmetic progressions in random sets of integers.
Our methods apply in the general framework of estimating the probability that a $p$-random subset of vertices in a $k$-uniform hypergraph induces significantly fewer hyperedges than expected. We show that under some simple structural conditions on the hypergraph and an upper bound on $p$ determined by a phase transition in the hard-core model on the infinite $k$-uniform, $Δ$-regular, linear hypertree, this probability can be accurately approximated by the Bethe free energy evaluated at the unique fixed point of a Belief Propagation operator on the hypergraph.
Equality in Fill's spectral gap problem
We study the adjacent-transposition chain on the symmetric group $\mathfrak{S}_n$ with a regular parameter vector $\vec{p} = (p_{i,j})_{i\neq j}$. Fill's spectral gap conjecture, recently resolved in the affirmative by Greaves-Zhu, states that among all regular parameter vectors, the spectral gap of the transition matrix is minimized by the uniform vector $p_{i,j}= 1/2$ for all $i\neq j$.
We prove the stronger statement that among all regular parameter vectors, the spectral gap is minimized if and only if $\vec{p}$ has a neutral label, i.e., there exists $c \in [n]$ such that $p_{c,i} = 1/2$ for all $i\neq c$. Moreover, in this case, we show that the multiplicity of the second largest eigenvalue is equal to the number of neutral labels, unless the number of neutral labels is $n-2$ or $n$, in which case the multiplicity is $n-1$. This confirms a conjecture of Fill.
The record statistic and forward stability of Schubert products
We initiate a probabilistic study of forward stability for products of Schubert polynomials through the record statistic (left-to-right maxima) of permutations. Building on the explicit record formula for forward stability obtained by Hardt and Wallach, we study random pairs of permutations drawn from three natural families: uniform permutations, Grassmannian permutations, and Boolean permutations. For each family, we determine record probabilities and use them to analyze the asymptotic behavior of forward stability. For uniform and Grassmannian permutations, we obtain asymptotics for the mean together with limiting distribution results. For Boolean permutations, we prove linear-order growth of the mean, and our analysis also produces an explicit time-inhomogeneous Markov chain that yields an exact linear-time uniform sampler. Beyond these cases, we prove that the record-set statistic is equidistributed on the avoidance classes of $132$ and $231$, and consequently the corresponding forward stability distributions coincide. We conclude with conjectures for numerous further permutation classes and a conjectural recursive criterion for when two avoidance classes have the same record-set distribution.
Semicircle laws with combined variance for non-uniform Erdős-Rényi hypergraphs
We consider Erdős-Rényi-type random hypergraphs that are non-uniform, in the sense that hyperedges of different sizes may coexist, and inhomogeneous, in that connection probabilities may depend on the hyperedge size. All parameters are allowed to scale with the hypergraph size. We study the random adjacency matrix whose $(u,v)$-entry counts the number of hyperedges containing both vertices $u$ and $v$, and characterize its expected limiting spectral distribution in terms of the connection probabilities and the hyperedge sizes. We provide a Pastur-type condition, in the sense of Chatterjee (2005), under which the matrix can be Gaussianized, as well as a more restrictive but simpler sufficient condition in terms of the generalized average degree of the model. As a second main result, based on such a Gaussianization, we characterize the limiting spectral distributions under non-sparse conditions as semicircle laws with an explicit parametric variance. The latter can be expressed as a convex combination of the variances arising in the uniform cases, with coefficients determined by the trade-off between the different sources of inhomogeneity.