math.PR ↗ arXiv
284 papers in this category
The number and structure of connected graphs with a fixed degree sequence
We study connected graphs with a fixed degree sequence, in the sparse setting where the number of edges grows linearly in the number of vertices. Using the relation to the configuration model, we identify the number of such connected graphs up to the exponential order. We do this by viewing a connected graph with a given degree distribution as the realization of the giant component in a larger configuration model, and carefully choosing the degree distribution of the larger graph so that it is likely that its giant component has the required degree distribution. To ensure that the connected graph has exactly the correct degrees, we use a switching argument. Additionally, we obtain results on rare event probabilities and describe the local structure of a uniform connected graph with a fixed degree sequence.
Algorithmic Phase Transition for Large Independent Sets in Dense Hypergraphs
We study the algorithmic tractability of finding large independent sets in dense random hypergraphs. In the sparse regime, much of the natural algorithms can be formulated within either the local or the low-degree polynomial (LDP) framework, and a rich literature has subsequently identified nearly sharp algorithmic thresholds within these classes by exploiting their stability. In the dense setting, however, the algorithmic paradigms are fundamentally different: they are online and thus need not be stable. Perhaps more crucially, even for the classical Erdős-Rényi random graph $G(n,p)$, LDPs are conjectured to fail in the 'easy' regime accessible to online algorithms, thereby challenging their viability for dense models.
Our focus is on two models: (i) finding large independent sets in dense $r$-uniform Erdős-Rényi hypergraphs, and (ii) the more challenging problem of finding large $γ$-balanced independent sets in dense $r$-uniform $r$-partite hypergraphs, where the $i$-th coordinate of $γ\in\mathbb{Q}^r$ specifies the proportion of vertices from $V_i$ in the independent set. For both models, we pinpoint the size of the largest independent set and design online algorithms that achieve a multiplicative approximation factor of $r^{1/(r-1)}$ in the uniform and $(\max_i γ_i)^{-1/(r-1)}$ in the $r$-partite model. Furthermore, we establish matching algorithmic lower bounds, showing that these computational gaps are sharp: no online algorithms can breach these gaps.
Leap generators for composition schemes
Leap generators have been introduced in [Duchon et al.'04] for exact-size random generation of structures in a class of the form $\mathcal{C}=\mathrm{Seq}(\mathcal{B})$ (sequence construction), in the supercritical case. We extend these generators to supercritical composition schemes $\mathcal{C}=\mathcal{A}\circ\mathcal{B}$. Compared to the sequence construction, the obtained exact-size random generator for $\mathcal{C}$ still has linear time complexity (under conditions on the sampling complexity in $\mathcal{A}$ and $\mathcal{B}$), but perfect uniformity of the distribution is lost in general. However the distribution on $\mathcal{C}_n$, called leap distribution, is asymptotically uniform, the total variation distance from the uniform distribution being $(c+o(1))n^{-1/2}$ for an explicit constant $c$. These generators are simple to implement and can be applied to several classes of walks and trees, in particular Pólya trees. Leap generators can also be given for certain critical composition schemes, those relating planar map families, where this time the total variation distance to the uniform distribution is $\sim c\,n^{-1/3}$ for an explicit constant $c$.
Small values of signed harmonic sums and logarithmic means of multiplicative functions
We construct sequences $\{a_n\}_{n\in\mathbb{N}}\in\{-1,1\}^{\mathbb{N}}$ with small values of signed harmonic sums \[ \sum_{n\in\mathcal{A}\cap[1,N]}\frac{a_n}{n}, \] for any reasonably dense subsets $\mathcal{A}\subset\mathbb{N}.$ We apply these methods to further construct completely multiplicative functions $f:\mathbb{N}\to\{-1,1\}$ with unusually small logarithmic partial sums, that is, \[ \sum_{n \leq N}\frac{f(n)}{n} \ll \exp\left(-c_0 \frac{N^{1/3}}{(\log N)^{1/3}} \right) \] holds for infinitely many $N\to\infty$. The proofs combine careful analysis of the small-scale distribution of random harmonic sums over subsets of $\mathbb{N}$, together with deterministic inductive arguments inspired by the ``anatomy" of integers.
Almost-Orthogonality in Lp Spaces: A Case Study with Grok
Carbery proposed the following sharpened form of triangle inequality for many functions: for any $p\ge 2$ and any finite sequence $(f_j)_j\subset L^p$ we have \[ \Big\|\sum_j f_j\Big\|_p \ \le\ \left(\sup_{j} \sum_{k} α_{jk}^{\,c}\right)^{1/p'} \Big(\sum_j \|f_j\|_p^p\Big)^{1/p}, \] where $c=2$, $1/p+1/p'=1$, and $α_{jk}=\sqrt{\frac{\|f_{j}f_{k}\|_{p/2}}{\|f_{j}\|_{p}\|f_{k}\|_{p}}}$. In the first part of this paper we construct a counterexample showing that this inequality fails for every $p>2$. We then prove that if an estimate of the above form holds, the exponent must satisfy $c\le p'$. Finally, at the critical exponent $c=p'$, we establish the inequality for all integer values $p\ge 2$.
In the second part of the paper we obtain a sharp three-function bound \[ \Big\|\sum_{j=1}^{3} f_j\Big\|_p \ \le\ \left(1+2Γ^{c(p)}\right)^{1/p'} \Big(\sum_{j=1}^{3} \|f_j\|_p^p\Big)^{1/p}, \] where $p \geq 3$, $c(p) = \frac{2\ln(2)}{(p-2)\ln(3)+2\ln(2)}$ and $Γ=Γ(f_1,f_2,f_3)\in[0,1]$ quantifies the degree of orthogonality among $f_1,f_2,f_3$. The exponent $c(p)$ is optimal, and improves upon the power $r(p) = \frac{6}{5p-4}$ obtained previously by Carlen, Frank, and Lieb. Some intermediate lemmas and inequalities appearing in this work were explored with the assistance of the large language model Grok.
Optimal Union Probability Interval Is NP-Hard
A problem dating back to Boole [Laws of Thought, Walton & Maberly,1854] is what can be computed about the probability of a finite union of events when given as input the probabilities of intersections of some of the events. The modern geometric study of the problem can be traced back to Hailperin [Amer. Math. Monthly 2 (1965) 343--359] who phrased the problem in the language of linear programming and generalized it to logical formulas of the events other than disjunction, heralding a substantial body of work in probabilistic logic [Nilsson, Artif.\ Intell.\ 28 (1986) 71--87], including the probabilistic satisfiability problem of Georgakopoulos, Kavvadis, and Papadimitriou [J.Complexity 4 (1988) 1--11], as well as fundamental connections to the geometry of metrics via cut and correlation polytopes [Deza and Laurent, Geometry of Cuts and Metrics, Springer, 1997] and to the study of marginal polytopes in graphical models of machine learning [Wainwright and Jordan, Found.\ Trends Mach.\ Learn. 1 (2008) 1--305]. This paper (i) describes the pertinent geometry of Boole's problem via coordinate projections of an elementary polytope arising essentially from Hailperin's linear program on the atoms of a Venn diagram, and (ii) shows that computing the optimal interval for the union probability is NP-hard, resolving an apparent gap in the literature highlighted by Pitowsky [Math.\ Programming 50 (1991) 395--414] and Boros et al. [Math.\ Oper.\ Res. 39 (2014) 1311--1329 and 51 (2026) 134--148].
On the partition function of a class of Mallows model
Let $\Sym{n}$ denote the set of all permutations on $n$ labels. Let $c:[0, 1]^2\to [0, \infty)$ be a twice continuously differentiable function. A subfamily of the Mallows model is the Gibbs probability measures on $\Sym{n}$ such that $\mathbb{P}(X=σ)=L_n^{-1} \prod_{i=1}^{n}\exp(-c(i/n, σ(i)/n))$. Mukherjee [Ann. Stat., Vol. 44(2), pp 853--875 (2016)] computed the limit of the log partition function and showed that $\lim_{n\to \infty}\frac{1}{n}\log L_n=-Γ_0$ where $Γ_0$ is the optimal cost associated with an entropy regularized optimal transport problem. In the KRP Memorial Volume of the Indian Journal of Pure and Applied Math, Pal conjectured an exact value for the limit $\lim_{n\to \infty} e^{-nΓ_0}L_n$ in terms of the Fredholm determinant of an integral operator and provided a partial proof. We give a complete proof of Pal's conjecture.
Optimal Hardness of Online Algorithms for Large Common Induced Subgraphs
We study the problem of efficiently finding large common induced subgraphs of two independent Erdős--Rényi random graphs $G_1, G_2 \sim \mathbb{G}(n,1/2)$. Recently, Chatterjee and Diaconis showed that the largest common induced subgraph of $G_1$ and $G_2$ has size $(4-o(1))\log_2 n$ with high probability. We first show that a simple greedy online algorithm finds a common induced subgraph of $G_1$ and $G_2$ of size $(2-o(1)) \log_2 n$ with high probability. Our main result shows that no online algorithm can find a common induced subgraph of $G_1$ and $G_2$ of size at least $(2+\varepsilon) \log_2 n$ with probability bounded away from $0$ as $n \to \infty$. Together, these results provide evidence that this problem exhibits a computation-to-optimization gap. To prove the impossibility result, we show that the solution space of the problem exhibits a version of the (multi) overlap gap property (OGP), and utilize an interpolation argument recently developed by Gamarnik, Kizildağ, and Warnke that connects OGP and online algorithms.
Dimer models on astroidal zig-zag graphs
On a finite weighted graph, the dimer model is a probability measure on its dimer covers, that assigns to any cover a probability proportional to the product of the weights of its edges. For planar bipartite graphs, dimer correlations are encoded by the inverse of the so-called Kasteleyn matrix; for a large graph, typically taken as a finite domain in a periodic graph, this inverse matrix is known explicitly only for a handful of examples. In all previously known examples, the Newton polygon -- a convex lattice polygon that classifies periodic graphs up to local moves -- is either a triangle or a quadrilateral.
Our main results are the following.
For any (minimal) periodic planar bipartite graph, we construct an $(n-3)$-dimensional family of finite subgraphs for which we obtain an explicit inverse Kasteleyn matrix; here $n$ is the number of sides of the Newton polygon. Their boundaries are formed by zig-zag paths and their overall shape is reminiscent of an astroid; we call them astroidal zig-zag graphs (AZ graphs). If the Newton polygon is the unit square then the corresponding AZ graph is the celebrated Aztec diamond with its size as the parameter.
Our inverse Kasteleyn matrices are given by a double contour integral on the corresponding spectral curve for any Fock weighting of the graph. This includes, in particular, all periodic weightings.
For periodic weightings, we asymptotically analyze the resulting inverse Kasteleyn matrices. We establish a phase separation in large AZ graphs into asymptotically frozen, rough (liquid), and smooth (gaseous) regions, and obtain an explicit parametrization of the `arctic curve'.
We also compute the deterministic limit of the height function, known as the limit shape, and prove the convergence of the local dimer correlations to the translation-invariant Gibbs measure of the slope predicted by the limit shape.
On the approximation of permutons
We study the optimal rectangular-discrepancy approximation of permutons by finite permutations. We transfer bounds from discrepancy theory to this more restricted setup. Moreover, we show that superlinear approximation can occur only for permutons supported by graphs of measure-preserving functions, and demonstrate how the local regularity of this function obstructs approximability. We also consider the biased Brownian separable permuton and prove a lower bound on its approximation error by showing that its supporting measure-preserving function has Lipschitz points almost surely.
Properties of tensorial free cumulants
In the past two years, several points of view have been proposed to address the question of the generalization of the theory of free probability to random tensors with different invariances, and it is unclear at this point whether they lead to the same notions of tensorial free cumulants and freeness. One way to approach this problem, developed by Collins, Gurau and the second named author for local unitary invariant random tensors, relies on finite size quantities involving averages over the invariance group, and whose asymptotics naturally possess the properties expected for tensorial generalizations of free cumulants of arbitrary orders. At this point, this approach has only been carried out for certain distributions, and for a subset of the moments that define such theories, and a more systematic and exhaustive study is lacking.
This is the program initiated in this paper: we link this approach to the one proposed by Nechita and Park; extend a number of their results as well as those of the aforementioned paper to arbitrary orders of fluctuations, thereby generalizing higher order free cumulants; push further the study of distributions with larger invariance groups; detail the link with the asymptotics of the free-energies of the tensor HCIZ and BGW integrals; and provide formulae for tensorial free cumulants of products of tensors.
Another important question is that of the definition of concrete distributions whose tensorial free-cumulants take non-trivial values. We compute the tensorial free cumulants for Gaussian random tensors with non-trivial covariances, and show that they provide such examples.
Primitive sets and von Mangoldt chains: Erdős Problem #1196 and beyond
A set of integers is primitive if no number in the set divides another. We introduce a new method for bounding Erdős sums of primitive sets, suggested from output of GPT-5.4 Pro, based on Markov chains with von Mangoldt weights. The method leads to a host of applications, yet seems to have been overlooked by the prior literature since Erdős's seminal 1935 paper.
As applications, we prove two 1966 conjectures of Erdős-Sárközy-Szemerédi, on primitive sets of large numbers (#1196) and on divisibility chains (#1217). The method also provides a short proof of the Erdős Primitive Set Conjecture (#164), as well as the related claim that 2 is an ''Erdős-strong'' prime. Moreover, the method resolves a revised form of the Banks-Martin conjecture, which has long been viewed as a unifying `master theorem' for the area.
Fibonacci numbers and the probability of polygon formation using random length sticks
We present two complementary proofs that, if the lengths of $n$ sticks are sampled at random, then the probability that no $p+1$ sticks can form a $(p+1)$-sided polygon can be expressed as the product of the reciprocals of a series of terms involving the $p$-step Fibonacci numbers. The first proof uses matrix algebra to extend the method previously used by Sudbury et al. to derive expressions for the probabilities of not being able to form triangles and quadrilaterals. The second alternative proof uses a different approach based on expressions for the minimum and maximum lengths of each stick that are compatible with the constraint of not being able to form a $(p+1)$-sided polygon, and provides insights into the structure of the probability expressions and the underlying reason that they include the Fibonacci numbers. Furthermore, the approach is developed in a generalised way that can, in principle, be applied to sticks randomly sampled from any probability distribution.
Gårding Polynomials
We introduce Gårding polynomials, a class of real multivariate polynomials defined via positivity regions invariant under translation by positive directions and closed under strictly positive affine transformations. We establish a structural theorem providing two complementary characterizations of this class: one via reduction to the multi-affine case through polarization, and another via a recursive condition involving partial derivatives. The class of Gårding polynomials strictly extends that of real stable polynomials while retaining many of their structural properties. In particular, multi-affine Gårding polynomials with nonnegative coefficients satisfy the Rayleigh property, and their positive univariate specializations yield ultra log-concave coefficient sequences. Moreover, the Gårding property for several matroid generating functions is preserved under natural matroid operations. As applications, we obtain new negative dependence results for generating functions associated with various classes of matroids and graphs--many of which lie beyond the reach of real stability or Lorentzian methods--as well as for characteristic polynomials of certain matrix classes.
The proportion of permutations fixing a $k$-set
Denote by $p(k)$ the limit, as $n \rightarrow \infty$, of the probability that a random permutation on a set of size $n$ has an invariant set of size $k$. We give an asymptotic formula for $p(k)$, showing that it is asymptotically $f(\{\log_2 k\}) k^{-δ} (\log k)^{-3/2}$ where $δ= 1 - \frac{1 + \log \log 2}{\log 2} \approx 0.086$ and $f$ is a smooth, positive, function on $\mathbb{R}/\mathbb{Z}$, which we will describe explicitly. The function $f$ satisfies $\frac{\max f}{\min f} < 1 + 2 \times 10^{-7}$ and we conjecture that it is not constant.
Estimating $p(k)$ is a model for the more well-known question which asks for an estimation of $M(n)$, the number of distinct elements in the $n$-by-$n$ multiplication table. By elaborating on the techniques in this paper, we will give an asymptotic for $M(n)$ in forthcoming work.
The Expiring Coupon Collector: Sliding-Window Surjection Flux and Rare-Entry Laws
We study the coupon collector with deterministic expiration: one coupon is drawn at each time, and each coupon remains active for exactly $M$ draws. Completion occurs when all $n$ coupon types are simultaneously active. Equivalently, the current length-$M$ sliding window of draws must contain all $n$ types.
The central object is not the one-time probability that a random window is onto, but the stationary flux of new entries into the onto-window set. We compute this flux exactly: \[
μ_{n,M}
=\Pbb(W_{t-1}\text{ is not onto},\ W_t\text{ is onto})
=\frac{(n-1)(n-1)!S(M-1,n-1)}{n^M}, \] where $S(\cdot,\cdot)$ denotes a Stirling number of the second kind. Under a quantitative subcritical separation condition, satisfied in particular by every fixed integer scale $M=\floor{αn\log n}$, $0<α<1$, we prove local declumping and obtain \[
μ_{n,M}T_{n,M}\Rightarrow \Exp(1). \] For the fixed subcritical scale $M=\floor{αn\log n}$, $0<α<1$, this gives the logarithmic scale \[
\log T_{n,M}=n^{1-α}+o_{\mathbb P}(n^{1-α}),
\qquad
\log \Ebb T_{n,M}=n^{1-α}+o(n^{1-α}), \] and, when $α>1/2$, the sharper normalization \[
n^{-α}e^{-n^{1-α}}T_{n,M}\Rightarrow \Exp(1),
\qquad
\Ebb T_{n,M}\sim n^αe^{n^{1-α}}. \] Thus the leading scale proposed in the Math StackExchange discussion is made rigorous; the exact finite-$n$ flux gives the canonical normalization throughout the subcritical range. The result is a sliding-window companion to rare-void entry-flux methods for nonmonotone coupon collectors.
Terminal Defects, Growing Multiplicity, and Variance Extremality in the Double Dixie Cup Problem
We develop a terminal-defect method for the double Dixie cup problem and use it to prove the finite-variance extremality conjecture of Doumas and Papanicolaou. For every \(m\ge1\) and \(N\ge2\), among all positive coupon probability vectors \(p=(p_1,\ldots,p_N)\), the variance of the time \(T_m(N)\) to collect \(m\) complete sets is uniquely minimized at the uniform vector. We prove the stronger radial statement that the variance is strictly increasing along every ray from the uniform vector. The proof is finite-\(N\) and exact: after Poissonization, the completion time is a maximum of independent Erlang variables, and the radial derivative of its distribution is compared to a size-biased law using a monotone-likelihood-ratio argument based on a log-scale monotonicity property of the Gamma reverse hazard.
The same framework gives a growing-multiplicity Gumbel theorem in the equal-probability case, with expectation and variance asymptotics on the inverse gamma-tail scale. This recovers the fixed-\(m\) equal-probability variance asymptotic stated as Conjecture 1 by Doumas and Papanicolaou, classically known for \(m=1\), and extends the mechanism to \(m=m_N\). We also illustrate the unequal-probability theory with endpoint-Laplace limits for power-law probabilities.
Counterexamples to an Extremal Conjecture for Random Cycle-Factors
Christoph, Draganić, Girão, Hurley, Michel, and Müyesser conjectured that, when $d\mid n$, the expected number of cycles in a uniformly random cycle-factor of a directed $d$-regular graph on $n$ vertices is uniquely maximised by the disjoint union of $n/d$ copies of the complete looped digraph $K_d^\circ$, with value $(n/d)H_d$ [FOCS 2025]. We disprove this conjecture in the strongest possible range. For every $d\ge 3$ and every multiple $n=kd$ with $k\ge 2$, we construct a directed $d$-regular graph on $n$ vertices whose uniformly random cycle-factor has expected cycle count strictly larger than $kH_d$. We also show that the conjectured extremal picture is correct in degree $d=2$, giving a sharp dichotomy between degree two and all higher degrees.
Exact Closed-Form Formulae for Linear and Circular Continuous Scan Statistics: $P_c(N - 1; N, w)$, $P_c(3; N, w)$, and $P(3; N, w)$
The continuous linear $P(k; N, w)$ and circular scan statistics $P_c(k; N, w)$ are fundamental tools in probability and spatial statistics, frequently used to detect clustering in uniform data. Let $X_1, X_2, \dots, X_N$ be independently and uniformly distributed random variables on a unit interval or unit ring. The exact distribution of these scan statistics relies on the minimum window width required to capture exactly $k$ points. Furthermore, the survival function $1 - P_c(k; N, w)$ directly corresponds to the geometric probability that if $N$ arcs of length $1 - w$ are uniformly and randomly placed on a unit circle, every point on the circle is covered at least $N + 1 - k$ times. Historically, evaluating the exact cumulative distribution functions, $P(k; N, w)$ and $P_c(k; N, w)$, relies heavily on complex recursive approximations. In this paper, we bypass these traditional recursive methods to derive direct, generalized closed-form expressions for some linear and circular continuous scan statistics. Specifically, we present the exact analytical solutions for $P_c(N - 1; N, w)$, $P_c(3; N, w)$, and $P(3; N, w)$ for arbitrary values of $N$ and window width $w$. These newly derived closed-form expressions not only provide exact baseline distributions for extreme spacings but also significantly simplify computational complexity compared to existing iterative approaches.
What is The Probability That A Random Graph With A Given Degree Sequence is Connected?
An $n$-tuple $D=(d(1),\dots,d(n))$ is a \emph{feasible degree sequence} if there is a graph on $\{1,\dots,n\}$ such that $i$ has degree $d(i)$. Any such graph will have $m=\sum_{i=1}^n d(i)/2$ edges. Letting $G(D)$ be a graph chosen uniformly from those with the given degree sequence, we upper-bound the probability that $G(D)$ is disconnected based on the number of vertices of degree $d$ for small $d$, and develop a powerful tool for proving such bounds. If there are any vertices of degree zero the probability $G$ is disconnected is $1$, so we assume there are no such vertices. Our results then imply that if there are $o(\sqrt{m})$ vertices of degree $1$ and $o(m)$ vertices of degree 2 then with high probability $G$ is connected, while if there are no vertices of degree 1 or 2 then the probability $G$ is disconnected is $O(\frac{n^4}{m^6})$.