arXiv++ Combinatorics

Browse math.CO papers from arXiv

gaussian distribution

38 papers tagged with this keyword
2026-03-11
An asymptotically optimal bound for the concentration function of a sum of independent integer random variables
For a random variable $X$ define $Q(X) = \sup_{x \in \mathbb{R}} \mathbb{P}(X=x)$. Let $X_1, \dots, X_n$ be independent integer random variables. Suppose $Q(X_i) \le α_i \in (0,1]$ for each $i \in \{1, \dots, n\}$. Juškevičius (2023) conjectured that $Q(X_1 + \dots +X_n) \le Q(Y_1 + \dots+ Y_n)$ where $Y_1, \dots, Y_n$ are independent and $Y_i$ is a random integer variable with $Q(Y_i) =α_i$ that has the smallest variance, i.e. the distribution of $Y_i$ has probabilities $α_i, \dots, α_i, β_i$ or probabilities $β_i, α_i, \dots, α_i$ on some interval of integers, where $0 \le β_i < α_i$. We prove this conjecture asymptotically: i.e., we show that for each $δ> 0$ there is $V_0 = V_0(δ)$ such that if ${\mathrm Var} (\sum Y_i) \ge V_0$ then $Q(\sum X_i) \le (1+δ) Q(\sum Y_i)$. This implies an analogous asymptotically optimal inequality for concentration at a point when $X_1$, $\dots$, $X_n$ take values in a separable Hilbert space. Our long and technical argument relies on several non-trivial previous results including an inverse Littlewood--Offord theorem and an approximation in total variation distance of sums of multivariate lattice random vectors by a discretized Gaussian distribution.
2026-02-08
Permanents of matrix ensembles: computation, distribution, and geometry
We report on a computational and experimental study of permanents. On the computational side, we use the GPU to greaatly accelerate the computation of permanents over $\mathbb{C},$ $\mathbb{R},$ $\mathbb{F}_p$ and $\mathbb{Q}.$ In particular, we use this to compute the permanents of DFT and Schur matrices far beyond the ranges hitherto known. On the experimental side, we present two new observations. First, for Haar-distributed unitary matrices~$U$, the permanent $\perm(U)$ follows a circularly-symmetric complex Gaussian distribution $\mathcal{CN}(0,σ^2)$ -- we confirm this via a number of tests for $n$ up to~23 with $50{,}000$ samples. The DFT matrix permanent is an extreme outlier for every prime $n\ge 7$. In contrast, for Haar-random \emph{orthogonal} matrices~$O$, the permanent $\perm(O)$ is approximately real Gaussian but with positive excess kurtosis that decays as~$O(1/n)$, indicating slower convergence. For matrices with Gaussian entries (GUE, GOE, Ginibre), the permanent follows an $α$-stable distribution with stability index $α\approx 1.0$--$1.4$, well below the Gaussian value $α=2$. Secondly, we study the permanent along geodesics on the unitary group. For the geodesic from the identity to the $n$-cycle permutation matrix, we find a universal scaling function $f(t)=\frac{1}{n}\ln|\perm(γ(t))|$ that is independent of~$n$ in the large-$n$ limit, with a midpoint value \[ \perm(γ({\textstyle\frac12})) = (-1)^{(n-1)/2}\cdot 2e^{-n}\bigl(1+\tfrac{1}{3n}+O(n^{-2})\bigr) \] for odd~$n$ and zero for even~$n$. For the geodesic to the DFT matrix, the permanent recovers $10$--$40$ times above its valley minimum when $n$ is prime, but not when $n$ is composite -- a geodesic fingerprint of primality.
2025-10-16
The asymptotic number of equivalence classes of linear codes with given dimension
We investigate the asymptotic number of equivalence classes of linear codes with prescribed length and dimension. While the total number of inequivalent codes of a given length has been studied previously, the case where the dimension varies as a function of the length has not yet been considered. We derive explicit asymptotic formulas for the number of equivalence classes under three standard notions of equivalence, for a fixed alphabet size and increasing length. Our approach also yields an exact asymptotic expression for the sum of all q-binomial coefficients, which is of independent interest and answers an open question in this context. Finally, we establish a natural connection between these asymptotic quantities and certain discrete Gaussian distributions arising from Brownian motion, providing a probabilistic interpretation of our results.
2025-10-16
Central limit theorem for the sine-$β$ point process at $β\le 2$
The purpose of this paper is to establish the analogue of the Soshnikov Central Limit Theorem for the sine-$β$ process at $β\le 2$. We consider regularized additive functionals, which correspond to 1-Sobolev regular functions $f(x/R)$, in the limit $R\to\infty$. Their convergence to the Gaussian distribution with respect to the Kolmogorov-Smirnov metric at the rate $(\ln R)^{-1/2}$ is established. The proof is based on the convergence of the circular $β$-ensemble to the sine-$β$ process, which was shown by Killip and Stoiciu. We find a convenient bound for the Laplace transforms of additive functionals under the circular $β$-ensemble, which holds under the scaling limit, suggested by Killip and Stoiciu. Further, we show that the additive functionals under the circular $β$-ensemble with $n$ particles converge to the gaussian distribution as $n\to\infty$ for all $1/2$-Sobolev regular functions for $β\le 2$, as was conjectured by Lambert. Finally, in order to prove the limit theorem for the circular $β$-ensemble we derive the connection between expectations of multiplicative functionals and the Jack measures, which generalizes the connection between the circular unitary ensemble and the Schur measures given by Gessel's theorem.
2025-02-20 v3
Coxeter codes: Extending the Reed-Muller family
Binary Reed-Muller (RM) codes are defined via evaluations of Boolean-valued functions on $\mathbb{Z}_2^m$. We introduce a class of binary linear codes that generalizes the RM family by replacing the domain $\mathbb{Z}_2^m$ with an arbitrary finite Coxeter group. Like RM codes, this class is closed under duality, forms a nested code sequence, satisfies a multiplication property, and has asymptotic rate determined by a Gaussian distribution. Coxeter codes also give rise to a family of quantum codes for which transversal diagonal $Z$ rotations can perform non-trivial logic.
2024-12-07
Hyperbolicity, slimness, and minsize, on average
A metric space $(X,d)$ is said to be $δ$-hyperbolic if $d(x,y)+d(z,w)$ is at most $\max(d(x,z)+d(y,w), d(x,w)+d(y,z))$ by $2 δ$. A geodesic space is $δ$-slim if every geodesic triangle $Δ(x,y,z)$ is $δ$-slim. It is well-established that the notions of $δ$-slimness, $δ$-hyperbolicity, $δ$-thinness and similar concepts are equivalent up to a constant factor. In this paper, we investigate these properties under an average-case framework and reveal a surprising discrepancy: while $\mathbb{E}δ$-slimness implies $\mathbb{E}δ$-hyperbolicity, the converse does not hold. Furthermore, similar asymmetries emerge for other definitions when comparing average-case and worst-case formulations of hyperbolicity. We exploit these differences to analyze the random Gaussian distribution in Euclidean space, random $d$-regular graph, and the random Erdős-Rényi graph model, illustrating the implications of these average-case deviations.
2024-09-06
Cumulants in rectangular finite free probability and beta-deformed singular values
Motivated by the $(q,γ)$-cumulants, introduced by Xu [arXiv:2303.13812] to study $β$-deformed singular values of random matrices, we define the $(n,d)$-rectangular cumulants for polynomials of degree $d$ and prove several moment-cumulant formulas by elementary algebraic manipulations; the proof naturally leads to quantum analogues of the formulas. We further show that the $(n,d)$-rectangular cumulants linearize the $(n,d)$-rectangular convolution from Finite Free Probability and that they converge to the $q$-rectangular free cumulants from Free Probability in the regime where $d\to\infty$, $1+n/d\to q\in[1,\infty)$. As an application, we employ our formulas to study limits of symmetric empirical root distributions of sequences of polynomials with nonnegative roots. One of our results is akin to a theorem of Kabluchko [arXiv:2203.05533] and shows that applying the operator $\exp(-\frac{s^2}{n}x^{-n}D_xx^{n+1}D_x)$, where $s>0$, asymptotically amounts to taking the rectangular free convolution with the rectangular Gaussian distribution of variance $qs^2/(q-1)$.
2024-07-04
Cumulants of threshold for Schensted row insertion into random tableaux
Schensted row insertion is a fundamental component of the Robinson-Schensted-Knuth (RSK) algorithm, a powerful tool in combinatorics and representation theory. This study examines the insertion of a deterministic number into a random tableau of a specified shape, focusing on the relationship between the value of the inserted number and the position of the new box created by the Schensted row insertion. Specifically, for a given tableau and a point on its boundary, we consider the threshold that separates values which, if inserted, would result in the new box being created above the point from those that would result in a new box below. We analyze a random tableau of fixed shape and study the corresponding random threshold value. Explicit combinatorial formulas for the cumulants of this random variable are provided, expressed in terms of Kerov's transition measure of the diagram. These combinatorial formulas involve summing over non-crossing alternating trees. As a first application of these results, we demonstrate that for random Young tableaux of prescribed large shape, the rightmost entry in the first row converges in distribution to an explicit Gaussian distribution.
2024-03-11
Untangling Gaussian Mixtures
Tangles were originally introduced as a concept to formalize regions of high connectivity in graphs. In recent years, they have also been discovered as a link between structural graph theory and data science: when interpreting similarity in data sets as connectivity between points, finding clusters in the data essentially amounts to finding tangles in the underlying graphs. This paper further explores the potential of tangles in data sets as a means for a formal study of clusters. Real-world data often follow a normal distribution. Accounting for this, we develop a quantitative theory of tangles in data sets drawn from Gaussian mixtures. To this end, we equip the data with a graph structure that models similarity between the points and allows us to apply tangle theory to the data. We provide explicit conditions under which tangles associated with the marginal Gaussian distributions exist asymptotically almost surely. This can be considered as a sufficient formal criterion for the separabability of clusters in the data.
2023-03-27
Limits of polyhedral multinomial distributions
We consider limits of certain measures supported on lattice points in lattice polyhedra defined as the intersection of half-spaces $\{m\in\mathbb{R}^n|\langle v_i,x\rangle+a_i \geq 0\}$, where $\sum_i v_i = 0$. The measures are densities associated to lattice random variables obtained by restriction of multinomial random variables. We find the limiting Gaussian distributions explicitly.
2022-08-10 v2
Computing the theta function
Let $f: {\Bbb R}^n \longrightarrow {\Bbb R}$ be a positive definite quadratic form and let $y \in {\Bbb R}^n$ be a point. We present a fully polynomial randomized approximation scheme (FPRAS) for computing $\sum_{x \in {\Bbb Z}^n} e^{-f(x)}$, provided the eigenvalues of $f$ lie in the interval roughly between $s$ and $e^{s}$ and for computing $\sum_{x \in {\Bbb Z}^n} e^{-f(x-y)}$, provided the eigenvalues of $f$ lie in the interval roughly between $e^{-s}$ and $s^{-1}$ for some $s \geq 3$. To compute the first sum, we represent it as the integral of an explicit log-concave function on ${\Bbb R}^n$, and to compute the second sum, we use the reciprocity relation for theta functions. We then apply our results to test the existence of many short integer vectors in a given subspace $L \subset {\Bbb R}^n$, to estimate the distance from a given point to a lattice, and to sample a random lattice point from the discrete Gaussian distribution.
Spectral large deviations of sparse random matrices
Published • View PublicationBIB
Eigenvalues of Wigner matrices has been a major topic of investigation. A particularly important subclass of such random matrices is formed by the adjacency matrix of an Erdős-Rényi graph $\mathcal{G}_{n,p}$ equipped with i.i.d. edge-weights. An observable of particular interest is the largest eigenvalue. In this paper, we study the large deviations behavior of the largest eigenvalue of such matrices, a topic that has received considerable attention over the years. We focus on the case $p = \frac{d}{n}$, where most known techniques break down. So far, results were known only for $\mathcal{G}_{n,\frac{d}{n}}$ without edge-weights (Krivelevich and Sudakov, '03), (Bhattacharya, Bhattacharya, and Ganguly, '21) and with Gaussian edge-weights (Ganguly and Nam, '21). In the present article, we consider the effect of general weight distributions. More specifically, we consider the entries whose tail probabilities decay at rate $e^{-t^α}$ with $α>0$, where the regimes $0<α<2$ and $α>2$ correspond to tails heavier and lighter than the Gaussian tail respectively. While in many natural settings the large deviations behavior is expected to depend crucially on the entry distribution, we establish a surprising and rare universal behavior showing that this is not the case when $α> 2.$ In contrast, in the $α< 2$ case, the large deviation rate function is no longer universal and is given by the solution to a variational problem, the description of which involves a generalization of the Motzkin-Straus theorem, a classical result from spectral graph theory. As a byproduct of our large deviation results, we also establish new law of large numbers results for the largest eigenvalue. In particular, we show that the typical value of the largest eigenvalue exhibits a phase transition at $α= 2$, i.e. the Gaussian distribution.
2022-01-27 v2
Cycles in Mallows random permutations
Published • View PublicationBIB
We study cycle counts in permutations of $1,\dots,n$ drawn at random according to the Mallows distribution. Under this distribution, each permutation $π\in S_n$ is selected with probability proportional to $q^{\text{inv}(π)}$, where $q>0$ is a parameter and $\text{inv}(π)$ denotes the number of inversions of $π$. For $\ell$ fixed, we study the vector $(C_1(Π_n),\dots,C_\ell(Π_n))$ where $C_i(π)$ denotes the number of cycles of length $i$ in $π$ and $Π_n$ is sampled according to the Mallows distribution. Here we show that if $0<q<1$ is fixed and $n\to\infty$ then there are positive constants $m_i$ such that each $C_i(Π_n)$ has mean $(1+o(1)) \cdot m_i\cdot n$ and the vector of cycle counts can be suitably rescaled to tend to a joint Gaussian distribution. Our results also show that when $q>1$ there is striking difference between the behaviour of the even and the odd cycles. The even cycle counts still have linear means, and when properly rescaled tend to a multivariate Gaussian distribution. For the odd cycle counts on the other hand, the limiting behaviour depends on the parity of $n$ when $q>1$. Both $(C_1(Π_{2n}),C_3(Π_{2n}),\dots)$ and $(C_1(Π_{2n+1}),C_3(Π_{2n+1}),\dots)$ have discrete limiting distributions -- they do not need to be renormalized -- but the two limiting distributions are distinct for all $q>1$. We describe these limiting distributions in terms of Gnedin and Olshanski's bi-infinite extension of the Mallows model. We also investigate these limiting distributions, and study the behaviour of the constants involved in the Gaussian limit laws. We for example show that as $q\downarrow 1$ the expected number of 1-cycles tends to $1/2$ -- which, curiously, differs from the value corresponding to $q=1$. In addition we exhibit an interesting "oscillating" behaviour in the limiting probability measures for $q>1$ and $n$ odd versus $n$ even.
2021-12-22 v2
Plücker Coordinates of the best-fit Stiefel Tropical Linear Space to a Mixture of Gaussian Distributions
Published • View PublicationBIB
In this research, we investigate a tropical principal component analysis (PCA) as a best-fit Stiefel tropical linear space to a given sample over the tropical projective torus for its dimensionality reduction and visualization. Especially, we characterize the best-fit Stiefel tropical linear space to a sample generated from a mixture of Gaussian distributions as the variances of the Gaussians go to zero. For a single Gaussian distribution, we show that the sum of residuals in terms of the tropical metric with the max-plus algebra over a given sample to a fitted Stiefel tropical linear space converges to zero by giving an upper bound for its convergence rate. Meanwhile, for a mixtures of Gaussian distribution, we show that the best-fit tropical linear space can be determined uniquely when we send variances to zero. We briefly consider the best-fit topical polynomial as an extension for the mixture of more than two Gaussians over the tropical projective space of dimension three. We show some geometric properties of these tropical linear spaces and polynomials.
Tropical Support Vector Machines: Evaluations and Extension to Function Spaces
Published • View PublicationBIB
Support Vector Machines (SVMs) are one of the most popular supervised learning models to classify using a hyperplane in an Euclidean space. Similar to SVMs, tropical SVMs classify data points using a tropical hyperplane under the tropical metric with the max-plus algebra. In this paper, first we show generalization error bounds of tropical SVMs over the tropical projective torus. While the generalization error bounds attained via Vapnik-Chervonenkis (VC) dimensions in a distribution-free manner still depend on the dimension, we also show numerically and theoretically by extreme value statistics that the tropical SVMs for classifying data points from two Gaussian distributions as well as empirical data sets of different neuron types are fairly robust against the curse of dimensionality. Extreme value statistics also underlie the anomalous scaling behaviors of the tropical distance between random vectors with additional noise dimensions. Finally, we define tropical SVMs over a function space with the tropical metric.
Quantitative Correlation Inequalities via Semigroup Interpolation
Published • View PublicationBIB
Most correlation inequalities for high-dimensional functions in the literature, such as the Fortuin-Kasteleyn-Ginibre (FKG) inequality and the celebrated Gaussian Correlation Inequality of Royen, are qualitative statements which establish that any two functions of a certain type have non-negative correlation. In this work we give a general approach that can be used to bootstrap many qualitative correlation inequalities for functions over product spaces into quantitative statements. The approach combines a new extremal result about power series, proved using complex analysis, with harmonic analysis of functions over product spaces. We instantiate this general approach in several different concrete settings to obtain a range of new and near-optimal quantitative correlation inequalities, including: $\bullet$ A quantitative version of Royen's celebrated Gaussian Correlation Inequality. Royen (2014) confirmed a conjecture, open for 40 years, stating that any two symmetric, convex sets must be non-negatively correlated under any centered Gaussian distribution. We give a lower bound on the correlation in terms of the vector of degree-2 Hermite coefficients of the two convex sets, analogous to the correlation bound for monotone Boolean functions over $\{0,1\}^n$ obtained by Talagrand (1996). $\bullet$ A quantitative version of the well-known FKG inequality for monotone functions over any finite product probability space, generalizing the quantitative correlation bound for monotone Boolean functions over $\{0,1\}^n$ obtained by Talagrand (1996). The only prior generalization of which we are aware is due to Keller (2008, 2009, 2012), which extended Talagrand's result to product distributions over $\{0,1\}^n$. We also give two different quantitative versions of the FKG inequality for monotone functions over the continuous domain $[0,1]^n$, answering a question of Keller (2009).
Angle sums of random polytopes
Published • View PublicationBIB
For two families of random polytopes we compute explicitly the expected sums of the conic intrinsic volumes and the Grassmann angles at all faces of any given dimension of the polytope under consideration. As special cases, we compute the expected sums of internal and external angles at all faces of any fixed dimension. The first family are the Gaussian polytopes defined as convex hulls of i.i.d. samples from a non-degenerate Gaussian distribution in $\mathbb R^d$. The second family are convex hulls of random walks with exchangeable increments satisfying certain mild general position assumption. The expected sums are expressed in terms of the angles of the regular simplices and the Stirling numbers, respectively. There are non-trivial analogies between these two settings. Further, we compute the angle sums for Gaussian projections of arbitrary polyhedral sets, of which the Gaussian polytopes are a special case. Also, we show that the expected Grassmann angle sums of a random polytope with a rotationally invariant law are invariant under affine transformations. Of independent interest may be also results on the faces of linear images of polyhedral sets. These results are well known but it seems that no detailed proofs can be found in the existing literature.
Higher rank motivic Donaldson-Thomas invariants of $\mathbb{A}^3$ via wall-crossing, and asymptotics
Published • View PublicationBIB
We compute, via motivic wall-crossing, the generating function of virtual motives of the Quot scheme of points on $\mathbb{A}^3$, generalising to higher rank a result of Behrend, Bryan and Szendrői. We show that this motivic partition function converges to a Gaussian distribution, extending a result of Morrison.
2020-03-18 v3
Convex Hulls of Random Order Types
Published • View PublicationBIB
We establish the following two main results on order types of points in general position in the plane (realizable simple planar order types, realizable uniform acyclic oriented matroids of rank $3$): (a) The number of extreme points in an $n$-point order type, chosen uniformly at random from all such order types, is on average $4+o(1)$. For labeled order types, this number has average $4- \frac{8}{n^2 - n +2}$ and variance at most $3$. (b) The (labeled) order types read off a set of $n$ points sampled independently from the uniform measure on a convex planar domain, smooth or polygonal, or from a Gaussian distribution are concentrated, i.e. such sampling typically encounters only a vanishingly small fraction of all order types of the given size. Result (a) generalizes to arbitrary dimension $d$ for labeled order types with the average number of extreme points $2d+o(1)$ and constant variance. We also discuss to what extent our methods generalize to the abstract setting of uniform acyclic oriented matroids. Moreover, our methods allow to show the following relative of the Erdős-Szekeres theorem: for any fixed $k$, as $n \to \infty$, a proportion $1 - O(1/n)$ of the $n$-point simple order types contain a triangle enclosing a convex $k$-chain over an edge. For the unlabeled case in (a), we prove that for any antipodal, finite subset of the $2$-dimensional sphere, the group of orientation preserving bijections is cyclic, dihedral or one of $A_4$, $S_4$ or $A_5$ (and each case is possible). These are the finite subgroups of $SO(3)$ and our proof follows the lines of their characterization by Felix Klein.
On the asymptotic behavior of the $q$-analog of Kostant's partition function
Published • View PublicationBIB
Kostant's partition function counts the number of distinct ways to express a weight of a classical Lie algebra $\mathfrak{g}$ as a sum of positive roots of $\mathfrak{g}$. We refer to each of these expressions as decompositions of a weight. Our main result considers an infinite family of weights, irrespective of Lie type, for which we establish a closed formula for the $q$-analog of Kostant's partition function and then prove that the (normalized) distribution of the number of positive roots in the decomposition of any of these weights converges to a Gaussian distribution as the rank of the Lie algebra goes to infinity. We also extend these results to the highest root of the classical Lie algebras and we end our analysis with some directions for future research.