arXiv++ Combinatorics

Browse math.CO papers from arXiv

random

6950 papers tagged with this keyword
2026-03-02
On the expected value of energy in groups
We obtain explicit upper and lower bounds for the expected action energy associated with a pair $({\sf A},{\sf Δ})$ of subsets sampled uniformly at random from a permutation group and its domain, respectively. We then specialize these bounds to multiplicative energy in several settings. In particular, we derive sharp asymptotic formulae for the expected energy of pairs of the form $({\sf A},{\sf A})$ and $({\sf A},{\sf A}^{-1})$. Finally, we apply these estimates to derive probabilistic results on the existence of subsets with large growth and to compare the typical behaviour of the cardinalities of the sets $|{\sf A}^{\ast 2}|$ and $|{\sf A}{\sf A}^{-1}|$.
2026-02-27
Combinatorial sufficient conditions for graph rigidity and applications to random graphs
A graph $G=(V,E)$ is called $d$-rigid if, for a generic embedding of its vertices in $\mathbb{R}^d$, every edge-length preserving continuous motion of the vertices preserves the distances between all pairs of non-adjacent vertices as well. In this paper, we present several new results on the rigidity of random graphs. In particular, we show that there exists $c>0$ such that, for $p\ge 2 \log{n}/n$, the binomial random graph $G(n,p)$ is with high probability (whp) $\lfloor c n p\rfloor$-rigid. This is sharp up to the constant $c$, and complements recent results of Peled and Peleg (in the regime $p= o(n^{-1/2})$), and of Jordán, Liu, and Villányi (in the constant $p$ regime). Moreover, we show that for every fixed $d\ge 2$ and $r\ge 501d$, a random $r$-regular graph is whp $d$-rigid, and that for $100/n\le p\le 2\log{n}/n$, the binomial random graph $G(n,p)$ contains whp an $\lfloor np/251\rfloor$-rigid subgraph with at least $(1-e^{-np/2})n$ vertices. Both results are sharp up to the multiplicative constant. In addition, we present a new sufficient condition for rigidity in terms of the minimum codegree of the graph (the minimum number of common neighbours of a pair of vertices in the graph). A main tool in our arguments is a new combinatorial sufficient condition for rigidity, which provides a common generalization to Whiteley's vertex-splitting lemmas, and to the "rigid partitions" method, developed in works by Crapo, Lindemann, Lew, Nevo, Peled and Raz, and by the present authors.
2026-02-27
Block-weighted random graphs: planar and beyond
We investigate random connected graphs from a block-stable class whose distribution is weighted based on the number of $2$-connected components, or blocks. This includes the class of planar graphs. For this, we develop a notion of a decorated block tree. Following similar ideas to Fleurat and the second author on block-weighted planar maps, we find a phase transition in the singular behaviour of the appropriate generating function and in the typical structure of the block tree. Moreover, for certain block-stable classes (including planar graphs), we obtain precise enumeration results and determine also the typical sizes of the largest blocks in subcritical, critical, and supercritical regimes. It strengthens previously known results on block sizes in uniform random planar graphs.
2026-02-27
Aldous-type Spectral Gaps in Unitary Groups
Aldous' spectral gap conjecture, proven by Caputo, Liggett and Richthammer, states the following: for any set of transpositions in the symmetric group $\mathrm{Sym}(n)$, the spectral gap of the corresponding random walk on the group -- an $n!$-state process -- coincides with that of the corresponding random walk of a single element -- an $n$-state process. This paper presents an analog of this conjecture in the unitary group $\mathrm{U}(n)$, and proves it in several non-trivial cases. The phenomenon we discover is that for some natural families of probability distributions on $\mathrm{U}(n)$, the spectral gap of the corresponding random walk, which has a continuous state space, is identical to that of a discrete KMP process (also known as the uniform reshuffling process) with two indistinguishable particles on a hypergraph on $n$ vertices -- a discrete Markov chain with $\binom{n+1}{2}$ states.
2026-02-23
Identification in Stochastic Choice
We characterize the identified sets of a wide range of stochastic choice models, including random utility, various models of boundedly-rational behavior, and dynamic discrete choice. In each of these settings, we show two distributions over choice rules are observationally equivalent if and only if they can be obtained from one another via a finite sequence of simple swapping transforms. We leverage this to obtain complete descriptions of both the defining inequalities and extreme points of these identified sets. In cases where choice frequencies vary smoothly with some parameters, we provide a novel global-inverse result for practically testing identification.
2026-02-22
Statistical Analysis of Hairpins and BasePairs in RNA Secondary Structures
We derive precise asymptotic expressions for the expectations, variances, covariance, and quite a few further mixed moments for the number of hairpins and the number of basepairs in RNA secondary structures, and give convincing evidence that the central-scaled distribution of the pair of random variables (hairpins, basepairs) tends in distribution to the bi-variate normal distribution with correlation $\sqrt{5 \sqrt{5} -11}/2= 0.2123322205\dots$
On constructing small subgraphs in the budget-constrained random graph process
Consider the budget-constrained random graph process introduced by Frieze, Krivelevich and Michaeli, where each time an edge is offered through the (standard) random graph process we must irrevocably decide whether to "purchase" this edge or not, with our goal being to construct a graph which satisfies some property within a given time $t$ and while purchasing at most $b$ edges. We consider the problem of constructing graphs containing certain fixed small subgraphs. We provide an optimal strategy for building a graph which contains a copy of $K_4$, showing that budget $b=ω(\max\{n^8/t^5,n^2/t\})$ suffices and that if $b=o(\max\{n^8/t^5,n^2/t\})$ then no strategy can a.a.s. produce a graph containing a copy of $K_4$. This resolves a problem raised by Iľkovič, León and Shu. More generally, we obtain analogously tight results for containing a wheel of any fixed size, or a graph consisting of a tree plus one additional universal vertex. We also tackle the problem of constructing graphs containing a copy of $K_5$, obtaining both lower and upper bounds on the optimal budget, though a gap remains in this case.
2026-02-19
Bilateral parking procedures
We introduce the class of bilateral parking procedures on the integer line. While cars try to park in the nearest available spot to their right in the classical case, we consider more general parking rules that allow cars to use the nearest available spot to their left. We show that for a natural subclass of local procedures, the number of corresponding parking functions of length $r$ is always equal to $(r+1)^{r-1}$. The setting can be extended to probabilistic procedures, in which the decision to park left or right is random. We finally describe how bilateral procedures can naturally be encoded by certain labeled binary forests, whose combinatorics shed light on several results from the literature.
Canonical labelling of random regular graphs
We prove that whenever $d=d(n)\to\infty$ and $n-d\to\infty$ as $n\to\infty$, then with high probability for any non-trivial initial colouring, the colour refinement algorithm distinguishes all vertices of the random regular graph $\mathcal{G}_{n,d}$. This, in particular, implies that with high probability $\mathcal{G}_{n,d}$ admits a canonical labelling computable in time $O(\min\{n^ω,nd^2+nd\log n\})$, where $ω<2.372$ is the matrix multiplication exponent.
2026-02-19
A logical approach to concentration
Concentration results say that a sequence of random variables becomes progressively concentrated around the mean. Such results are common in the study of functions of random graphs. We introduce a real-valued logic with various aggregate operators on graphs, including summation, and prove that every term in the language, seen as a random variable on random graphs within the classical Erdős-Rényi random graph model, is concentrated. We prove this for dense and sparse variants of Erdős-Rényi graphs. On the one hand, our results extend the line of work originating with Fagin and Glebskii et al. on zero-one laws for dense random graphs, as well as the zero-one law of Shelah and Spencer for sparse random graphs. On the other hand, they can be seen as a meta-theorem for inferring concentration results on random graphs, and we give examples of such applications.
2026-02-19
On Sets of Monochromatic Objects in Bicolored Point Sets
Let $P$ be a set of $n$ points in the plane, not all on a line, each colored \emph{red} or \emph{blue}. The classical Motzkin--Rabin theorem guarantees the existence of a \emph{monochromatic} line. Motivated by the seminal work of Green and Tao (2013) on the Sylvester-Gallai theorem, we investigate the quantitative and structural properties of monochromatic geometric objects, such as lines, circles, and conics. We first show that if no line contains more than three points, then for all sufficiently large $n$ there are at least $n^{2}/24 - O(1)$ monochromatic lines. We then show a converse of a theorem of Jamison (1986): Given $n\ge 6$ blue points and $n$ red points, if the blue points lie on a conic and every line through two blue points contains a red point, then all red points are collinear. We also settle the smallest nontrivial case of a conjecture of Milićević (2018) by showing that if we have $5$ blue points with no three collinear and $5$ red points, if the blue points lie on a conic and every line through two blue points contains a red point, then all $10$ points lie on a cubic curve. Further, we analyze the random setting and show that, for any non-collinear set of $n\ge 10$ points independently colored red or blue, the expected number of monochromatic lines is minimized by the \emph{near-pencil} configuration. Finally, we examine monochromatic circles and conics, and exhibit several natural families in which no such monochromatic objects exist.
2026-02-18
Anticoncentration of Random Sums in $\mathbb{Z}_p$
In this paper we investigate the probability distribution of the sum $Y$ of $\ell$ independent identically distributed random variables taking values in $\mathbb{Z}_p$. Our main focus is the regime of small values of $\ell$, which is less explored compared to the asymptotic case $\ell \to \infty$. Starting with the case $\ell=3$, we prove that if the distributions of the $Y_i$ are uniformly bounded by $λ< 1$ and $p > 2/λ$, then there exists a constant $C_{3,λ} < 1$ such that \[ \max_{x \in \mathbb{Z}_p} \mathbb{P}[Y = x] \leq C_{3,λ}λ. \] Moreover, when the distributions are uniformly separated from $1$, the constant $C_{3,λ}$ can be made explicit. By iterating this argument, we obtain effective anticoncentration bounds for larger values of $\ell$, yielding nontrivial estimates already in small and moderate regimes where asymptotic results do not apply.
2026-02-18
Comparability of random permutations in the strong Bruhat order
The (strong) Bruhat order for permutations provides a partial ordering defined as follows: two permutations are comparable if one can be obtained from the other by a sequence of adjacent transpositions that each increase the number of inversions by $1$. Given two random permutations, what is the probability that they are comparable in the Bruhat order? This problem was first considered in a 2006 work of Hammett and Pittel, which showed an exponential lower bound and a polynomial upper bound. The lower bound was very recently improved to the subexponential bound of $\exp(-n^{1/2 + o(1)})$ by Boretsky, Cornejo, Hodges, Horn, Lesnevich, and McAllister. Hammett and Pittel predicted that the probability should decrease polynomially. We show that the probability decreases faster than any polynomial and is on the order of $\exp(-Θ(\log^2 n))$.
2026-02-16
Large expander subgraphs in high genus triangulations
We prove that random triangulations of high genus contain very large expander subgraphs, answering a question of Benjamini. Our approach relies on new general criteria for arbitrary graphs to contain large expander subgraphs.
2026-02-15
Vertex operators, infinite wedge representations, and correlation functions of the t-Schur measure
We study the $t$-Schur measure on partitions, defined by $ \mathbb{P}(λ)=Z^{-1}S_λ(x;t)s_λ(y) $, where $S_λ(x;t)$ denotes the $t$-Schur symmetric functions and $s_λ(y)$ the ordinary Schur functions, and $Z$ is the normalising constant. Using vertex operator calculus, we realise $S_λ(x;t)$ in the charged free-fermion Fock space, yielding a $t$-deformation of the classical boson-fermion correspondence. These realisations give vertex-algebraic proofs of the $t$-Cauchy identities and $t$-Gessel identity. Building on this framework, we compute the correlation functions of the $t$-Schur measure and show that the associated point process is determinantal, with an explicit correlation kernel. The Poissonised $t$-Plancherel measure appears as a specialisation of our construction, so its correlation functions follow as a corollary. As an application, we derive the limiting distribution for the length of the longest ascent pair in a random permutation. Our results interpolate the Schur case at $t=0$, connect to the Schur-$Q$ theory at $t=-1$, and provide a probabilistic interpretation of a natural $t$-refinement of increasing subsequences via a generalised RSK correspondence.
Sidorenko property and forcing in regular tournaments
We give a complete characterization of tournaments H that have the Sidorenko property with respect to nearly regular tournaments, i.e., the homomorphism density of H among all nearly regular tournaments is minimized by a random tournament. Corollaries of our result are a positive answer to the question of Noel, Ranganathan and Simbaqueba whether there exist infinitely many non-transitive tournaments that are quasirandom forcing for nearly regular tournaments, and a negative answer to their question whether almost every tournament is quasirandom forcing for nearly regular tournaments.
Finite-sample confidence regions for spectral clustering and graph centrality
Let a graph be observed through a finite random sampling mechanism. Spectral methods are routinely applied to such graphs, yet their outputs are treated as deterministic objects. This paper develops finite-sample inference for spectral graph procedures. The primary result constructs explicit confidence regions for latent eigenspaces of graph operators under an explicit sampling model. These regions propagate to confidence regions for spectral clustering assignments and for smooth graph centrality functionals. All bounds are nonasymptotic and depend explicitly on the sample size, noise level, and spectral gap. The analysis isolates a failure of common practice: asymptotic perturbation arguments are often invoked without a finite-sample spectral gap, leading to invalid uncertainty claims. Under verifiable gap and concentration conditions, the present framework yields coverage guarantees and certified stability regions. Several corollaries address fairness-constrained post-processing and topological summaries derived from spectral embeddings.
2026-02-11
Note on the trace of random walks on pseudorandom graphs
We study the graph-theoretic properties of the trace of random walks on pseudorandom graphs. We show that for any $\varepsilon>0$, there exists a constant $C$ such that the cover time of an $(n,d,λ)$-graph $G$ with $d/λ\ge C$ is at most $(1+\varepsilon)n\log n$, meaning the expected number of steps needed to reach all vertices at least once is at most $(1+\varepsilon)n\log n$ regardless of the starting vertex. Furthermore, we prove that with high probability, the trace of a random walk of length $(1+\varepsilon)n\log n$ on $G$ is Hamiltonian, regardless of the starting vertex. These results also hold for random $d$-regular graphs with sufficiently large $d$. These findings answer two questions proposed by Frieze, Krivelevich, Michaeli, and Peled [PLMS, 2018]. Notably, our results imply a bound on a stronger version of the cover time: with high probability, all vertices are covered after $(1+\varepsilon)n\log n$ steps, regardless of the starting vertex. Our proofs rely on the spectral properties of the adjacency matrix and the graph expansion. All results are asymptotically optimal.
2026-02-11
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
We introduce a mathematical framework for the linear representation hypothesis (LRH), which asserts that intermediate layers of language models store features linearly. We separate the hypothesis into two claims: linear representation (features are linearly embedded in neuron activations) and linear accessibility (features can be linearly decoded). We then ask: How many neurons $d$ suffice to both linearly represent and linearly access $m$ features? Classical results in compressed sensing imply that for $k$-sparse inputs, $d = O(k\log (m/k))$ suffices if we allow non-linear decoding algorithms (Candes and Tao, 2006; Candes et al., 2006; Donoho, 2006). However, the additional requirement of linear decoding takes the problem out of the classical compressed sensing, into linear compressed sensing. Our main theoretical result establishes nearly-matching upper and lower bounds for linear compressed sensing. We prove that $d = Ω_ε(\frac{k^2}{\log k}\log (m/k))$ is required while $d = O_ε(k^2\log m)$ suffices. The lower bound establishes a quantitative gap between classical and linear compressed setting, illustrating how linear accessibility is a meaningfully stronger hypothesis than linear representation alone. The upper bound confirms that neurons can store an exponential number of features under the LRH, giving theoretical evidence for the "superposition hypothesis" (Elhage et al., 2022). The upper bound proof uses standard random constructions of matrices with approximately orthogonal columns. The lower bound proof uses rank bounds for near-identity matrices (Alon, 2003) together with Turán's theorem (bounding the number of edges in clique-free graphs). We also show how our results do and do not constrain the geometry of feature representations and extend our results to allow decoders with an activation function and bias.
2026-02-11
An Improved Upper Bound for the Euclidean TSP Constant Using Band Crossovers
Consider $n$ points generated uniformly at random in the unit square, and let $L_n$ be the length of their optimal traveling salesman tour. Beardwood, Halton, and Hammersley (1959) showed $L_n / \sqrt n \to β$ almost surely as $n\to \infty$ for some constant $β$. The exact value of $β$ is unknown but estimated to be approximately $0.71$ (Applegate, Bixby, Chvátal, Cook 2011). Beardwood et al. further showed that $0.625 \leq β\leq 0.92116.$ Currently, the best known bounds are $0.6277 \leq β\leq 0.90380$, due to Gaudio and Jaillet (2019) and Carlsson and Yu (2023), respectively. The upper bound was derived using a computer-aided approach that is amenable to lower bounds with improved computation speed. In this paper, we show via simulation and concentration analysis that future improvement of the $0.90380$ is limited to $\sim0.88$. Moreover, we provide an alternative tour-constructing heuristic that, via simulation, could potentially improve the upper bound to $\sim0.85$. Our approach builds on a prior \emph{band-traversal} strategy, initially proposed by Beardwood et al. (1959) and subsequently refined by Carlsson and Yu (2023): divide the unit square into bands of height $Θ(1/\sqrt{n})$, construct paths within each band, and then connect the paths to create a TSP tour. Our approach allows paths to cross bands, and takes advantage of pairs of points in adjacent bands which are close to each other. A rigorous numerical analysis improves the upper bound to $0.90367$.