probability distribution
284 papers tagged with this keyword
Simple Eigenvalues and Non-vanishing Eigenvectors of the Anderson Model
We consider the Anderson model on the finite grid $G = \mathbb Z/L_1\mathbb Z\times\cdots\times\mathbb Z/L_d\mathbb Z$, defined by the random Hamiltonian $H_t=Δ+tV$, where $Δ$ is the discrete Laplacian and $V=\mathrm{diag}(\{ω_{x}\}_{x\in G})$ is a random onsite potential with $ω_x\simμ$ i.i.d. We ask the natural question of when $H_t$ has simple eigenvalues and non-vanishing eigenvectors. We prove that, when $μ$ is a continuous probability distribution, $H_t$ has this property for all but finitely many $t$ values with probability $1$. However, when $μ$ is a Bernoulli distribution, the conditions fail with positive probability, for which we give a lower bound. We also calculate the exact probability of these conditions being met in the Bernoulli case when $d = 1$ and $L = L_1$ is prime.
Interactive Proofs For Distribution Testing With Conditional Oracles
We revisit the framework of interactive proofs for distribution testing, first introduced by Chiesa and Gur (ITCS 2018), which has recently experienced a surge in interest, accompanied by notable progress (e.g., Herman and Rothblum, STOC 2022, FOCS 2023; Herman, RANDOM~2024). In this model, a data-poor verifier determines whether a probability distribution has a property of interest by interacting with an all-powerful, data-rich but untrusted prover bent on convincing them that it has the property. While prior work gave sample-, time-, and communication-efficient protocols for testing and estimating a range of distribution properties, they all suffer from an inherent issue: for most interesting properties of distributions over a domain of size $N$, the verifier must draw at least $Ω(\sqrt{N})$ samples of its own. While sublinear in $N$, this is still prohibitive for large domains encountered in practice.
In this work, we circumvent this limitation by augmenting the verifier with the ability to perform an exponentially smaller number of more powerful (but reasonable) \emph{pairwise conditional} queries, effectively enabling them to perform ``local comparison checks'' of the prover's claims. We systematically investigate the landscape of interactive proofs in this new setting, giving polylogarithmic query and sample protocols for (tolerantly) testing all \emph{label-invariant} properties, thus demonstrating exponential savings without compromising on communication, for this large and fundamental class of testing tasks.
On the Expected Duration of a Generalized Bingo Game
We investigate the expected number of calls required to achieve Bingo in a generalized (n,m)-Bingo game, where each n x n card is filled by sampling n numbers from m possible values per column. Using the inclusion-exclusion principle, we derive exact formulas for the probability distribution and the expected game length. Our main theoretical result proves that the expected number of calls is a linear function of m.
Maximal Cells in Shifted Staircase Tableaux and a Quarter-Circle Law
In this note, we explicitly compute the probability that a given cell in a random standard Young tableau of the shifted staircase shape $(2n-1, 2n-3, \ldots, 3,1)$ contains the maximal label. We also show that the asymptotic distribution of the cell containing the maximal label is governed by the quarter-circle law. The bijection between the tableaux and thereduced decompositions of the longest element of the group $B_n$ of the signed permutations yields the probability distribution of the first (and any) letter of the random reduced decompositions. We also show the results of some computational experiments on the random sorting networks of $B_n$.
Horton-Strahler numbers for binary butterfly trees: exact analysis
Peca suggested in a recent paper on the arxiv to consider binary butterfly trees and their Horton-Strahler numbers.
The trees are obtained by glueing two binary trees together in a special way; the results are again binary trees but with
a different probability distribution. A thorough combinatorial analysis is provided and leads asymptotically to the same results as for classical binary trees.
Rearrangements of distributions on integers that minimize variance
Which permutations of a probability distribution on integers minimize variance?
Let $X$ be a random variable on a set of integers $\{x_1, \dots, x_N\}$ such that $\mathbb{P}(X_i = x_i) = p_i$, $i \in \{1,\dots,N\}$. Let $(p^{(1)}, \dots, p^{(N)})$ be the sequence $(p_1, \dots, p_N)$ ordered non-increasingly. Let $X^+$ be the random variable defined by $\mathbb{P}(X^+=0)=p^{(1)}$, $\mathbb{P}(X^+=1) = p^{(2)}$, $\mathbb{P}(X^+=-1)=p^{(3)}, \dots, \mathbb{P}(X^+=(-1)^N \lfloor \frac {N} 2 \rfloor)=p^{(N)}$. In this short note we generalize and prove the inequality $\mathrm{Var}\, X^+ \le \mathrm{Var}\, X$.
Trickle-down Theorems via C-Lorentzian Polynomials II: Pairwise Spectral Influence and Improved Dobrushin's Condition
Let $μ$ be a probability distribution on a multi-state spin system on a set $V$ of sites. Equivalently, we can think of this as a $d$-partite simplical complex with distribution $μ$ on maximal faces. For any pair of vertices $u,v\in V$, define the pairwise spectral influence $\mathcal{I}_{u,v}$ as follows. Let $σ$ be a choice of spins $s_w\in S_w$ for every $w\in V \setminus \{u,v\}$, and construct a matrix in $\mathbb{R}^{(S_u\cup S_v)\times (S_u\cup S_v)}$ where for any $s_u\in S_u, s_v\in S_v$, the $(us_u,vs_v)$-entry is the probability that $s_v$ is the spin of $v$ conditioned on $s_u$ being the spin of $u$ and on $σ$. Then $\mathcal{I}_{u,v}$ is the maximal second eigenvalue of this matrix, over all choices of spins for all $w \in V \setminus \{u,v\}$. Equivalently, $\mathcal{I}_{u,v}$ is the maximum local spectral expansion of links of codimension $2$ that include a spin for every $w \in V \setminus \{u,v\}$.
We show that if the largest eigenvalue of the pairwise spectral influence matrix with entries $\mathcal{I}_{u,v}$ is bounded away from 1, i.e. $λ_{\max}(\mathcal{I})\leq 1-ε$ (and $X$ is connected), then the Glauber dynamics mixes rapidly and generate samples from $μ$. This improves/generalizes the classical Dobrushin's influence matrix as the $\mathcal{I}_{u,v}$ lower-bounds the classical influence of $u\to v$. As a by-product, we also prove improved/almost optimal trickle-down theorems for partite simplicial complexes. The proof builds on the trickle-down theorems via $\mathcal{C}$-Lorentzian polynomials machinery recently developed by the authors and Lindberg.
Graph entropy, degree assortativity, and hierarchical structures in networks
Published in Physical Review E 112(6): 064315, 2025
• View Publication
• BIB
We connect several notions relating the structural and dynamical properties of a graph. Among them are the topological entropy coming from the vertex shift, which is related to the spectral radius of the graph's adjacency matrix, the Randić index, and the degree assortativity. We show that, among all connected graphs with the same degree sequence, the graph having maximum entropy is characterized by a hierarchical structure; namely, it satisfies a breadth-first search ordering with decreasing degrees (BFD-ordering for short). Consequently, the maximum-entropy graph necessarily has high degree assortativity; furthermore, for such a graph the degree centrality and eigenvector centrality coincide. Moreover, the notion of assortativity is related to the general Randić index. We prove that the graph that maximizes the Randić index satisfies a BFD-ordering. For trees, the converse holds as well. We also define a normalized Randić function and show that its maximum value equals the difference of Shannon entropies of two probability distributions defined on the edges and vertices of the graph based on degree correlations.
Matroid bingo
We investigate some natural probability distributions associated with the game of matroid bingo.
Identifiability of Large Phylogenetic Mixtures for Many Phylogenetic Model Structures
Identifiability of phylogenetic models is a necessary condition to ensure that the model parameters can be uniquely determined from data. Mixture models are phylogenetic models where the probability distributions in the model are convex combinations of distributions in simpler phylogenetic models. Mixture models are used to model heterogeneity in the substitution process in DNA sequences. While many basic phylogenetic models are known to be identifiable, mixture models in generality have only been shown to be identifiable in certain cases. We expand the main theorem of [Rhodes, Sullivant 2012] to prove identifiability of mixture models in equivariant phylogenetic models, specifically the Jukes-Cantor, Kimura 2-parameter model, Kimura 3-parameter model and the Strand Symmetric model.
Temporal Exploration of Random Spanning Tree Models
The Temporal Graph Exploration problem (TEXP) takes as input a temporal graph, i.e., a sequence of graphs $(G_i)_{i\in \mathbb{N}}$ on the same vertex set, and asks for a walk of shortest length visiting all vertices, where the $i$-th step uses an edge from $G_i$. If each such $G_i$ is connected, then an exploration of length $n^2$ exists, and this is known to be the best possible up to a constant. More fine-grained lower and upper bounds have been obtained for restricted temporal graph classes, however, for several fundamental classes, a large gap persists between known bounds, and it remains unclear which properties of a temporal graph make it inherently difficult to explore.
Motivated by this limited understanding and the central role of the Temporal Graph Exploration problem in temporal graph theory, we study the problem in a randomised setting. We introduce the Random Spanning Tree (RST) model, which consists of a set of $n$-vertex trees together with an arbitrary probability distribution $μ$ over this set. A random temporal graph generated by the RST model is a sequence of independent samples drawn from $μ$.
We initiate a systematic study of the Temporal Graph Exploration problem in such random temporal graphs and establish tight general bounds on exploration time. Our first main result proves that any RST model can, with high probability (w.h.p.), be explored in $O(n^{3/2})$ time, and we show that this bound is tight up to a constant factor. This demonstrates a fundamental difference between the adversarial and random settings. Our second main result shows that if all trees of an RST are subgraphs of a fixed graph with $m$ edges then, w.h.p.\ , it can be explored in $O(m)$ time.
Global fluctuations for standard Young tableaux
We introduce the notion of a Young generating function for a probability measure on integer partitions. We use this object to characterize probability distributions over integer partitions satisfying a law of large numbers and those that satisfy a central limit theorem. We further establish a multilevel central limit theorem, which enables the study of random standard Young tableaux. As applications of these results, we describe the fluctuations of height functions associated with (i) the Plancherel growth process, (ii) random standard Young tableaux of fixed shape, and (iii) probability distributions induced by extreme characters of the infinite symmetric group $S_\infty$. In all cases, we identify the limiting fluctuations as a conditioned Gaussian Free Field.
Discrete Boltzmann distributions via multisets and their coefficients
This paper investigates the combinatorics that gives rise to the Boltzmann probability distribution. Despite being one of the most important distributions in physics and other fields of science, the mathematics of the underlying model of particles at different energy levels is underexplored. This paper gives a reconstruction, using multisets with fixed sums as mathematical representations. Counting (the coefficients of) such multisets gives a general description of binomial, trinomial, quadrinomial etc.\ coefficients, here called N-nomials. These coefficients give rise to multiple discrete Boltzmann distributions that are linked to explanations in the physics literature.
How to Learn a Star: Binary Classification with Starshaped Polyhedral Sets
We consider binary classification restricted to a class of continuous piecewise linear functions whose decision boundaries are (possibly nonconvex) starshaped polyhedral sets, supported on a fixed polyhedral simplicial fan. We investigate the expressivity of these function classes and describe the combinatorial and geometric structure of the loss landscape, most prominently the sublevel sets, for two loss-functions: the 0/1-loss (discrete loss) and a log-likelihood loss function. In particular, we give explicit bounds on the VC dimension of this model, and concretely describe the sublevel sets of the discrete loss as chambers in a hyperplane arrangement. For the log-likelihood loss, we give sufficient conditions for the optimum to be unique, and describe the geometry of the optimum when varying the rate parameter of the underlying exponential probability distribution.
Partial results for union-closed conjectures on the weighted cube
The celebrated union-closed conjecture is concerned with the cardinalities of various subsets of the Boolean $d$-cube. The cardinality of such a set is equivalent, up to a constant, to its measure under the uniform distribution, so we can pose more general conjectures by choosing a different probability distribution on the cube. In particular, for any sequence of probabilities $(p_i)_{i=1}^d$ we can consider the product of $d$ independent Bernoulli random variables, with success probabilities $p_i$. In this short note, we find a generalised form of Karpas' special case of the union-closed conjecture for families $\mathcal{F}$ with density at least half. We also generalise Knill's logarithmic lower bound.
A New Representation of Ewens-Pitman's Partition Structure and Its Characterization via Riordan Array Sums
Ewens-Pitman's partition structure arises as a system of sampling consistent probability distributions on set partitions induced by the Pitman-Yor process. It is widely used in statistical applications, particularly in species sampling models in Bayesian nonparametrics. Drawing references from the area of representation theory of the infinite symmetric group, we view Ewens-Pitman's partition structure as an example of a non-extreme harmonic function on a branching graph, specifically, the Kingman graph. Taking this perspective enables us to obtain combinatorial and algebraic constructions of this distribution using the interpolation polynomial approach proposed by Borodin and Olshanski (The Electronic Journal of Combinatorics, 7, 2000). We provide a new explicit representation of Ewens-Pitman's partition structure using modern umbral interpolation based on Sheffer polynomial sequences. In addition, we show that a certain type of marginals of this distribution can be computed using weighted row sums of a Riordan array. In this way, we show that some summary statistics and estimators derived from Ewens-Pitman's partition structure can be obtained using methods of generating functions. This approach simplifies otherwise cumbersome calculations of these quantities often involving various special combinatorial functions. In addition, it has the added benefit of being amenable to symbolic computation.
Statistics for random representations of Lie algebras
In this paper we investigate how a typical, large-dimensional representation looks for a complex Lie algebra. In particular, we study the family $\mathfrak{sl}_{r+1}(\mathbb{C})$ of Lie algebras for $r \geq 2$ and derive asymptotic probability distributions for the multiplicity of small irreducible representations, as well as the largest dimension, the largest height, and the total number of irreducible representations appearing in the decomposition of a representation sampled uniformly from all representations with the same dimension. This provides a natural generalization to the similar statistical studies of integer partitions, which forms the case $r=1$ of our considerations and where one has a rich toolkit ranging from combinatorial methods to approaches utilizing the theory of modular forms. We perform our analysis by extending the statistical mechanics inspired approaches in the case of partitions to the infinite family here.
Matching adjacent cards
Published in Mathematics Magazine 97 (2024) 471-483
• View Publication
• BIB
In a well-shuffled deck of cards, what is the probability that somewhere in the deck there are adjacent cards of the same rank? What is the average number of adjacent matches? What is the probability distribution for the number of matches? We answer these and related questions for both the standard $52$-card deck with four suits and $13$ ranks and for generalized decks with $k$ suits and $n$ ranks. We also determine the limiting distribution as $n$ goes to infinity with $k$ fixed.
Local Shearer bound
We prove the following local strengthening of Shearer's classic bound on the independence number of triangle-free graphs: For every triangle-free graph $G$ there exists a probability distribution on its independent sets such that every vertex $v$ of $G$ is contained in a random independent set drawn from the distribution with probability $(1-o(1))\frac{\ln d(v)}{d(v)}$. This resolves the main conjecture raised by Kelly and Postle (2018) about fractional coloring with local demands, which in turn confirms a conjecture by Cames van Batenburg et al. (2018) stating that every $n$-vertex triangle-free graph has fractional chromatic number at most $(\sqrt{2}+o(1))\sqrt{\frac{n}{\ln(n)}}$. Addressing another conjecture posed by Cames van Batenburg et al., we also establish an analogous upper bound in terms of the number of edges.
To prove these results we establish a more general technical theorem that works in a weighted setting. As a further application of this more general result, we obtain a new spectral upper bound on the fractional chromatic number of triangle-free graphs: We show that every triangle-free graph $G$ satisfies $χ_f(G)\le (1+o(1))\frac{ρ(G)}{\ln ρ(G)}$ where $ρ(G)$ denotes the spectral radius. This improves the bound implied by Wilf's classic spectral estimate for the chromatic number by a $\ln ρ(G)$ factor and makes progress towards a conjecture of Harris on fractional coloring of degenerate graphs.
On the satisfiability of random $3$-SAT formulas with $k$-wise independent clauses
The problem of identifying the satisfiability threshold of random $3$-SAT formulas has received a lot of attention during the last decades and has inspired the study of other threshold phenomena in random combinatorial structures. The classical assumption in this line of research is that, for a given set of $n$ Boolean variables, each clause is drawn uniformly at random among all sets of three literals from these variables, independently from other clauses. Here, we keep the uniform distribution of each clause, but deviate significantly from the independence assumption and consider richer families of probability distributions. For integer parameters $n$, $m$, and $k$, we denote by $\DistFamily_k(n,m)$ the family of probability distributions that produce formulas with $m$ clauses, each selected uniformly at random from all sets of three literals from the $n$ variables, so that the clauses are $k$-wise independent. Our aim is to make general statements about the satisfiability or unsatisfiability of formulas produced by distributions in $\DistFamily_k(n,m)$ for different values of the parameters $n$, $m$, and $k$.