arXiv++ Combinatorics

Browse math.CO papers from arXiv

probability distribution

284 papers tagged with this keyword
A refined graph container lemma and applications to the hard-core model on bipartite expanders
Published in Random Structures & Algorithms 68 (2026): e70041 • View PublicationBIB
We establish a refined version of a graph container lemma due to Galvin and discuss several applications related to the hard-core model on bipartite expander graphs. Given a graph $G$ and $λ>0$, the hard-core model on $G$ at activity $λ$ is the probability distribution $μ_{G,λ}$ on independent sets in $G$ given by $μ_{G,λ}(I)\propto λ^{|I|}$. As one of our main applications, we show that the hard-core model at activity $λ$ on the hypercube $Q_d$ exhibits a `structured phase' for $λ= Ω( \log^2 d/d^{1/2})$ in the following sense: in a typical sample from $μ_{Q_d,λ}$, most vertices are contained in one side of the bipartition of $Q_d$. This improves upon a result of Galvin which establishes the same for $λ=Ω(\log d/ d^{1/3})$. As another application, we establish a fully polynomial-time approximation scheme (FPTAS) for the hard-core model on a $d$-regular bipartite $α$-expander, with $α>0$ fixed, when $λ= Ω( \log^2 d/d^{1/2})$. This improves upon the bound $λ=Ω(\log d/ d^{1/4})$ due to the first author, Perkins and Potukuchi. We discuss similar improvements to results of Galvin-Tetali, Balogh-Garcia-Li and Kronenberg-Spinka.
2024-10-03
Fractional list packing for layered graphs
The fractional list packing number $χ_{\ell}^{\bullet}(G)$ of a graph $G$ is a graph invariant that has recently arisen from the study of disjoint list-colourings. It measures how large the lists of a list-assignment $L:V(G)\rightarrow 2^{\mathbb{N}}$ need to be to ensure the existence of a `perfectly balanced' probability distribution on proper $L$-colourings, i.e., such that at every vertex $v$, every colour appears with equal probability $1/|L(v)|$. In this work we give various bounds on $χ_{\ell}^{\bullet}(G)$, which admit strengthenings for correspondence and local-degree versions. As a corollary, we improve theorems on the related notion of flexible list colouring. In particular we study Cartesian products and $d$-degenerate graphs, and we prove that $χ_{\ell}^{\bullet}(G)$ is bounded from above by the pathwidth of $G$ plus one. The correspondence analogue of the latter is false for treewidth instead of pathwidth.
2024-09-30
Combinatorics of a dissimilarity measure for pairs of draws from discrete probability vectors on finite sets of objects
Motivated by a problem in population genetics, we examine the combinatorics of dissimilarity for pairs of random unordered draws of multiple objects, with replacement, from a collection of distinct objects. Consider two draws of size $K$ taken with replacement from a set of $I$ objects, where the two draws represent samples from potentially distinct probability distributions over the set of $I$ objects. We define the set of \emph{identity states} for pairs of draws via a series of actions by permutation groups, describing the enumeration of all such states for a given $K \geq 2$ and $I \geq 2$. Given two probability vectors for the $I$ objects, we compute the probability of each identity state. From the set of all such probabilities, we obtain the expectation for a dissimilarity measure, finding that it has a simple form that generalizes a result previously obtained for the case of $K=2$. We determine when the expected dissimilarity between two draws from the same probability distribution exceeds that of two draws taken from different probability distributions. We interpret the results in the setting of the genetics of polyploid organisms, those whose genetic material contains many copies of the genome ($K > 2$).
2024-09-12
A Successive Refinement Algorithm for Tri-Level Stochastic Defender-Attacker Problems with Decision-Dependent Probability Distributions
Tri-level defender-attacker game models are a well-studied method for determining how best to protect a system (e.g., a transportation network) from attacks. Existing models assume that defender and attacker actions have a perfect effect, i.e., system components hardened by a defender cannot be destroyed by the attacker, and attacked components always fail. Because of these assumptions, these models produce solutions in which defended components are never attacked, a result that may not be realistic in some contexts. This paper considers an imperfect defender-attacker problem in which defender decisions (e.g., hardening) and attacker decisions (e.g., interdiction) have an imperfect effect such that the probability distribution of a component's capacity depends on the amount of defense and attack resource allocated to the component. Thus, this problem is a stochastic optimization problem with decision-dependent probabilities and is challenging to solve because the deterministic equivalent formulation has many high-degree multilinear terms. To address the challenges in solving this problem, we propose a successive refinement algorithm that dynamically refines the support of the random variables as needed, leveraging the fact that a less-refined support has fewer scenarios and multilinear terms and is, therefore, easier to solve. A comparison of the successive refinement algorithm versus the deterministic equivalent formulation on a tri-level stochastic maximum flow problem indicates that the proposed method solves many more problem instances and is up to $66$ times faster. These results indicate that it is now possible to solve tri-level problems with imperfect hardening and attacks.
2024-09-10 v2
The number of solutions of a random system of polynomials over a finite field
We study the probability distribution of the number of common zeros of a system of $m$ random $n$-variate polynomials over a finite commutative ring $R$. We compute the expected number of common zeros of a system of polynomials over $R$. Then, in the case that $R$ is a field, under a necessary-and-sufficient condition on the sample space, we show that the number of common zeros is binomially distributed.
2024-08-22
The distribution of the length of the longest path in random acyclic orientations of a complete bipartite graph
Randomly sampling an acyclic orientation on the complete bipartite graph $K_{n,k}$ with parts of size $n$ and $k$, we investigate the length of the longest path. We provide a probability generating function for the distribution of the longest path length, and we use Analytic Combinatorics to perform asymptotic analysis of the probability distribution in the case of equal part sizes $n = k$ tending toward infinity. We show that the distribution is asymptotically Gaussian, and we obtain precise asymptotics for the mean and variance. These results address a question asked by Peter J. Cameron. Keywords: bipartite graph, directed graph, random graph, acyclic orientation, poly-Bernoulli numbers, lonesum matrices, generating function, analytic combinatorics, asymptotics.
2024-08-11
Möbius inversion and the bootstrap
Estimating nonlinear functionals of probability distributions from samples is a fundamental statistical problem. The "plug-in" estimator obtained by applying the target functional to the empirical distribution of samples is biased. Resampling methods such as the bootstrap derive artificial datasets from the original one by resampling. Comparing the outcome of the plug-in estimator in the original and resampled datasets allows estimating and thus correcting the bias. In the asymptotic setting, iterations of this procedure attain an arbitrarily high order of bias correction, but finite sample results are scarce. This work develops a new theoretical understanding of bootstrap bias correction by viewing it as an iterative linear solver for the combinatorial operation of Möbius inversion. It sharply characterizes the regime of linear convergence of the bootstrap bias reduction for moment polynomials. It uses these results to show its superalgebraic convergence rate for band-limited functionals. Finally, it derives a modified bootstrap iteration enabling the unbiased estimation of unknown order-$m$ moment polynomials in $m$ bootstrap iterations.
2024-06-26
Equilibria in a Hypercube Spatial Voting Model
We give conditions for equilibria in the following Voronoi game on the discrete hypercube. Two players position themselves in $\{0,1\}^d$ and each receives payoff equal to the measure (under some probability distribution) of their Voronoi cell (the set of all points which are closer to them than to the other player). This game can be thought of as a discrete analogue of the Hotelling--Downs spatial voting model in which the political spectrum is determined by $d$ binary issues rather than a continuous interval. We observe that if an equilibrium does exist then it must involve the two players co-locating at the majority point (ie the point representing majority opinion on each separate issue). Our main result is that a sufficient condition for an equilibrium is that on each issue the majority option is held by at least $\frac{3}{4}$ of voters. The value $\frac{3}{4}$ can be improved slightly in a way that depends on $d$ and with this improvement the result is best possible. We give similar sufficient conditions for the existence of a local equilibrium. We also analyse the situation where the distribution is a mix of two product measures. We show that either there is an equilibrium or the best response to the majority point is its antipode.
2024-06-12
On the equivalence of quasirandomness and exchangeable representations independent from lower-order variables
It is often convenient to represent a process for randomly generating a graph as a graphon. (More precisely, these give \emph{vertex exchangeable} processes -- those processes in which each vertex is treated the same way.) Other structures can be treated by generalizations like hypergraphons, permutatons, and, for a very general class, theons. These representations are not unique: different representations can lead to the same probability distribution on graphs. This naturally leads to questions (going back at least to Hoover's proof of the Aldous--Hoover Theorem on the existence of such representations) that ask when quasirandomness properties on the distribution guarantee the existence of particularly simple representations. We extend the usual theon representation by adding an additional datum of a random permutation to each tuple, which we call a $\ast$-representation. We show that if a process satisfies the \emph{unique coupling} property UCouple[$\ell$], which says roughly that all $\ell$-tuples of vertices ``look the same'', then the process is $\ast$-$\ell$-independent: there is a $\ast$-representation that does not make use of any random information about $\ell$-tuples (including tuples of length $<\ell$). Simple examples show that the use of $\ast$-representations is necessary. This resolves a question of Coregliano and Razborov, since it easily follows that UCouple[l] implies Independence[\ell'] (the existence of an $\ell'$-independent ordinary representation) for $\ell'<\ell$.
2024-06-09
Probabilistic Approach to Black-Box Binary Optimization with Budget Constraints: Application to Sensor Placement
We present a fully probabilistic approach for solving binary optimization problems with black-box objective functions and with budget constraints. In the probabilistic approach, the optimization variable is viewed as a random variable and is associated with a parametric probability distribution. The original optimization problem is replaced with an optimization over the expected value of the original objective, which is then optimized over the probability distribution parameters. The resulting optimal parameter (optimal policy) is used to sample the binary space to produce estimates of the optimal solution(s) of the original binary optimization problem. The probability distribution is chosen from the family of Bernoulli models because the optimization variable is binary. The optimization constraints generally restrict the feasibility region. This can be achieved by modeling the random variable with a conditional distribution given satisfiability of the constraints. Thus, in this work we develop conditional Bernoulli distributions to model the random variable conditioned by the total number of nonzero entries, that is, the budget constraint. This approach (a) is generally applicable to binary optimization problems with nonstochastic black-box objective functions and budget constraints; (b) accounts for budget constraints by employing conditional probabilities that sample only the feasible region and thus considerably reduces the computational cost compared with employing soft constraints; and (c) does not employ soft constraints and thus does not require tuning of a regularization parameter, for example to promote sparsity, which is challenging in sensor placement optimization problems. The proposed approach is verified numerically by using an idealized bilinear binary optimization problem and is validated by using a sensor placement experiment in a parameter identification setup.
2024-05-28
Upper Bounds on the Average Height of Random Binary Trees
We study the average height of random trees generated by leaf-centric binary tree sources as introduced by Zhang, Yang and Kieffer. A leaf-centric binary tree source induces for every $n \geq 2$ a probability distribution on the set of binary trees with $n$ leaves. Our results generalize a result by Devroye, according to which the average height of a random binary search tree of size $n$ is in $\mathcal{O}(\log n)$.
2024-05-15
Ahead of the Count: An Algorithm for Probabilistic Prediction of Instant Runoff (IRV) Elections
How can we probabilistically predict the winner in a ranked-choice election without all ballots being counted? In this study, we introduce a novel algorithm designed to predict outcomes in Instant Runoff Voting (IRV) elections. The algorithm takes as input a set of discrete probability distributions describing vote totals for each candidate ranking and calculates the probability that each candidate will win the election. In fact, we calculate all possible sequences of eliminations that might occur in the IRV rounds and assign a probability to each. The discrete probability distributions can be arbitrary and, in applications, could be measured empirically from pre-election polling data or from partial vote tallies of an in-progress election. The algorithm is effective for elections with a small number of candidates (five or fewer), with fast execution on typical consumer computers. The run-time is short enough for our method to be used for real-time election night modeling where new predictions are made continuously as more and more vote information becomes available. We demonstrate the algorithm in abstract examples, and also using real data from the 2022 Alaska state elections to simulate election-night predictions and also predictions of election recounts.
2024-05-10 v2
Positive formula for the product of conjugacy classes on the unitary group
The convolution product of two conjugacy classes of the unitary group $U_n$ is described by a probability distribution on the space of central measures. Relating this convolution to the quantum cohomology of Grassmannians and using recent results describing the structure constants of the latter, we give a manifestly positive formula for the density of the probability distribution for the product of generic conjugacy classes. In the same flavor as the hive model of Knutson and Tao, this formula is given in terms of a subtraction-free sum of volumes of explicit polytopes. As a consequence, this expression also provides a positive and explicit formula for the volume of $SU_n$-valued flat connections on the three-holed two dimensional sphere, which was first given by Witten in terms of an infinite sum of characters.
2024-05-09 v2
The largest subgraph without a forbidden induced subgraph
We initiate the systematic study of the following Turán-type question. Suppose $Γ$ is a graph with $n$ vertices such that the edge density between any pair of subsets of vertices of size at least $t$ is at most $1 - c$, for some $t$ and $c > 0$. What is the largest number of edges in a subgraph $G \subseteq Γ$ which does not contain a fixed graph $H$ as an induced subgraph or, more generally, which belongs to a hereditary property $\mathcal{P}$? This provides a common generalization of two recently studied cases, namely $Γ$ being a (pseudo-)random graph and a graph without a large complete bipartite subgraph. We focus on the interesting case where $H$ is a bipartite graph. We determine the answer up to a constant factor with respect to $n$ and $t$, for certain bipartite $H$ and for $Γ$ either a dense random graph or a Paley graph with a square number of vertices. In particular, our bounds match if $H$ is a tree, or if one part of $H$ has $d$ vertices complete to the other part, all other vertices in that part have degree at most $d$, and the other part has sufficiently many vertices. As applications of the latter result, we answer a question of Alon, Krivelevich, and Samotij on the largest subgraph with a hereditary property which misses a bipartite graph, and determine up to a constant factor the largest number of edges in a string subgraph of $Γ$. The proofs are based on a variant of the dependent random choice and a novel approach for finding induced copies by inductively defining probability distributions supported on induced copies of smaller subgraphs.
2024-04-15
On the geometry of exponential random graphs and applications
In a seminal paper in 2009, Borcea, Brändén, and Liggett described the connection between probability distributions and the geometry of their generating polynomials. Namely, they characterized that stable generating polynomials correspond to distributions with the strongest form of negative dependence. This motivates us to investigate other distributions that can have this property, and our focus is on random graph models. In this article, we will lay the groundwork to investigate Markov random graphs, and more generally exponential random graph models (ERGMs), from this geometric perspective. In particular, by determining when their corresponding generating polynomials are either stable and/or Lorentzian. The Lorentzian property was first described in 2020 by Brändén and Huh and independently by Anari, Oveis-Gharan, and Vinzant where the latter group called it the completely log-concave property. The theory of stable polynomials predates this, and is commonly thought of as the multivariate notion of real-rootedness. Brändén and Huh proved that stable polynomials are always Lorentzian. Although it is a strong condition, verifying stability is not always feasible. We will characterize when certain classes of Markov random graphs are stable and when they are only Lorentzian. We then shift our attention to applications of these properties to real-world networks.
2024-04-11 v2
Glauber dynamics for the hard-core model on bounded-degree $H$-free graphs
Published in Combinator. Probab. Comp. 34 (2025) 803-814 • View PublicationBIB
The hard-core model has as its configurations the independent sets of some graph instance $G$. The probability distribution on independent sets is controlled by a `fugacity' $λ>0$, with higher $λ$ leading to denser configurations. We investigate the mixing time of Glauber (single-site) dynamics for the hard-core model on restricted classes of bounded-degree graphs in which a particular graph $H$ is excluded as an induced subgraph. If $H$ is a subdivided claw then, for all $λ$, the mixing time is $O(n\log n)$, where $n$ is the order of $G$. This extends a result of Chen and Gu for claw-free graphs. When $H$ is a path, the set of possible instances is finite. For all other $H$, the mixing time is exponential in $n$ for sufficiently large $λ$, depending on $H$ and the maximum degree of $G$.
Faithlessness in Gaussian graphical models
The implication problem for conditional independence (CI) asks whether the fact that a probability distribution obeys a given finite set of CI relations implies that a further CI statement also holds in this distribution. This problem has a long and fascinating history, cumulating in positive results about implications now known as the semigraphoid axioms as well as impossibility results about a general finite characterization of CI implications. Motivated by violation of faithfulness assumptions in causal discovery, we study the implication problem in the special setting where the CI relations are obtained from a directed acyclic graphical (DAG) model along with one additional CI statement. Focusing on the Gaussian case, we give a complete characterization of when such an implication is graphical by using algebraic techniques. Moreover, prompted by the relevance of strong faithfulness in statistical guarantees for causal discovery algorithms, we give a graphical solution for an approximate CI implication problem, in which we ask whether small values of one additional partial correlation entail small values for yet a further partial correlation.
2024-03-12 v2
$λ$-shaped random matrices, $λ$-plane trees, and $λ$-Dyck paths
Published in Electron. J. Probab. 30: 1-24, #11 (2025) • View PublicationBIB
We consider random matrices whose shape is the dilation $Nλ$ of a self-conjugate Young diagram $λ$. In the large-$N$ limit, the empirical distribution of the squared singular values converges almost surely to a probability distribution $F^λ$. The moments of $F^λ$ enumerate two combinatorial objects: $λ$-plane trees and $λ$-Dyck paths, which we introduce and show to be in bijection. We also prove that the distribution $F^λ$ is algebraic, in the sense of Rao and Edelman. In the case of fat hook shapes we provide explicit formulae for $F^λ$ and we express it as a free convolution of two measures involving a Marchenko-Pastur and a Bernoulli distribution.
2024-02-26 v2
Marginal Independence and Partial Set Partitions
We establish a bijection between marginal independence models on $n$ random variables and split closed order ideals in the poset of partial set partitions. We also establish that every discrete marginal independence model is toric in cdf coordinates. This generalizes results of Boege, Petrovic, and Sturmfels and Drton and Richardson, and provides a unified framework for discussing marginal independence models. Additionally, we provide an axiomatic characterization of marginal independence and we show that our set of axioms are sound and complete in the set of probability distributions. This follows the work of Geiger, Paz and Pearl who provided an analogous characterization of independence for statements involving 2 sets of random variables.
2024-01-15 v3
Probability Mass Function, Moments and Factorial Moments of the Negative Binomial Distribution NB$(k,r)$
The negative binomial distribution NB$(k,r)$ of Type I is the probability distribution for a sequence of independent Bernoulli trials (with success parameter $p\in(0,1)$) with $r$ nonoverlapping success runs of length $\ge k$. We present a new, more concise, expression for its probability mass function. We show it can also be succinctly written using hypergeometric functions. We also present new expressions (combinatorial sums) for its moments and factorial moments, as opposed to only the mean and variance (which are already known). Next, we present an alternative non-combinatorial viewpoint, which yields expressions for the factorial moments not only for nonoverlapping success runs, but also for runs with an overlap of $\ell$, where $\ell\in[0,k-1]$. The case $\ell=k-1$ is the negative binomial distribution NB$(k,r)$ of Type III. The results also yield the solution for the negative binomial distribution NB$(k,r)$ with a minimum gap between the success runs (explained in the text). Addendum 1/23/2024: The probability mass function and factorial moments are derived from the probability generating function. Addendum 1/26/2024: Alternative expressions are presented for the negative binomial distribution NB$(k,r)$ of Type II.