Papers by Elchanan Mossel
52 paper(s) by this author
· All BibTeX
Asymptotics for the harmonic descent chain and applications to critical beta-splitting trees
Motivated by the connection to a probabilistic model of phylogenetic trees introduced by Aldous, we study the recursive sequence governed by the rule $x_n = \sum_{i=1}^{n-1} \frac{1}{h_{n-1}(n-i)} x_i$ where $h_{n-1} = \sum_{j=1}^{n-1} 1/j$, known as the harmonic descent chain. While it is known that this sequence converges to an explicit limit $x$, not much is known about the rate of convergence. We first show that a class of recursive sequences including the above are decreasing and use this to bound the rate of convergence. Moreover, for the harmonic descent chain we prove the asymptotic $x_n - x = n^{-γ_* + o(1)}$ for an implicit exponent $γ_*$. As a consequence, we deduce central limit theorems for various statistics of the critical beta-splitting random tree. This answers a number of questions of Aldous, Janson, and Pittel.
Noise Sensitivity and Learning Lower Bounds for Hierarchical Functions
Recent works explore deep learning's success by examining functions or data with hierarchical structure. To study the learning complexity of functions with hierarchical structure, we study the noise stability of functions with tree hierarchical structure on independent inputs. We show that if each function in the hierarchy is $\varepsilon$-far from linear, the noise stability is exponentially small in the depth of the hierarchy.
Our results have immediate applications for agnostic learning. In the Boolean setting using the results of Dachman-Soled, Feldman, Tan, Wan and Wimmer (2014), our results provide Statistical Query super-polynomial lower bounds for agnostically learning classes that are based on hierarchical functions.
We also derive similar SQ lower bounds based on the indicators of crossing events in critical site percolation. These crossing events are not formally hierarchical as we define but still have some hierarchical features as studied in mathematical physics.
Using the results of Abbe, Bengio, Cornacchiam, Kleinberg, Lotfi, Raghu and Zhang (2022), our results imply sample complexity lower bounds for learning hierarchical functions with gradient descent on fully connected neural networks.
Finally in the Gaussian setting, using the results of Diakonikolas, Kane, Pittas and Zarifis (2021), our results provide super-polynomial lower bounds for agnostic SQ learning.
Random Subwords and Billiard Walks in Affine Weyl Groups
Let $W$ be an irreducible affine Weyl group, and let $\mathsf{b}$ be a finite word over the alphabet of simple reflections of $W$. Fix a probability $p\in(0,1)$. For each integer $K\geq 0$, let $\mathsf{sub}_p(\mathsf{b}^K)$ be the random subword of $\mathsf{b}^K$ obtained by deleting each letter independently with probability $1-p$. Let $v_p(\mathsf{b}^K)$ be the element of $W$ represented by $\mathsf{sub}_p(\mathsf{b}^K)$. One can view $v_p(\mathsf{b}^K)$ geometrically as a random alcove; in many cases, this alcove can be seen as the location after a certain amount of time of a random billiard trajectory that, upon hitting a hyperplane in the Coxeter arrangement of $W$, reflects off of the hyperplane with probability $1-p$. We show that the asymptotic distribution of $v_p(\mathsf{b}^K)$ is a central spherical multivariate normal distribution with some variance $σ_{\mathsf{b}}^2$ depending on $\mathsf{b}$ and $p$. We provide a formula to compute $σ_{\mathsf{b}}^2$ that is remarkably simple when $\mathsf{b}$ contains only one occurrence of the simple reflection that is not in the associated finite Weyl group. As a corollary, we provide an asymptotic formula for $\mathbb{E}[\ell(v_p(\mathsf{b}^K))]$, the expected Coxeter length of $v_p(\mathsf{b}^K)$. For example, when $W=\widetilde A_{r}$ and $\mathsf{b}$ contains each simple reflection exactly once, we find that \[\lim_{K\to\infty}\frac{1}{\sqrt{K}}\mathbb{E}[\ell(v_p(\mathsf{b}^K))]=\sqrt{\frac{2}πr(r+1)\frac{p}{1-p}}.\]
Multiplayer Games of War
A recent paper by Bhatia, Chin, Mani, and Mossel (2026) defined stochastic processes aimed at modeling the game of War for {\em two players} with $n$ cards. That paper showed that these models, assuming uniform random decks, are equivalent to the Gambler's Ruin problem and therefore have an expected termination time of $Θ(n^2)$. In this paper, we generalize these models to {\em any number of players} $m$. We prove that the game with $m$ players is equivalent to a sticky random walk on an $m$-simplex; therefore, the termination time is the same as the absorption time of the sticky random walk. Interestingly, it seems that this absorption time has not been analyzed before. We show that the absorption time of the walk and the termination time of the game are both $Θ(n^2)$ for any number of players.
Weak recovery, hypothesis testing, and mutual information in stochastic block models and planted factor graphs
The stochastic block model is a canonical model of communities in random graphs. It was introduced in the social sciences and statistics as a model of communities, and in theoretical computer science as an average case model for graph partitioning problems under the name of the ``planted partition model.'' Given a sparse stochastic block model, the two standard inference tasks are: (i) Weak recovery: can we estimate the communities with non trivial overlap with the true communities? (ii) Detection/Hypothesis testing: can we distinguish if the sample was drawn from the block model or from a random graph with no community structure with probability tending to $1$ as the graph size tends to infinity?
In this work, we show that for sparse stochastic block models, the two inference tasks are equivalent except at a critical point. That is, weak recovery is information theoretically possible if and only if detection is possible. We thus find a strong connection between these two notions of inference for the model. We further prove that when detection is impossible, an explicit hypothesis test based on low degree polynomials in the adjacency matrix of the observed graph achieves the optimal statistical power. This low degree test is efficient as opposed to the likelihood ratio test, which is not known to be efficient. Moreover, we prove that the asymptotic mutual information between the observed network and the community structure exhibits a phase transition at the weak recovery threshold.
Our results are proven in much broader settings including the hypergraph stochastic block models and general planted factor graphs. In these settings we prove that the impossibility of weak recovery implies contiguity and provide a condition which guarantees the equivalence of weak recovery and detection.
Monotonicity, Topology, and Convexity of Recurrence in Random Walks
We consider non-homogeneous random walks on the two-dimensional positive quadrant $\mathbb{N}^2$ and the one-dimensional slab $\{0,1,\dots,k\}\times\mathbb{N}$. In the 1960's the following question was asked for $\mathbb{N}^2$: is it true if such a random walk $X$ is recurrent and $Y$ is another random walk that at every point is more likely to go down and more likely to go left than $Y$, then $Y$ is also recurrent?
We provide an example showing that the answer is negative. We also show, via a coupling argument, that if either the random walk $X$ or $Y$ is sufficiently homogeneous then the answer is in fact positive. In addition, we show using the Rayleigh monotonicity principle that the analogous question for random walks on trees is positive.
These results show that the subset of parameter space that yields recurrent random walks possesses some geometric properties, in this case the structure of an order ideal. Motivated by this perspective, we consider the more symmetric setting of homogeneous random walks on finitely generated abelian groups, and ask when this subset possesses other geometric properties, namely various topological properties and convexity. We answer some of these questions: in particular, we show that this subset is closed, and under a symmetric support condition, show it is path-connected and additionally show it is convex if and only if its effective dimension is at most 2. We also show its complement is in some sense typically path-connected but not convex. We finally propose some related open problems.
Influences in Mixing Measures
The theory of influences in product measures has profound applications in theoretical computer science, combinatorics, and discrete probability. This deep theory is intimately connected to functional inequalities and to the Fourier analysis of discrete groups. Originally, influences of functions were motivated by the study of social choice theory, wherein a Boolean function represents a voting scheme, its inputs represent the votes, and its output represents the outcome of the elections. Thus, product measures represent a scenario in which the votes of the parties are randomly and independently distributed, which is often far from the truth in real-life scenarios.
We begin to develop the theory of influences for more general measures under mixing or correlation decay conditions. More specifically, we prove analogues of the KKL and Talagrand influence theorems for Markov Random Fields on bounded degree graphs with correlation decay. We show how some of the original applications of the theory of in terms of voting and coalitions extend to general measures with correlation decay. Our results thus shed light both on voting with correlated voters and on the behavior of general functions of Markov Random Fields (also called ``spin-systems") with correlation decay.
When will (game) wars end?
We study several variants of the classical card game war. As anyone who played this game knows, the game can take some time to terminate, but it usually does. Here, we analyze a number of asymptotic variants of the game, where the number of cards is $n$, and show that all have expected termination time of order $n^2$. This is the same expected termination time as in the game where at each turn a fair coin toss decides which player wins a card, known as Gambler's Ruin and studied by Pascal, Fermat and others in the seventeenth century.
Exact Phase Transitions for Stochastic Block Models and Reconstruction on Trees
Published
• View Publication
• BIB
In this paper we continue to rigorously establish the predictions in ground breaking work in statistical physics by Decelle, Krzakala, Moore, Zdeborová (2011) regarding the block model, in particular in the case of $q=3$ and $q=4$ communities.
We prove that for $q=3$ and $q=4$ there is no computational-statistical gap if the average degree is above some constant by showing it is information theoretically impossible to detect below the Kesten-Stigum bound. The proof is based on showing that for the broadcast process on Galton-Watson trees, reconstruction is impossible for $q=3$ and $q=4$ if the average degree is sufficiently large. This improves on the result of Sly (2009), who proved similar results for regular trees for $q=3$. Our analysis of the critical case $q=4$ provides a detailed picture showing that the tightness of the Kesten-Stigum bound in the antiferromagnetic case depends on the average degree of the tree. We also prove that for $q\geq 5$, the Kestin-Stigum bound is not sharp.
Our results prove conjectures of Decelle, Krzakala, Moore, Zdeborová (2011), Moore (2017), Abbe and Sandon (2018) and Ricci-Tersenghi, Semerjian, and Zdeborová (2019). Our proofs are based on a new general coupling of the tree and graph processes and on a refined analysis of the broadcast process on the tree.
A second moment proof of the spread lemma
Published
• View Publication
• BIB
This note concerns a well-known result which we term the ``spread lemma,'' which establishes the existence (with high probability) of a desired structure in a random set. The spread lemma was central to two recent celebrated results: (a) the improved bounds of Alweiss, Lovett, Wu, and Zhang (2019) on the Erdős-Rado sunflower conjecture; and (b) the proof of the fractional Kahn--Kalai conjecture by Frankston, Kahn, Narayanan and Park (2019). While the lemma was first proved (and later refined) by delicate counting arguments, alternative proofs have also been given, via Shannon's noiseless coding theorem (Rao, 2019), and also via manipulations of Shannon entropy bounds (Tao, 2020).
In this note we present a new proof of the spread lemma, that takes advantage of an explicit recasting of the proof in the language of Bayesian statistical inference. We show that from this viewpoint the proof proceeds in a straightforward and principled probabilistic manner, leading to a truncated second moment calculation which concludes the proof. The proof can also be viewed as a demonstration of the ``planting trick'' introduced by Achlioptas and Coga-Oghlan (2008) in the study of random constraint satisfaction problems.
On the Second Kahn--Kalai Conjecture
For any given graph $H$, we are interested in $p_\mathrm{crit}(H)$, the minimal $p$ such that the Erdős-Rényi graph $G(n,p)$ contains a copy of $H$ with probability at least $1/2$. Kahn and Kalai (2007) conjectured that $p_\mathrm{crit}(H)$ is given up to a logarithmic factor by a simpler "subgraph expectation threshold" $p_\mathrm{E}(H)$, which is the minimal $p$ such that for every subgraph $H'\subseteq H$, the Erdős-Rényi graph $G(n,p)$ contains \emph{in expectation} at least $1/2$ copies of $H'$. It is trivial that $p_\mathrm{E}(H) \le p_\mathrm{crit}(H)$, and the so-called "second Kahn-Kalai conjecture" states that $p_\mathrm{crit}(H) \lesssim p_\mathrm{E}(H) \log e(H)$ where $e(H)$ is the number of edges in $H$.
In this article, we present a natural modification $p_\mathrm{E, new}(H)$ of the Kahn--Kalai subgraph expectation threshold, which we show is sandwiched between $p_\mathrm{E}(H)$ and $p_\mathrm{crit}(H)$. The new definition $p_\mathrm{E, new}(H)$ is based on the simple observation that if $G(n,p)$ contains a copy of $H$ and $H$ contains \emph{many} copies of $H'$, then $G(n,p)$ must also contain \emph{many} copies of $H'$. We then show that $p_\mathrm{crit}(H) \lesssim p_\mathrm{E, new}(H) \log e(H)$, thus proving a modification of the second Kahn--Kalai conjecture. The bound follows by a direct application of the set-theoretic "spread" property, which led to recent breakthroughs in the sunflower conjecture by Alweiss, Lovett, Wu and Zhang and the first fractional Kahn--Kalai conjecture by Frankston, Kahn, Narayanan and Park.
Approximate polymorphisms
Published
• View Publication
• BIB
For a function $g\colon\{0,1\}^m\to\{0,1\}$, a function $f\colon \{0,1\}^n\to\{0,1\}$ is called a $g$-polymorphism if their actions commute: $f(g(\mathsf{row}_1(Z)),\ldots,g(\mathsf{row}_n(Z))) = g(f(\mathsf{col}_1(Z)),\ldots,f(\mathsf{col}_m(Z)))$ for all $Z\in\{0,1\}^{n\times m}$. The function $f$ is called an approximate polymorphism if this equality holds with probability close to $1$, when $Z$ is sampled uniformly.
We study the structure of exact polymorphisms as well as approximate polymorphisms. Our results include:
- We prove that an approximate polymorphism $f$ must be close to an exact polymorphism;
- We give a characterization of exact polymorphisms, showing that besides trivial cases, only the functions $g = \mathsf{AND}, \mathsf{XOR}, \mathsf{OR}, \mathsf{NXOR}$ admit non-trivial exact polymorphisms.
We also study the approximate polymorphism problem in the list-decoding regime (i.e., when the probability equality holds is not close to $1$, but is bounded away from some value). We show that if $f(x \land y) = f(x) \land f(y)$ with probability larger than $s_\land \approx 0.815$ then $f$ correlates with some low-degree character, and $s_\land$ is the optimal threshold for this property.
Our result generalize the classical linearity testing result of Blum, Luby and Rubinfeld, that in this language showed that the approximate polymorphisms of $g = \mathsf{XOR}$ are close to XOR's, as well as a recent result of Filmus, Lifshitz, Minzer and Mossel, showing that the approximate polymorphisms of AND can only be close to AND functions.
A Phase Transition in Arrow's Theorem
Published
• View Publication
• BIB
Arrow's Theorem concerns a fundamental problem in social choice theory: given the individual preferences of members of a group, how can they be aggregated to form rational group preferences? Arrow showed that in an election between three or more candidates, there are situations where any voting rule satisfying a small list of natural "fairness" axioms must produce an apparently irrational intransitive outcome. Furthermore, quantitative versions of Arrow's Theorem in the literature show that when voters choose rankings in an i.i.d.\ fashion, the outcome is intransitive with non-negligible probability.
It is natural to ask if such a quantitative version of Arrow's Theorem holds for non-i.i.d.\ models. To answer this question, we study Arrow's Theorem under a natural non-i.i.d.\ model of voters inspired by canonical models in statistical physics; indeed, a version of this model was previously introduced by Raffaelli and Marsili in the physics literature. This model has a parameter, temperature, that prescribes the correlation between different voters. We show that the behavior of Arrow's Theorem in this model undergoes a striking phase transition: in the entire high temperature regime of the model, a Quantitative Arrow's Theorem holds showing that the probability of paradox for any voting rule satisfying the axioms is non-negligible; this is tight because the probability of paradox under pairwise majority goes to zero when approaching the critical temperature, and becomes exponentially small in the number of voters beyond it. We prove this occurs in another natural model of correlated voters and conjecture this phenomena is quite general.
AND Testing and Robust Judgement Aggregation
Published
• View Publication
• BIB
A function $f\colon\{0,1\}^n\to \{0,1\}$ is called an approximate AND-homomorphism if choosing ${\bf x},{\bf y}\in\{0,1\}^n$ randomly, we have that $f({\bf x}\land {\bf y}) = f({\bf x})\land f({\bf y})$ with probability at least $1-ε$, where $x\land y = (x_1\land y_1,\ldots,x_n\land y_n)$. We prove that if $f\colon \{0,1\}^n \to \{0,1\}$ is an approximate AND-homomorphism, then $f$ is $δ$-close to either a constant function or an AND function, where $δ(ε) \to 0$ as $ε\to0$. This improves on a result of Nehama, who proved a similar statement in which $δ$ depends on $n$.
Our theorem implies a strong result on judgement aggregation in computational social choice. In the language of social choice, our result shows that if $f$ is $ε$-close to satisfying judgement aggregation, then it is $δ(ε)$-close to an oligarchy (the name for the AND function in social choice theory). This improves on Nehama's result, in which $δ$ decays polynomially with $n$.
Our result follows from a more general one, in which we characterize approximate solutions to the eigenvalue equation $\mathrm T f = λg$, where $\mathrm T$ is the downwards noise operator $\mathrm T f(x) = \mathbb{E}_{\bf y}[f(x \land {\bf y})]$, $f$ is $[0,1]$-valued, and $g$ is $\{0,1\}$-valued. We identify all exact solutions to this equation, and show that any approximate solution in which $\mathrm T f$ and $λg$ are close is close to an exact solution.
Regular graphs with linearly many triangles
A $d$-regular graph on $n$ nodes has at most $T_{\max} = \frac{n}{3} \tbinom{d}{2}$ triangles. We compute the leading asymptotics of the probability that a large random $d$-regular graph has at least $c \cdot T_{\max}$ triangles, and provide a strong structural description of such graphs.
When $d$ is fixed, we show that such graphs typically consist of many disjoint $d+1$-cliques and an almost triangle-free part. When $d$ is allowed to grow with $n$, we show that such graphs typically consist of $d+o(d)$ sized almost cliques together with an almost triangle-free part.
This confirms a conjecture of Collet and Eckmann from 2002 and considerably strengthens their observation that the triangles cannot be totally scattered in typical instances of regular graphs with many triangles.
The Mean-Field Approximation: Information Inequalities, Algorithms, and Complexity
The mean field approximation to the Ising model is a canonical variational tool that is used for analysis and inference in Ising models. We provide a simple and optimal bound for the KL error of the mean field approximation for Ising models on general graphs, and extend it to higher order Markov random fields. Our bound improves on previous bounds obtained in work in the graph limit literature by Borgs, Chayes, Lovász, Sós, and Vesztergombi and another recent work by Basak and Mukherjee. Our bound is tight up to lower order terms. Building on the methods used to prove the bound, along with techniques from combinatorics and optimization, we study the algorithmic problem of estimating the (variational) free energy for Ising models and general Markov random fields. For a graph $G$ on $n$ vertices and interaction matrix $J$ with Frobenius norm $\| J \|_F$, we provide algorithms that approximate the free energy within an additive error of $εn \|J\|_F$ in time $\exp(poly(1/ε))$. We also show that approximation within $(n \|J\|_F)^{1-δ}$ is NP-hard for every $δ> 0$. Finally, we provide more efficient approximation algorithms, which find the optimal mean field approximation, for ferromagnetic Ising models and for Ising models satisfying Dobrushin's condition.
The Vertex Sample Complexity of Free Energy is Polynomial
We study the following question: given a massive Markov random field on $n$ nodes, can a small sample from it provide a rough approximation to the free energy $\mathcal{F}_n = \log{Z_n}$?
Results in graph limit literature by Borgs, Chayes, Lovász, Sós, and Vesztergombi show that for Ising models on $n$ nodes and interactions of strength $Θ(1/n)$, an $ε$ approximation to $\log Z_n / n$ can be achieved by sampling a randomly induced model on $2^{O(1/ε^2)}$ nodes. We show that the sampling complexity of this problem is {\em polynomial in} $1/ε$. We further show a polynomial dependence on $ε$ cannot be avoided.
Our results are very general as they apply to higher order Markov random fields. For Markov random fields of order $r$, we obtain an algorithm that achieves $ε$ approximation using a number of samples polynomial in $r$ and $1/ε$ and running time that is $2^{O(1/ε^2)}$ up to polynomial factors in $r$ and $ε$. For ferromagnetic Ising models, the running time is polynomial in $1/ε$.
Our results are intimately connected to recent research on the regularity lemma and property testing, where the interest is in finding which properties can tested within $ε$ error in time polynomial in $1/ε$. In particular, our proofs build on results from a recent work by Alon, de la Vega, Kannan and Karpinski, who also introduced the notion of polynomial vertex sample complexity. Another critical ingredient of the proof is an effective bound by the authors of the paper relating the variational free energy and the free energy.
Gaussian Bounds for Noise Correlation of Resilient Functions
Published
• View Publication
• BIB
Gaussian bounds on noise correlation of functions play an important role in hardness of approximation, in quantitative social choice theory and in testing. The author (2008) obtained sharp gaussian bounds for the expected correlation of $\ell$ low influence functions $f^{(1)},\ldots, f^{(\ell)} : Ω^n \to [0,1]$, where the inputs to the functions are correlated via the $n$-fold tensor of distribution $\mathcal{P}$ on $Ω^{\ell}$.
It is natural to ask if the condition of low influences can be relaxed to the condition that the function has vanishing Fourier coefficients. Here we answer this question affirmatively. For the case of two functions $f$ and $g$, we further show that if $f,g$ have a noisy inner product that exceeds the gaussian bound, then the Fourier supports of their large coefficients intersect.
Shotgun Assembly of Random Jigsaw Puzzles
Published
• View Publication
• BIB
In a recent work, Mossel and Ross considered the shotgun assembly problem for a random jigsaw puzzle. Their model consists of a puzzle - an $n\times n$ grid, where each vertex is viewed as a center of a piece. They assume that each of the four edges adjacent to a vertex, is assigned one of $q$ colors (corresponding to "jigs", or cut shapes) uniformly at random. Mossel and Ross asked: how large should $q = q(n)$ be so that with high probability the puzzle can be assembled uniquely given the collection of individual tiles? They showed that if $q = ω(n^2)$, then the puzzle can be assembled uniquely with high probability, while if $q = o(n^{2/3})$, then with high probability the puzzle cannot be uniquely assembled. Here we improve the upper bound and show that for any $\eps > 0$, the puzzle can be assembled uniquely with high probability if $q \geq n^{1+\eps}$. The proof uses an algorithm of $n^{Θ(1/\eps)}$ running time.
On the Correlation of Increasing Families
Published
• View Publication
• BIB
The classical correlation inequality of Harris asserts that any two monotone increasing families on the discrete cube are nonnegatively correlated. In 1996, Talagrand established a lower bound on the correlation in terms of how much the two families depend simultaneously on the same coordinates. Talagrand's method and results inspired a number of important works in combinatorics and probability theory.
In this paper we present stronger correlation lower bounds that hold when the increasing families satisfy natural regularity or symmetry conditions. In addition, we present several new classes of examples for which Talagrand's bound is tight.
A central tool in the paper is a simple lemma asserting that for monotone events noise decreases correlation. This lemma gives also a very simple derivation of the classical FKG inequality for product measures, and leads to a simplification of part of Talagrand's proof.