Papers by Elchanan Mossel
52 paper(s) by this author
· All BibTeX
Evolutionary Trees and the Ising Model on the Bethe Lattice: a Proof of Steel's Conjecture
Published
• View Publication
• BIB
A major task of evolutionary biology is the reconstruction of phylogenetic trees from molecular data. The evolutionary model is given by a Markov chain on a tree. Given samples from the leaves of the Markov chain, the goal is to reconstruct the leaf-labelled tree.
It is well known that in order to reconstruct a tree on $n$ leaves, sample sequences of length $Ω(\log n)$ are needed. It was conjectured by M. Steel that for the CFN/Ising evolutionary model, if the mutation probability on all edges of the tree is less than $p^{\ast} = (\sqrt{2}-1)/2^{3/2}$, then the tree can be recovered from sequences of length $O(\log n)$. The value $p^{\ast}$ is given by the transition point for the extremality of the free Gibbs measure for the Ising model on the binary tree. Steel's conjecture was proven by the second author in the special case where the tree is "balanced." The second author also proved that if all edges have mutation probability larger than $p^{\ast}$ then the length needed is $n^{Ω(1)}$. Here we show that Steel's conjecture holds true for general trees by giving a reconstruction algorithm that recovers the tree from $O(\log n)$-length sequences when the mutation probabilities are discretized and less than $p^\ast$. Our proof and results demonstrate that extremality of the free Gibbs measure on the infinite binary tree, which has been studied before in probability, statistical physics and computer science, determines how distinguishable are Gibbs measures on finite binary trees.
Noise stability of functions with low influences: invariance and optimality
Published
• View Publication
• BIB
In this paper we study functions with low influences on product probability spaces. The analysis of boolean functions with low influences has become a central problem in discrete Fourier analysis. It is motivated by fundamental questions arising from the construction of probabilistically checkable proofs in theoretical computer science and from problems in the theory of social choice in economics.
We prove an invariance principle for multilinear polynomials with low influences and bounded degree; it shows that under mild conditions the distribution of such polynomials is essentially invariant for all product spaces. Ours is one of the very few known non-linear invariance principles. It has the advantage that its proof is simple and that the error bounds are explicit. We also show that the assumption of bounded degree can be eliminated if the polynomials are slightly ``smoothed''; this extension is essential for our applications to ``noise stability''-type problems.
In particular, as applications of the invariance principle we prove two conjectures: the ``Majority Is Stablest'' conjecture from theoretical computer science, which was the original motivation for this work, and the ``It Ain't Over Till It's Over'' conjecture from social choice theory.
Non-interactive correlation distillation, inhomogeneous Markov chains, and the reverse Bonami-Beckner inequality
Published
• View Publication
• BIB
In this paper we study non-interactive correlation distillation (NICD), a generalization of the study of noise sensitivity of boolean functions. We extend the model to NICD on trees. In this model there is a fixed undirected tree with players at some of the nodes. One node is given a uniformly random string and this string is distributed throughout the network, with the edges of the tree acting as independent binary symmetric channels. The goal of the players is to agree on a shared random bit without communicating. Our new contributions include the following: 1. In the case of a k-leaf star graph, we resolve the open question of whether the success probability must go to zero as k goes to infinity. We show that this is indeed the case and provide matching upper and lower bounds on the asymptotically optimal rate (a slowly-decaying polynomial). 2. In the case of the k-vertex path graph, we show that it is always optimal for all players to use the same 1-bit function. 3. In the general case we show that all players should use monotone functions. 4. For certain trees it is better if not all players use the same function. Our techniques include the use of the reverse Bonami-Beckner inequality.
Random autocatalytic networks
We determine conditions under which a random biochemical system is likely to contain a subsystem that is both autocatalytic and able to survive on some ambient `food' source. Such systems have previously been investigated for their relevance to origin-of-life models. In this paper we extend earlier work, by finding precisely the order of catalysation required for the emergence of such self-sustaining autocatalytic networks. This answers questions raised in earlier papers, yet also allows for a more general class of models. We also show that a recently-described polynomial time algorithm for determining whether a catalytic reaction system contains an autocatalytic, self-sustaining subsystem is unlikely to adapt to allow inhibitory catalysation - in this case we show that the associated decision problem is NP-complete.
Coin flipping from a cosmic source: On error correction of truly random bits
Published
• View Publication
• BIB
We study a problem related to coin flipping, coding theory, and noise sensitivity. Consider a source of truly random bits $x \in \bits^n$, and $k$ parties, who have noisy versions of the source bits $y^i \in \bits^n$, where for all $i$ and $j$, it holds that $\Pr[y^i_j = x_j] = 1 - \eps$, independently for all $i$ and $j$. That is, each party sees each bit correctly with probability $1-ε$, and incorrectly (flipped) with probability $ε$, independently for all bits and all parties. The parties, who cannot communicate, wish to agree beforehand on {\em balanced} functions $f_i : \bits^n \to \bits$ such that $\Pr[f_1(y^1) = ... = f_k(y^k)]$ is maximized. In other words, each party wants to toss a fair coin so that the probability that all parties have the same coin is maximized. The functions $f_i$ may be thought of as an error correcting procedure for the source $x$.
When $k=2,3$ no error correction is possible, as the optimal protocol is given by $f_i(x^i) = y^i_1$. On the other hand, for large values of $k$, better protocols exist. We study general properties of the optimal protocols and the asymptotic behavior of the problem with respect to $k$, $n$ and $\eps$. Our analysis uses tools from probability, discrete Fourier analysis, convexity and discrete symmetrization.
A Law of Large Numbers for Weighted Majority
Consider an election between two candidates in which the voters' choices are random and independent and the probability of a voter choosing the first candidate is $p>1/2$. Condorcet's Jury Theorem which he derived from the weak law of large numbers asserts that if the number of voters tends to infinity then the probability that the first candidate will be elected tends to one. The notion of influence of a voter or its voting power is relevant for extensions of the weak law of large numbers for voting rules which are more general than simple majority. In this paper we point out two different ways to extend the classical notions of voting power and influences to arbitrary probability distributions. The extension relevant to us is the ``effect'' of a voter, which is a weighted version of the correlation between the voter's vote and the election's outcomes. We prove an extension of the weak law of large numbers to weighted majority games when all individual effects are small and show that this result does not apply to any voting rule which is not based on weighted majority.
Robust reconstruction on trees is determined by the second eigenvalue
Published in Annals of Probability 2004, Vol. 32, No. 3, 2630-2649
• View Publication
• BIB
Consider a Markov chain on an infinite tree T=(V,E) rooted at ρ. In such a chain, once the initial root state σ(ρ) is chosen, each vertex iteratively chooses its state from the one of its parent by an application of a Markov transition rule (and all such applications are independent). Let μ_j denote the resulting measure for σ(ρ)=j. The resulting measure μ_j is defined on configurations σ=(σ(x))_{x\in V}\in A^V, where A is some finite set. Let μ_j^n denote the restriction of μto the sigma-algebra generated by the variables σ(x), where x is at distance exactly n from ρ. Letting α_n=max_{i,j\in A}d_{TV}(μ_i^n,μ_j^n), where d_{TV} denotes total variation distance, we say that the reconstruction problem is solvable if lim inf_{n\to\infty}α_n>0. Reconstruction solvability roughly means that the nth level of the tree contains a nonvanishing amount of information on the root of the tree as n\to\infty. In this paper we study the problem of robust reconstruction. Let νbe a nondegenerate distribution on A and ε>0. Let σbe chosen according to μ_j^n and σ' be obtained from σby letting for each node independently, σ(v)=σ'(v) with probability 1-εand σ'(v) be an independent sample from νotherwise. We denote by μ_j^n[ν,ε] the resulting measure on σ'. The measure μ_j^n[ν,ε] is a perturbation of the measure μ_j^n.
Shuffling by semi-random transpositions
Published
• View Publication
• BIB
In the cyclic-to-random shuffle, we are given n cards arranged in a circle. At step k, we exchange the k'th card along the circle with a uniformly chosen random card. The problem of determining the mixing time of the cyclic-to-random shuffle was raised by Aldous and Diaconis in 1986. Recently, Mironov used this shuffle as a model for the cryptographic system known as ``RC4'' and proved an upper bound of O(n log n) for the mixing time. We prove a matching lower bound, thus establishing that the mixing time is indeed of order $Θ(n \log n)$. We also prove an upper bound of O(n log n) for the mixing time of any ``semi-random transposition shuffle'', i.e., any shuffle in which a random card is exchanged with another card chosen according to an arbitrary (deterministic or random) rule. To prove our lower bound, we exhibit an explicit complex-valued test function which typically takes very different values for permutations arising from the cyclic-to-random-shuffle and for uniform random permutations; we expect that this test function may be useful in future analysis of RC4. Perhaps surprisingly, the proof hinges on the fact that the function exp(z)-1 has nonzero fixed points in the complex plane. A key insight from our work is the importance of complex analysis tools for uncovering structure in nonreversible Markov chains.
Distorted metrics on trees and phylogenetic forests
Published
• View Publication
• BIB
We study distorted metrics on binary trees in the context of phylogenetic reconstruction. Given a binary tree $T$ on $n$ leaves with a path metric $d$, consider the pairwise distances $\{d(u,v)\}$ between leaves. It is well known that these determine the tree and the $d$ length of all edges. Here we consider distortions $\d$ of $d$ such that for all leaves $u$ and $v$ it holds that $|d(u,v) - \d(u,v)| < f/2$ if either $d(u,v) < M$ or $\d(u,v) < M$, where $d$ satisfies $f \leq d(e) \leq g$ for all edges $e$. Given such distortions we show how to reconstruct in polynomial time a forest $T_1,...,T_α$ such that the true tree $T$ may be obtained from that forest by adding $α-1$ edges and $α-1 \leq 2^{-Ω(M/g)} n$.
Metric distortions arise naturally in phylogeny, where $d(u,v)$ is defined by the log-det of a covariance matrix associated with $u$ and $v$. of a covariance matrix associated with $u$ and $v$. When $u$ and $v$ are ``far'', the entries of the covariance matrix are small and therefore $\d(u,v)$, which is defined by log-det of an associated empirical-correlation matrix may be a bad estimate of $d(u,v)$ even if the correlation matrix is ``close'' to the covariance matrix.
Our metric results are used in order to show how to reconstruct phylogenetic forests with small number of trees from sequences of length logarithmic in the size of the tree. Our method also yields an independent proof that phylogenetic trees can be reconstructed in polynomial time from sequences of polynomial length under the standard assumptions in phylogeny. Both the metric result and its applications to phylogeny are almost tight.
New coins from old: computing with unknown bias
Published
• View Publication
• BIB
Suppose that we are given a function f : (0,1) -> (0,1) and, for some unknown p in (0,1), a sequence of independent tosses of a p-coin (i.e., a coin with probability p of ``heads'').
For which functions f is it possible to simulate an f(p)-coin?; This question was raised by S. Asmussen and J. Propp. A simple simulation scheme for the constant function 1/2 was described by von Neumann (1951); this scheme can be easily implemented using a finite automaton. We prove that in general, an f(p)-coin can be simulated by a finite automaton for all p in (0,1), if and only if f is a rational function over Q. We also show that if an f(p)-coin can be simulated by a pushdown automaton, then f is an algebraic function over Q; however, pushdown automata can simulate f(p)-coins for certain non-rational functions such as the square root of p. These results complement the work of Keane and O'Brien (1994), who determined the functions $f$ for which an f(p)-coin can be simulated when there are no computational restrictions on the simulation scheme.
Information flow on trees
Published
• View Publication
• BIB
Consider a tree network $T$, where each edge acts as an independent copy of a given channel $M$, and information is propagated from the root. For which $T$ and $M$ does the configuration obtained at level $n$ of $T$ typically contain significant information on the root variable? This problem arose independently in biology, information theory and statistical physics.
For all $b$, we construct a channel for which the variable at the root of the $b$-ary tree is independent of the configuration at level 2 of that tree, yet for sufficiently large $B>b$, the mutual information between the configuration at level $n$ of the $B$-ary tree and the root variable is bounded away from zero. This is related to certain secret-sharing protocols.
We improve the upper bounds on information flow for asymmetric binary channels (which correspond to the Ising model with an external field) and for symmetric $q$-ary channels (which correspond to Potts models).
Let $\lam_2(M)$ denote the second largest eigenvalue of $M$, in absolute value. A CLT of Kesten and Stigum~(1966) implies that if $b |\lam_2(M)|^2 >1$, then the {\em census} of the variables at any level of the $b$-ary tree, contains significant information on the root variable. We establish a converse: if $b |\lam_2(M)|^2 < 1$, then the census of the variables at level $n$ of the $b$-ary tree is asymptotically independent of the root variable. This contrasts with examples where $b |\lam_2(M)|^2 <1$, yet the {\em configuration} at level $n$ is not asymptotically independent of the root variable.
On the mixing time of simple random walk on the super critical percolation cluster
Published
• View Publication
• BIB
We study the robustness under perturbations of mixing times, by studying mixing times of random walks in percolation clusters inside boxes in $\Z^d$. We show that for $d \geq 2$ and $p > p_c(\Z^d)$, the mixing time of simple random walk on the largest cluster inside $\{-n,...,n\}^d$ is $Θ(n^2)$ - thus the mixing time is robust up to constant factor.