arXiv++ Combinatorics

Browse math.CO papers from arXiv

Papers by Amir Yehudayoff

21 paper(s) by this author · All BibTeX
2025-09-01
A LYM Inequality For Product Measures
This note proves a version of Lubell-Yamamoto-Meshalkin inequality for general product measures.
2025-07-10
Approximation Depth of Convex Polytopes
We study approximations of polytopes in the standard model for computing polytopes using Minkowski sums and (convex hulls of) unions. Specifically, we study the ability to approximate a target polytope by polytopes of a given depth. Our main results imply that simplices can only be ``trivially approximated''. On the way, we obtain a characterization of simplices as the only ``outer additive'' convex bodies.
Better Neural Network Expressivity: Subdividing the Simplex
This work studies the expressivity of ReLU neural networks with a focus on their depth. A sequence of previous works showed that $\lceil \log_2(n+1) \rceil$ hidden layers are sufficient to compute all continuous piecewise linear (CPWL) functions on $\mathbb{R}^n$. Hertrich, Basu, Di Summa, and Skutella (NeurIPS'21 / SIDMA'23) conjectured that this result is optimal in the sense that there are CPWL functions on $\mathbb{R}^n$, like the maximum function, that require this depth. We disprove the conjecture and show that $\lceil\log_3(n-1)\rceil+1$ hidden layers are sufficient to compute all CPWL functions on $\mathbb{R}^n$. A key step in the proof is that ReLU neural networks with two hidden layers can exactly represent the maximum function of five inputs. More generally, we show that $\lceil\log_3(n-2)\rceil+1$ hidden layers are sufficient to compute the maximum of $n\geq 4$ numbers. Our constructions almost match the $\lceil\log_3(n)\rceil$ lower bound of Averkov, Hojny, and Merkert (ICLR'25) in the special case of ReLU networks with weights that are decimal fractions. The constructions have a geometric interpretation via polyhedral subdivisions of the simplex into ``easier'' polytopes.
On the Depth of Monotone ReLU Neural Networks and ICNNs
We study two models of ReLU neural networks: monotone networks (ReLU$^+$) and input convex neural networks (ICNN). Our focus is on expressivity, mostly in terms of depth, and we prove the following lower bounds. For the maximum function MAX$_n$ computing the maximum of $n$ real numbers, we show that ReLU$^+$ networks cannot compute MAX$_n$, or even approximate it. We prove a sharp $n$ lower bound on the ICNN depth complexity of MAX$_n$. We also prove depth separations between ReLU networks and ICNNs; for every $k$, there is a depth-2 ReLU network of size $O(k^2)$ that cannot be simulated by a depth-$k$ ICNN. The proofs are based on deep connections between neural networks and polyhedral geometry, and also use isoperimetric properties of triangulations.
2023-09-15
The discrepancy of greater-than
The discrepancy of the $n \times n$ greater-than matrix is shown to be $\fracπ{2 \ln n}$ up to lower order terms.
2023-04-11 v2
Dual Systolic Graphs
We define a family of graphs we call dual systolic graphs. This definition comes from graphs that are duals of systolic simplicial complexes. Our main result is a sharp (up to constants) isoperimetric inequality for dual systolic graphs. The first step in the proof is an extension of the classical isoperimetric inequality of the boolean cube. The isoperimetric inequality for dual systolic graphs, however, is exponentially stronger than the one for the boolean cube. Interestingly, we know that dual systolic graphs exist, but we do not yet know how to efficiently construct them. We, therefore, define a weaker notion of dual systolicity. We prove the same isoperimetric inequality for weakly dual systolic graphs, and at the same time provide an efficient construction of a family of graphs that are weakly dual systolic. We call this family of graphs clique products. We show that there is a non-trivial connection between the small set expansion capabilities and the threshold rank of clique products, and believe they can find further applications.
2023-04-07 v2
Replicability and stability in learning
Replicability is essential in science as it allows us to validate and verify research findings. Impagliazzo, Lei, Pitassi and Sorrell (`22) recently initiated the study of replicability in machine learning. A learning algorithm is replicable if it typically produces the same output when applied on two i.i.d. inputs using the same internal randomness. We study a variant of replicability that does not involve fixing the randomness. An algorithm satisfies this form of replicability if it typically produces the same output when applied on two i.i.d. inputs (without fixing the internal randomness). This variant is called global stability and was introduced by Bun, Livni and Moran ('20) in the context of differential privacy. Impagliazzo et al. showed how to boost any replicable algorithm so that it produces the same output with probability arbitrarily close to 1. In contrast, we demonstrate that for numerous learning tasks, global stability can only be accomplished weakly, where the same output is produced only with probability bounded away from 1. To overcome this limitation, we introduce the concept of list replicability, which is equivalent to global stability. Moreover, we prove that list replicability can be boosted so that it is achieved with probability arbitrarily close to 1. We also describe basic relations between standard learning-theoretic complexity measures and list replicable numbers. Our results, in addition, imply that besides trivial cases, replicable algorithms (in the sense of Impagliazzo et al.) must be randomized. The proof of the impossibility result is based on a topological fixed-point theorem. For every algorithm, we are able to locate a "hard input distribution" by applying the Poincaré-Miranda theorem in a related topological setting. The equivalence between global stability and list replicability is algorithmic.
2021-05-28
A lower bound for essential covers of the cube
Published • View PublicationBIB
Essential covers were introduced by Linial and Radhakrishnan as a model that captures two complementary properties: (1) all variables must be included and (2) no element is redundant. In their seminal paper, they proved that every essential cover of the $n$-dimensional hypercube must be of size at least $Ω(n^{0.5})$. Later on, this notion found several applications in complexity theory. We improve the lower bound to $Ω(n^{0.52})$, and describe two applications.
2021-02-10 v2
Slicing the hypercube is not easy
We prove that at least $Ω(n^{0.51})$ hyperplanes are needed to slice all edges of the $n$-dimensional hypercube. We provide a couple of applications: lower bounds on the computational complexity of parity, and a lower bound on the cover number of the hypercube by skew hyperplanes.
2018-07-02 v3
An isoperimetric inequality for Hamming balls and local expansion in hypercubes
Published in The Electronic Journal of Combinatorics, volume 29, issue 1, #P1.15, January 2022 • View PublicationBIB
We prove a vertex isoperimetric inequality for the $n$-dimensional Hamming ball $\mathcal{B}_n(R)$ of radius $R$. The isoperimetric inequality is sharp up to a constant factor for sets that are comparable to $\mathcal{B}_n(R)$ in size. A key step in the proof is a local expansion phenomenon in hypercubes.
2017-07-17 v3
On weak $ε$-nets and the Radon number
Published • View PublicationBIB
We show that the Radon number characterizes the existence of weak nets in separable convexity spaces (an abstraction of the euclidean notion of convexity). The construction of weak nets when the Radon number is finite is based on Helly's property and on metric properties of VC classes. The lower bound on the size of weak nets when the Radon number is large relies on the chromatic number of the Kneser graph. As an application, we prove a boosting-type result for weak $ε$-nets.
2016-10-12 v2
On statistical learning via the lens of compression
This work continues the study of the relationship between sample compression schemes and statistical learning, which has been mostly investigated within the framework of binary classification. The central theme of this work is establishing equivalences between learnability and compressibility, and utilizing these equivalences in the study of statistical learning theory. We begin with the setting of multiclass categorization (zero/one loss). We prove that in this case learnability is equivalent to compression of logarithmic sample size, and that uniform convergence implies compression of constant size. We then consider Vapnik's general learning setting: we show that in order to extend the compressibility-learnability equivalence to this case, it is necessary to consider an approximate variant of compression. Finally, we provide some applications of the compressibility-learnability equivalences: (i) Agnostic-case learnability and realizable-case learnability are equivalent in multiclass categorization problems (in terms of sample complexity). (ii) This equivalence between agnostic-case learnability and realizable-case learnability does not hold for general learning problems: There exists a learning problem whose loss function takes just three values, under which agnostic-case and realizable-case learnability are not equivalent. (iii) Uniform convergence implies compression of constant size in multiclass categorization problems. Part of the argument includes an analysis of the uniform convergence rate in terms of the graph dimension, in which we improve upon previous bounds. (iv) A dichotomy for sample compression in multiclass categorization problems: If a non-trivial compression exists then a compression of logarithmic size exists. (v) A compactness theorem for multiclass categorization problems.
Geometric stability via information theory
Published • View PublicationBIB
The Loomis-Whitney inequality, and the more general Uniform Cover inequality, bound the volume of a body in terms of a product of the volumes of lower-dimensional projections of the body. In this paper, we prove stability versions of these inequalities, showing that when they are close to being tight, the body in question is close in symmetric difference to a 'box'. Our results are best possible up to a constant factor depending upon the dimension alone. Our approach is information theoretic. We use our stability result for the Loomis-Whitney inequality to obtain a stability result for the edge-isoperimetric inequality in the infinite $d$-dimensional lattice. Namely, we prove that a subset of $\mathbb{Z}^d$ with small edge-boundary must be close in symmetric difference to a $d$-dimensional cube. Our bound is, again, best possible up to a constant factor depending upon $d$ alone.
2015-03-26 v2
Sign rank versus VC dimension
Published • View PublicationBIB
This work studies the maximum possible sign rank of $N \times N$ sign matrices with a given VC dimension $d$. For $d=1$, this maximum is {three}. For $d=2$, this maximum is $\tildeΘ(N^{1/2})$. For $d >2$, similar but slightly less accurate statements hold. {The lower bounds improve over previous ones by Ben-David et al., and the upper bounds are novel.} The lower bounds are obtained by probabilistic constructions, using a theorem of Warren in real algebraic topology. The upper bounds are obtained using a result of Welzl about spanning trees with low stabbing number, and using the moment curve. The upper bound technique is also used to: (i) provide estimates on the number of classes of a given VC dimension, and the number of maximum classes of a given VC dimension -- answering a question of Frankl from '89, and (ii) design an efficient algorithm that provides an $O(N/\log(N))$ multiplicative approximation for the sign rank. We also observe a general connection between sign rank and spectral gaps which is based on Forster's argument. Consider the $N \times N$ adjacency matrix of a $Δ$ regular graph with a second eigenvalue of absolute value $λ$ and $Δ\leq N/2$. We show that the sign rank of the signed version of this matrix is at least $Δ/λ$. We use this connection to prove the existence of a maximum class $C\subseteq\{\pm 1\}^N$ with VC dimension $2$ and sign rank $\tildeΘ(N^{1/2})$. This answers a question of Ben-David et al.~regarding the sign rank of large VC classes. We also describe limitations of this approach, in the spirit of the Alon-Boppana theorem. We further describe connections to communication complexity, geometry, learning theory, and combinatorics.
2013-05-14
Grounded Lipschitz functions on trees are typically flat
Published • View PublicationBIB
A grounded M-Lipschitz function on a rooted d-ary tree is an integer-valued map on the vertices that changes by at most along edges and attains the value zero on the leaves. We study the behavior of such functions, specifically, their typical value at the root v_0 of the tree. We prove that the probability that the value of a uniformly chosen random function at v_0 is more than M+t is doubly-exponentially small in t. We also show a similar bound for continuous (real-valued) grounded Lipschitz functions.
2011-08-12
Monotone expansion
Published • View PublicationBIB
This work, following the outline set in [B2], presents an explicit construction of a family of monotone expanders. The family is essentially defined by the Mobius action of SL_2(R) on the real line. For the proof, we show a product-growth theorem for SL_2(R).
2010-09-22 v2
Rank Bounds for Design Matrices with Applications to Combinatorial Geometry and Locally Correctable Codes
Published • View PublicationBIB
A (q,k,t)-design matrix is an m x n matrix whose pattern of zeros/non-zeros satisfies the following design-like condition: each row has at most q non-zeros, each column has at least k non-zeros and the supports of every two columns intersect in at most t rows. We prove that the rank of any (q,k,t)-design matrix over a field of characteristic zero (or sufficiently large finite characteristic) is at least n - (qtn/2k)^2 . Using this result we derive the following applications: (1) Impossibility results for 2-query LCCs over the complex numbers: A 2-query locally correctable code (LCC) is an error correcting code in which every codeword coordinate can be recovered, probabilistically, by reading at most two other code positions. Such codes have numerous applications and constructions (with exponential encoding length) are known over finite fields of small characteristic. We show that infinite families of such linear 2-query LCCs do not exist over the complex numbers. (2) Generalization of results in combinatorial geometry: We prove a quantitative analog of the Sylvester-Gallai theorem: Let $v_1,...,v_m$ be a set of points in $\C^d$ such that for every $i \in [m]$ there exists at least $δm$ values of $j \in [m]$ such that the line through $v_i,v_j$ contains a third point in the set. We show that the dimension of $\{v_1,...,v_m \}$ is at most $O(1/δ^2)$. Our results generalize to the high dimensional case (replacing lines with planes, etc.) and to the case where the points are colored (as in the Motzkin-Rabin Theorem).
Entropy of Random Walk Range
Published • View PublicationBIB
We study the entropy of the set traced by an $n$-step random walk on $\Z^d$. We show that for $d \geq 3$, the entropy is of order $n$. For $d = 2$, the entropy is of order $n/\log^2 n$. These values are essentially governed by the size of the boundary of the trace.
The Player's Effect
In a function that takes its inputs from various players, the effect of a player measures the variation he can cause in the expectation of that function. In this paper we prove a tight upper bound on the number of players with large effect, a bound that holds even when the players' inputs are only known to be pairwise independent. We also study the effect of a set of players, and show that there always exists a "small" set that, when eliminated, leaves every set with little effect. Finally, we ask whether there always exists a player with positive effect. We answer this question differently in various scenarios, depending on the properties of the function and the distribution of players' inputs. More specifically, we show that if the function is non-monotone or the distribution is only known to be pairwise independent, then it is possible that all players have 0 effect. If the distribution is pairwise independent with minimal support, on the other hand, then there must exist a player with "large" effect.
2007-12-31 v3
The Maximal Probability that k-wise Independent Bits are All 1
Published in Random Struct. Alg., 38, 502-525, 2011 • View PublicationBIB
A k-wise independent distribution on n bits is a joint distribution of the bits such that each k of them are independent. In this paper we consider k-wise independent distributions with identical marginals, each bit has probability p to be 1. We address the following question: how high can the probability that all the bits are 1 be, for such a distribution? For a wide range of the parameters n,k and p we find an explicit lower bound for this probability which matches an upper bound given by Benjamini et al., up to multiplicative factors of lower order. The question we investigate can be seen as a relaxation of a major open problem in error-correcting codes theory, namely, how large can a linear error correcting code with given parameters be? The question is a type of discrete moment problem, and our approach is based on showing that bounds obtained from the theory of the classical moment problem provide good approximations for it. The main tool we use is a bound controlling the change in the expectation of a polynomial after small perturbation of its zeros.