cs.DS ↗ arXiv
158 papers in this category
Palette Sparsification for General Uniform Hypergraphs
We prove a palette sparsification theorem for general $r$-uniform hypergraphs. For all sufficiently large $n$, every $r\ge 3$, and every $α\ge 7.1$, we show that an $n$-vertex $r$-uniform hypergraph of maximum degree $Δ$ is w.h.p. colorable from independently sampled lists of size $O(\sqrt{\log n})$ drawn from an ambient palette of size $\lceil αΔ^{1/(r-1)}\rceil$. The $\sqrt{\log n}$ dependence is asymptotically tight.
A Canonical m-Atomic Decomposition of Bipartite Graphs via a Grid Model
We study finite, connected, simple bipartite graphs in a grid model, in which a graph is drawn as a rectangular array and its structure is read off from empty subrectangles, called holes. In this model we attach to every brick a numerical invariant, its characteristic m, the difference between the number of rows and the largest proper independent set. A brick is excessive if m > 0.
Our main results concern this invariant. We determine the characteristic of a disconnected excessive brick from those of its components, showing that m = min_i min{m_i, imb(W_i)} while the imbalance is additive; and we prove that an m-excessive brick is m-extendable, that is, every matching of size m extends to a maximum matching. Since Plummer's notion of n-extendability is defined only for graphs carrying a perfect matching, and our proof nowhere uses balance, the characteristic extends that notion canonically to unbalanced bipartite graphs. Using the characteristic we partition bipartite graphs into eleven structural classes.
The underlying decomposition into atomic blocks is the classical decomposition into elementary components, and the description of the maximum proper independent sets by ideals of the block poset is likewise classical; the paper states precisely which results are classical and are not claimed here. What the grid model adds is a single geometric framework in which holes, characteristics and the block triangular form are read off from one picture.
Fast Algorithms for Stoquastic Spin Systems
We establish a general framework for developing fast sampling and counting algorithms for stoquastic spin systems at high temperature. Our framework is based on a rapidly mixing Markov chain for polymer models and a subcritical percolation process for sampling individual polymers. We apply our framework to obtain fast algorithms for approximating the partition function and sampling from the thermal distribution of (1) general stoquastic spin systems, (2) ferromagnetic Heisenberg models, and (3) antiferromagnetic Heisenberg models on bipartite graphs. For the Heisenberg models, we obtain an improved bound on the inverse temperature by using their respective cycle and loop representations.
Online Permutation Embedding: Optimal Stopping and Scaling Laws
We study optimal online algorithms for embedding a permutation $π$ of $[k]$ into an iid stream of uniform $[0,1]$ random variables. This problem is a broad generalization of the classical online monotone subsequence selection problem, recovered in the special case $π=\mathrm{Id}_k$. Our first contribution is an efficiently solvable dynamic program for the optimal embedding time of any $k$-permutation $π$. This dynamic program also yields an explicit optimal online embedding algorithm. We then investigate the asymptotic scaling of the optimal embedding time for uniformly random target permutations, as well as the extremal problem of identifying the permutations with largest expected online embedding time. Our second main result shows that, to first order, random permutations are strictly faster to embed than monotone permutations, which in turn are strictly faster to embed than the extremal permutations. This separation stands in sharp contrast to prevailing conjectures and heuristics in the offline theory of permutation embeddings.
Cluster-Graph Edit Distance: Metric Proxies, Multiscale Embeddings, and Complexity
The cluster graphs on $n$ vertices, the disjoint unions of complete graphs, have the integer partitions of $n$ as their isomorphism classes, and the quotient edit distance $q^*(λ,μ)=\min_{σ\in S_n}|E(G_λ)\triangleσE(G_μ)|$ makes that set a metric space. Its metric geometry and its computational complexity both issue from one identity: $q^*$ is an affine function of the maximum of $\lVert X\rVert_F^2$ over the contingency tables with margins $λ$ and $μ$. Combinatorially, it yields two explicit $\ell_1$ models: the vertex-mass metric $δ_1$ on sorted degree sequences, with $\frac12δ_1\le q^*<\frac32δ_1$ and both constants optimal, and the block-energy metric $B$ on the vectors $\bigl(\binom{λ_i}2\bigr)_i$, with $q^*\le B\le2q^*-1$ by a per-table refinement measuring how far an alignment is from a block bijection. Hence $c_1(\mathcal K_n)\le2$, and an $O(n\log n)$-time algorithm returns an alignment of cost below $2q^*$ with the certificate $q^*\in[\lceil(B+1)/2\rceil,B]$. The Euclidean distortion of the class is $c_2(\mathcal K_n)=Θ(n^{1/4})$; against it we measure the weighted dyadic sums $F^{(γ)}$ of the Ferrers staircase, of dimension below $4n$ and computable in $O(n)$ time. The unweighted member has distortion exactly $Θ(n^{1/4}\sqrt{\log n})$, while the critical weight $γ=\frac14$ improves this unconditionally to $O(n^{1/4}(\log n)^{1/4})$ through an inverse energy inequality proved from the quantization of staircase jumps; removing the residual $(\log n)^{1/4}$ is reduced to one inverse inequality on the realizable cone. Computationally, the same identity gives a classification: deciding $q^*(λ,μ)\le Q$ is strongly NP-complete, evaluation is strongly NP-hard and admits no FPTAS unless $\mathrm P=\mathrm{NP}$, while the farthest alignment is polynomial-time solvable.
Derandomizing Karger's Contraction Algorithm for Matroids
Karger's randomized contraction algorithm finds a minimum-weight cocircuit of a matroid whenever the cogirth-density ratio is bounded. We prove that the same hypothesis yields a deterministic algorithm with the same exponent. If every contraction minor of rank at least $r_0$ of a matroid $M$ has cogirth-density ratio at most $c$, then a minimum-weight cocircuit of $M$ is computable deterministically in $m^{O(r_0)} n^{O(c)}$ time when the contraction minors of bounded rank have at most $m$ parallel classes, by an algorithm that knows neither $r_0$ nor $c$. As a consequence, we give a deterministic algorithm computing the cogirth of rank-$p$ perturbed graphic matroids in $2^{O(p^2)} n^{O(1)}$ time, fixed-parameter tractable in $p$, settling the cogirth side of a question of Geelen and Kapadia (2018). The extensions of the contraction method carry over deterministically: enumerating all near-minimum 1-cocycles, computing a minimum-weight $k$-cocycle, and computing the Pareto frontier under several positive criteria.
Isomorphism of tournaments with bounded VC dimension
The tournament isomorphism problem is one of the two fundamental bottlenecks to designing better algorithms for the graph isomorphism problem. Though the problem has been investigated for more than five decades, compared to graphs, there are only very few results on the isomorphism problem of tournaments. For most classes of tournaments neither hardness nor polynomial-time solvability is known.
Tournaments of bounded VC dimension are such a class for which no results are available, even though the VC dimension is arguably one of the most robust and central notions of combinatorial tameness. Resolving an open problem of Neuen and Grohe, we show that the isomorphism problem for tournaments of VC dimension $d$ can be decided in time $n^{O(d\log d)}$. Consequently, automorphism groups of tournaments of bounded VC dimension can be computed in polynomial time. To this end, we develop a new method to isomorphism-invariantly decompose tournaments. To facilitate recursion, we introduce the notion of a patched tournament and analyze bounded VC dimension in patched tournaments. We design a recursive algorithm that balances the size of the decomposed pieces against their number and makes use of the structure of near twins.
In an orthogonal direction, it is known that a hereditary class of tournaments has unbounded VC dimension if and only if it contains all 2-colorable tournaments. As a second result, we show that also this class does not form an obstruction towards polynomial-time isomorphism testing and indeed show that isomorphism of tournaments of bounded chromatic number is polynomial-time decidable.
Online balancing of vectors with small coordinates
Let $v_1,\ldots,v_T\in B_2^m$ be fixed in advance and revealed sequentially, and assume that $\|v_t\|_\infty\leqslant d^{-1/2}$ for some $d\geqslant 1$ and every $1\leqslant t\leqslant T$. There are absolute constants $L,C,c>0$ and a randomized online signing such that $$\mathbb{P}\left\{\max_{k\leqslant T}\left\|\sum_{t=1}^k\varepsilon_t v_t\right\|_\infty>6L\right\} \leqslant CT\exp\left(-\frac{cd}{\ln^2(ed)}\right).$$ Consequently, constant prefix discrepancy holds with probability at least $1-\varepsilon$ once $d$ is at least $C\ln\frac{3T}{\varepsilon}\left[\ln\left(e+\ln\frac{3T}{\varepsilon}\right)\right]^2$. In particular, every fixed sequence of vectors $a_t\in[-1,1]^m$ with at most $d$ nonzero coordinates admits an online signing with prefix discrepancy $O(\sqrt d)$ and failure probability at most $CT\exp[-cd/\ln^2(ed)]$. We also prove a nonuniform version in which the failure probability depends on the individual parameters $d_t=\|v_t\|_\infty^{-2}$, and a lower bound showing that a universal constant prefix discrepancy is impossible when $d=o(\ln T)$. We identify the corresponding $\ln^2 d$ barrier for the compact-potential method and extend the argument to general symmetric target bodies admitting a quadratic smoothness estimate.
Online Interval Selection on a Simple Chain
Published
• View Publication
• BIB
A set of intervals $I = \{ I_1, I_2, \dots, I_n \}$ forms a simple chain if, for every $2\leq i \leq n-1$, interval $I_i$ overlaps only with $I_{i-1}$ and $I_{i+1}$. We show that a deterministic memoryless one-directional revoking algorithm achieves a competitive ratio of $2(1 - 1/\sqrt{e}) \approx 0.786$ on the simple chain in the random order model, hence performs worse than the basic greedy algorithm without revoking that has a competitive ratio of $(1 - 1/e^2) \approx 0.864$, but better than any deterministic revoking algorithm in the adversarial model that has a competitive ratio of at most $0.75$. The proof of the latter also leads to a lower bound of $n/4$ for the advice complexity.
A 5/4 bound for graphic $s$-$t$ path TSP on subcubic graphs
We study the graphic $s$-$t$ path TSP on subcubic graphs (maximum degree 3): given two vertices $s,t$, find a shortest walk from $s$ to $t$ that visits every vertex. Our main result is that the optimal $5/4$ coefficient is attained for every terminal pair -- including the difficult case where deleting both $s$ and $t$ disconnects the graph. Concretely, every pair of distinct vertices $s,t$ in a simple 2-connected subcubic graph $G$ admits a spanning $s$-$t$ walk of length at most $\lfloor(5n+n_2(G))/4\rfloor-1$, where $n=|V(G)|$ and $n_2(G)$ is the number of degree-2 vertices; the asymptotic coefficient $5/4$ cannot be improved, and a simple $O(n^2)$ algorithm finds a walk of length at most $\lfloor(5n+n_2(G))/4\rfloor$.
An edge-rooted even-cover theorem of Wigal, Yoo, and Yu, combined with a short conversion lemma proved here, gives a bound of this form only when $s$ and $t$ are the two endpoints of a given edge; we remove that adjacency restriction. For cubic graphs ($n_2(G)=0$) the bound reads $\lfloor 5n/4\rfloor-1$, to our knowledge the first $5/4$ bound for cubic path TSP proved directly rather than through the general path-to-tour reduction.
Online Discrepancy Minimization for Sub-Gaussian Inputs via Regularization and Restriction
We study online discrepancy minimization: vectors $v_1,\ldots,v_T\in\mathbb{R}^n$ arrive sequentially, and each must immediately be assigned a sign $x_t\in\{\pm1\}$, with the aim of minimizing $\|\sum_{t=1}^T x_t v_t\|_\infty$. We give a polynomial-time potential-based algorithm combining a regularization of the $\ell_\infty$-norm with restriction to an adaptively chosen coordinate set. For i.i.d. inputs with independent, symmetric, centered, unit-variance sub-Gaussian coordinates of sub-Gaussian norm at most $σ$, the algorithm achieves terminal discrepancy $O(σ^8\sqrt{n})$ with probability at least $1-\exp(-Ω(σ^3\sqrt{n}))$. If the coordinates are independently masked by Bernoulli variables with mean $k/n$, where $k\gtrsim(\log n)^2$, the bound improves to $O(σ^8\sqrt{k})$, with failure probability $\exp(-Ω(σ^3\sqrt{k}))$. Both guarantees hold for every prescribed finite horizon $T$, with no dependence on $T$. The dense result substantially generalizes a theorem of Bansal and Spencer (2020) for Rademacher inputs and gives an efficient $O(\sqrt{n})$ bound for Gaussian inputs, as conjectured by Gamarnik et al. (2022). When $T$ is polynomially larger than $n$, this is conditionally close to optimal: under worst-case hardness assumptions for standard approximate lattice problems, Vafa and Vaikuntanathan (2025) showed that no polynomial-time algorithm, even offline, can improve the $\sqrt{n}$ scale by a fixed polynomial factor in $T/n$.
A Necessary and Sufficient Hall Condition for Hypergraphs
We prove a necessary and sufficient Hall condition for a family $A=(A_e)_{e\in E(G)}$ of hypergraphs, possibly with loops, indexed by the edges of a forest $G$. We also show that acyclicity of the index graph is sharp for this Hall characterization. As an application, we prove that every $5$-tough chordal graph is Hamilton-connected, improving earlier sufficient toughness bounds for Hamiltonicity of $18$ in 1998 and $10$ in 2017.
On the self-intersection time of non-backtracking random walks
We study the self-intersection time of the non-backtracking random walk on connected undirected graphs. For every fixed $Δ\geq 3$ we show that the expected self-intersection time is $O(\sqrt{n} \log n)$ on $n$-vertex graphs with minimum degree at least $3$ and maximum degree at most $Δ$. For regular graphs with a uniform spectral gap, we improve this to $O(\sqrt{n})$. We also show an $Ω(\sqrt{n})$ lower bound on a class of regular expanders. Our upper bound on the expected self-intersection time implies an improved mixing time bound on Glauber dynamics for the Ising model on $Δ$-regular graphs at the tree uniqueness threshold.
Proper $\{a,b\}$-edge-weightings of trees
Let $a$ and $b$ be distinct real weights. An $\{a,b\}$-edge-weighting of a tree assigns one of these weights to each edge and is proper if adjacent vertices have different sums of incident edge weights. For every such pair, we give an explicit structural characterization of the trees that do not admit a proper $\{a,b\}$-edge-weighting. If $ab(a+b)\neq0$, then $K_2$ is the only tree without such a weighting. If $a+b=0$, then a tree has no proper $\{a,b\}$-edge-weighting exactly when every vertex has degree $1$ or $3$ and the subgraph induced by the degree-$3$ vertices has a perfect matching. For the remaining case $ab=0$, form the spanning forest consisting of the edges whose deletion leaves two odd-order components. A tree $T$ has no proper $\{a,b\}$-edge-weighting exactly when both bipartition classes have odd order and every component of this forest satisfies two conditions. First, every component satisfies the preceding degree-and-matching condition. Second, within each component, the degree of a vertex $v$ in the forest plus twice the number of incident edges $e$ outside the forest for which the component of $T-e$ not containing $v$ has an odd number of vertices from each bipartition class is independent of $v$. For every fixed pair of distinct real weights, the proofs yield a linear-time algorithm that decides whether a proper $\{a,b\}$-edge-weighting exists and constructs one when it does.
A Degree Threshold for Independent Domination in Generalized Prisms
We study per-colour independent (k)-rainbow domination and its connection with independent domination in generalized prisms. Building on the known prism identity and the trivial regime above the maximum degree, we focus on the boundary case where the number of colours equals the maximum degree.
For every fixed (k\ge 3), we prove that the decision problem remains NP-complete even on a highly restricted class of graphs: (C_4)-free, bipartite, ((k,2))-biregular subdivision graphs arising from simple (k)-regular graphs. The reduction gives an exact correspondence between optimal rainbow-independent dominating functions on the subdivision graph and proper (k)-edge-colourings of the original graph.
We also introduce an excess parameter measuring how far the domination number lies above its natural lower bound. For cubic graphs, this excess coincides with the classical edge-colouring degree and therefore with standard resistance parameters for subcubic graphs.
These results reveal a sharp one-unit threshold: above the maximum degree the problem becomes trivial for every graph, while at the boundary NP-hard instances already occur within a very narrow structural family.
Coatom Enumeration in Hypergraph Horn Functions: Rank-Three Representations of Horn Model Posets
For a finite hypergraph H, the complements of the models of its associated definite Horn CNF are exactly the stopping sets of H; hence the complements of its coatoms are the inclusion-minimal nonempty stopping sets. We study their output-sensitive enumeration from the hypergraph incidence representation. Our main result is a representation of arbitrary Horn model posets whose incidence size is linear in the incidence length of the normalized Horn input. Given a Horn CNF $Γ$, we construct a hypergraph $C(Γ)$ of rank at most three whose proper-model poset is inclusion-order isomorphic to the model poset of $Γ$; equivalently, each source model has a unique extension to a proper target model. Thus maximal models of $Γ$ correspond bijectively to target coatoms. Combining this representation with the maximal-model construction of Kavvadias, Sideri, and Stavropoulos shows that coatom enumeration is not in OutputP unless P=NP, even when all hyperedges have size two or three. Incidence splitting reduces maximum element frequency to three while preserving the stopping-set poset, and a local replacement of two-element hyperedges yields the same lower bound for three-uniform hypergraphs of maximum element frequency at most three. These thresholds are conditionally sharp for arbitrary-order enumeration: rank at most two and maximum element frequency at most two both admit output-linear total-time generation; in the frequency-two case, a polynomial-delay, polynomial-space algorithm is also available. In contrast, coatom extension is NP-complete already for three-uniform hypergraphs of exact element frequency two.
Quadratic Degree Sequence Optimization and the Critical Roots of a Graph
The degree sequence optimization problem is to find a subgraph of a given graph which maximizes the sum over all vertices of a given function evaluated at the subgraph degree of that vertex. Here we study this problem and its complexity for quadratic functions. In particular, we introduce the critical roots of a graph, and show they define intervals over which the optimal value of the problem, as the quadratic root varies, is convex piecewise affine.
Simultaneous Graph Parameters and How to Bound Them
Beisegel et al. [SWAT 2024] introduced the concept of simultaneous $\mathcal{C}$-numbers which associate a graph class $\mathcal{C}$ with a graph parameter. Given a graph $G$, the simultaneous $\mathcal{C}$-number is the smallest number $d$ for which there is a graph $H \in \mathcal{C}$ and a function $L : V(G) \to \mathcal{P}(\{1,\dots,d\})$ such that two vertices $u$ and $v$ are adjacent in $G$ if and only if they are adjacent in $H$ and their sets $L(u)$ and $L(v)$ are not disjoint. We study the relation of these simultaneous $\mathcal{C}$-numbers to other graph parameters. In particular, we investigate which parameters fulfill the following property: Parameter $p$ is bounded on class $\mathcal{C}$ if and only if $p$ is bounded on the class of graphs of simultaneous $\mathcal{C}$-number $d$ for any fixed $d$. We show that many well-known graph parameters have this property. Examples are cliquewidth, twin-width, mim-width, tree independence number, thinness as well as boxicity. We furthermore present some parameters, including modular-width and tree-length, that do no have this property. We also study when a parameter forms an upper bound on a simultaneous $\mathcal{C}$-number. We characterize those graph classes $\mathcal{C}$ for which the parameters treewidth, pathwidth, bandwidth, and treedepth upper bound the simultaneous $\mathcal{C}$-number. Furthermore, we present sufficient conditions on a class $\mathcal{C}$, such that $\mathcal{P}$-modular cardinality upper bounds the simultaneous $\mathcal{C}$-number, where $\mathcal{P}$ is replaced by the complete graphs, the edgeless graphs, cographs, or the class $\mathcal{C}$ itself. On the contrary, we show that modular width never forms an upper bound on a non-trivial simultaneous $\mathcal{C}$-number. Finally, we present some general algorithmic results on the clique problem and computation of simultaneous $\mathcal{C}$-numbers.
Bicriteria Approximation Algorithms for Demand Matching
The demand matching problem generalizes both the knapsack problem and the $b$-matching problem. In this problem, each edge of a graph has a demand and a weight, each vertex has a capacity, and the goal is to find a maximum weight subset of edges whose total incident demand at every vertex does not exceed its capacity. We study $(α, β)$-bicriteria approximation algorithms, which return a solution of weight at least $1/α$ times the optimum while allowing an additive capacity violation of at most $β$ times the maximum edge demand.
We give an iterative relaxation algorithm for the demand matching problem that exploits a structural characterization of strictly fractional extreme points of the natural LP relaxation, which reduces the residual rounding problem to odd-cycle instances. Combined with a better-of-two rounding strategy, this yields $(7/6, 1)$- and $(1, 1)$-bicriteria approximation algorithms for general and bipartite graphs, respectively. We further generalize this approach to obtain a parametric family of algorithms, including a $(1, 4/3)$-bicriteria approximation. Separately, for the more general $k$-hypergraph demand matching problem, we give a greedy, combinatorial $(k, 1)$-bicriteria approximation algorithm.
We complement these algorithmic results with matching lower bounds relative to the natural LP relaxation for $β= 0$ and all $β\geq 1$, completely characterizing the trade-off between weight approximation and additive capacity violation in this range.
Quality Control Algorithms for Pattern Counting
In recent work, Marcussen, Rubinfeld, and Sudan introduced the notion of quality control problems, which aim to capture the task of determining if a given input is truly random. Formally, their goal is to accept typical inputs from the specified distribution while rejecting every input whose value of a specified statistic is far from the distributional baseline. This captures the empirical practice of using specified statistics as a proxy for the quality of randomness. Empirical algorithms, however, have not exploited the asymmetry in the definition of quality control problems, which require soundness guarantees in the worst-case while only seeking average-case completeness. Their work abstracted a problem definition emphasizing this asymmetry and used it to give efficient quality control algorithms for assessing the randomness of graphs.
In this work, we introduce and study quality control problems over sequences, where the goal is to distinguish a sequence of i.i.d. characters from sequences where some specified pattern appears too often (or too infrequently) as a subsequence. We consider this problem in both the finite-alphabet setting and for real-valued sequences. We refer to the former setting as the pattern counting problem. In the latter case, the natural notion of a pattern is to consider the relative ordering of the characters in the subsequence, and we refer to this as the permutation pattern counting problem. Algorithms to approximately count (permutation) patterns of length $k$ in a worst-case sequence of length $n$ can provably require exponential in $k$ queries into the sequence. In contrast, we show that by taking advantage of the asymmetry in the definition of quality control, we give algorithms that run in poly$(k)$ time to solve these problems. We also prove that any quality control algorithm (over some natural distributions) requires superlinear queries in $k$.