edit distance
51 papers tagged with this keyword
Algebraic Geometry Codes Approach the Half-Singleton Bound with Constant Field Size
We study linear codes for insertion and deletion (insdel) errors through the lens of evaluation codes. We develop a general framework for analyzing random puncturings of evaluation codes, where the edit distance is controlled by only the size of the evaluation domain and the maximum number of zeros of a nonzero function in the underlying function space. Our proof generalizes the results of Con, Guo, Li, and Zhang (ICALP 2025), and simultaneously simplifies their arguments by avoiding an in-depth analysis of longest common subsequences. We demonstrate the applicability of our core theorem by instantiating it with random puncturings of Reed--Muller codes. We then recover the result that random Reed--Solomon codes approach the half-Singleton bound over linear-sized fields while also improving the dependence on the additive gap $\varepsilon$ from $2^{O(1/\varepsilon^2)}$ to $2^{O(1/\varepsilon)}$. Finally, by applying the framework to algebraic geometry codes arising from asymptotically good towers of function fields, we show that there exist randomized families of structured linear codes over constant-sized fields that approach the half-Singleton bound.
$L_2$ Turán Problems for Small Tournaments and Stability
We investigate the $L_2$ Turán problems for various small directed graphs, specifically focusing on self-converse tournaments and stability versions. First, we determine the exact maximum $L_2$ norm squared of the out-degree sequence for digraphs avoiding the transitive tournament $TT_4$ and the strongly connected tournament $R_4$, answering open questions from recent paper. We prove that the complete directed 3-partite Turán graph $T_3(m)$ exactly maximizes the $L_2$ norm squared for $TT_4$-free digraphs. For $R_4$-free digraphs, the maximum is achieved by $T_3(m)$ except when $m \equiv 1 \pmod 3$, where peeling off a terminal sink vertex to form $T_3(m-1) \to v$ strictly increases the objective. We complement these results with exact values and a conjecture for the regular tournament $Reg_5$. Furthermore, we prove a stability version for $\vec{C}_3$-free digraphs: any sequence of digraphs asymptotically achieving the maximum $L_2$ density must have an edit distance of $O(δ^{1/2})m^2$ to the extremal ordered digon-chain $\vec{F}_{m,2}$.
The Cayley Completion of a Graph
A finite connected graph is rarely a Cayley graph. We measure how far it is from being one: given $G$ with $n$ vertices and $m$ edges, how few edges must be added, or added and deleted, before the result is a Cayley graph of an abelian group of order $n$ on the same vertex set? This defines two invariants, the completion number $γ^{+}$ (additions only) and the Cayley edit distance $γ_{\triangle}$ (both), each normalized by $m$. We show that deciding the edit version is NP-complete already for a fixed cyclic host, by a reduction from Hamiltonian Cycle in which the edit cost of a labeling is $n+m-2k$ when it realizes a longest path with $k$ edges; the optimal cost is $m-n+2pp(G)$, bounded in polynomial time by the matching number. We prove that irregularity alone forces $γ^{+}(G)\ge nΔ^{*}/(2m)-1$, where $Δ^{*}$ is the least $d\geΔ$ with $nd$ even, computable in linear time from the degree sequence; we characterize equality exactly. It is attained on the star, where $γ^{+}(K_{1,q})=(q-1)/2$ and the star maximizes $γ^{+}$, while $γ_{\triangle}$ stays bounded by an absolute constant. We determine paths and grids exactly, $γ^{+}(P_n)=γ^{+}(P_n\,\square\,P_n)=1/(n-1)$, and show $γ_{\triangle}(K_{1,q})\to 2$, not the $3/2$ suggested by the additive case. We report an exhaustive certified census of all $995$ connected graphs on at most seven vertices. The degree bound is attained on $89.4\%$ and the two invariants separate strictly on $84.7\%$, though both rates vary sharply with order: attainment $100\%,100\%,84.8\%,89.7\%$ and separation $0\%,61.9\%,73.2\%,87.7\%$ for $n=4,5,6,7$, dominated by the $853$ graphs on seven vertices. The star uniquely maximizes both. Edit count and the bi-Lipschitz distortion of the completed host are independent, moving oppositely on stars and paths.Data and certificates at doi:10.5281/zenodo.21852006.
Cluster-Graph Edit Distance: Metric Proxies, Multiscale Embeddings, and Complexity
The cluster graphs on $n$ vertices, the disjoint unions of complete graphs, have the integer partitions of $n$ as their isomorphism classes, and the quotient edit distance $q^*(λ,μ)=\min_{σ\in S_n}|E(G_λ)\triangleσE(G_μ)|$ makes that set a metric space. Its metric geometry and its computational complexity both issue from one identity: $q^*$ is an affine function of the maximum of $\lVert X\rVert_F^2$ over the contingency tables with margins $λ$ and $μ$. Combinatorially, it yields two explicit $\ell_1$ models: the vertex-mass metric $δ_1$ on sorted degree sequences, with $\frac12δ_1\le q^*<\frac32δ_1$ and both constants optimal, and the block-energy metric $B$ on the vectors $\bigl(\binom{λ_i}2\bigr)_i$, with $q^*\le B\le2q^*-1$ by a per-table refinement measuring how far an alignment is from a block bijection. Hence $c_1(\mathcal K_n)\le2$, and an $O(n\log n)$-time algorithm returns an alignment of cost below $2q^*$ with the certificate $q^*\in[\lceil(B+1)/2\rceil,B]$. The Euclidean distortion of the class is $c_2(\mathcal K_n)=Θ(n^{1/4})$; against it we measure the weighted dyadic sums $F^{(γ)}$ of the Ferrers staircase, of dimension below $4n$ and computable in $O(n)$ time. The unweighted member has distortion exactly $Θ(n^{1/4}\sqrt{\log n})$, while the critical weight $γ=\frac14$ improves this unconditionally to $O(n^{1/4}(\log n)^{1/4})$ through an inverse energy inequality proved from the quantization of staircase jumps; removing the residual $(\log n)^{1/4}$ is reduced to one inverse inequality on the realizable cone. Computationally, the same identity gives a classification: deciding $q^*(λ,μ)\le Q$ is strongly NP-complete, evaluation is strongly NP-hard and admits no FPTAS unless $\mathrm P=\mathrm{NP}$, while the farthest alignment is polynomial-time solvable.
A Proof of Nash-Williams' Conjecture
A central open question in extremal design theory is Nash-Williams' Conjecture from 1970 that every triangle-divisible graph on $n$ vertices (for $n$ large enough) with minimum degree at least $0.75 n$ has a triangle decomposition. In this paper, we prove this conjecture in full.
In 2016, Barber, Kühn, Lo, and Osthus proved that if the fractional relaxation of Nash-Williams' Conjecture holds for minimum degree $cn$ for some constant $c\ge 0.75$, then Nash-Williams' Conjecture holds for any constant $c' > c$. The previously best-known bound on the fractional relaxation was due to Delcourt and Postle from 2021 with $c= \frac{7+\sqrt{21}}{14} \approx 0.82733$. This bound on the fractional relaxation has grown in importance over the years as it has been directly tied to bounds for a number of other problems in extremal design theory.
This paper consists of three parts. In Part I, our first main result is a proof of the Fractional Nash-Williams' Conjecture: if $G$ is a graph on $n$ vertices with minimum degree at least $\frac{3n}{4}$, then $G$ has a fractional triangle decomposition.
In Part II, our second main result is a Fractional Stability Theorem for Nash-Williams' Conjecture: if a graph $G$ on $n$ vertices has minimum degree close to $\frac{3n}{4}$ but no fractional $K_3$-decomposition, then $G$ is close (in edit distance) to the join of two $\frac{n}{4}$-regular graphs each on $\frac{n}{2}$ vertices. We use this to prove that if a triangle-divisible graph $G$ on $n$ vertices has minimum degree close to $\frac{3n}{4}$ but no $K_3$-decomposition, then $G$ is close (in edit distance) to the join of two $\frac{n}{4}$-regular graphs each on $\frac{n}{2}$ vertices.
In Part III, our final main result is a proof of Nash-Williams' Conjecture in full.
The edit distance of word-representable and comparability graphs
In this paper, we establish that the maximum edit distance of an $n$-vertex graph from the hereditary property of word-representable graphs is $n^2/8-o(n^2)$. In addition, we establish that the maximum edit distance of an $n$-vertex graph from the hereditary property of poset comparability graphs is $5n^2/32-o(n^2)$.
In fact, we determine the edit distance function over all edge densities $p\in [0,1]$ for the property of word-representable graphs, for the property of $k$-word-representable graphs for each $k\geq 2$, and for the property comparability graphs. The latter has a peculiar structure that requires an infinite sequence of colored regularity graphs.
Efficient graph similarity assessment method based on vectors of topological indices
Measuring similarity between complex objects is a fundamental task in many scientific fields. When objects are represented as graphs, graph similarity/distance measures offer a powerful framework for quantifying structural resemblance. Those comparative measures play a key role in domains such as network science, chemoinformatics, and social network analysis. While methods like graph edit distance and graph kernels are widely used, they can be computationally intensive or fail to capture fine structural variations, since they require graphs without any structural uncertainty. Another class of methods is based on using topological indices to encode structural information of the graphs, followed by the application of distance or similarity measures for real numbers to obtain corresponding graph-level metrics. In this paper, we introduce a novel class of distance/similarity measures which are based on multiple topological indices. Since they are generally computed in polynomial time, our method is computationally efficient in practice. We demonstrate its effectiveness through comparisons and show that it captures subtle structural information meaningfully. Additionally, we explore its applicability in two domains: analyzing random graph models in network theory and assessing molecular similarity among isomers in chemoinformatics. These preliminary results suggest that our approach holds promise for graph comparison across disciplines.
The Turán density of short tight cycles
The $3$-uniform tight $\ell$-cycle $C_\ell^{3}$ is the $3$-graph on $\{1,\dots,\ell\}$ consisting of all $\ell$ consecutive triples in the cyclic order. Let $\mathcal{C}$ be either the pair $\{C_{4}^{3}, C_{5}^{3}\}$ or the single tight $\ell$-cycle $C_{\ell}^{3}$ for some $\ell\ge 7$ not divisible by $3$.
We show that the Turán density of $\mathcal{C}$, that is, the asymptotically maximal edge density of a large $\mathcal{C}$-free $3$-graph, is equal to $2\sqrt{3} - 3$. We also establish the corresponding Erdős-Simonovits-type stability result, informally stating that all almost maximum $\mathcal{C}$-free graphs are close in the edit distance to a 2-part recursive construction. This extends the earlier analogous results of Kamčev-Letzter-Pokrovskiy ["The Turán density of tight cycles in three-uniform hypergraphs", Int. Math. Res. Not. 6 (2024), 4804-4841] that apply for sufficiently large $\ell$ only.
Additionally, we prove a finer structural result that allows us to determine the maximum number of edges in a $\{C_{4}^{3}, C_{5}^{3}\}$-free $3$-graph with a given number of vertices up to an additive $O(1)$ error term.
Constant Rate Isometric Embeddings of Hamming Metric into Edit Metric
A function $\varphi: \{0,1\}^n \to \{0,1\}^N$ is called an isometric embedding of the $n$-dimensional Hamming metric space to the $N$-dimensional edit metric space if, for all $x, y \in \{0,1\}^n$, the Hamming distance between $x$ and $y$ is equal to the edit distance between $\varphi(x)$ and $\varphi(y)$. The rate of such an embedding is defined as the ratio $n/N$. It is well known in the literature how to construct isometric embeddings with a rate of $Ω(\frac{1}{\log n})$. However, achieving even near-isometric embeddings with a positive constant rate has remained elusive until now.
In this paper, we present an isometric embedding with a rate of 1/8 by discovering connections to synchronization strings, which were studied in the context of insertion-deletion codes (Haeupler-Shahrasbi [JACM'21]). At a technical level, we introduce a framework for obtaining high-rate isometric embeddings using a novel object called misaligners. As an immediate consequence of our constant rate isometric embedding, we improve known conditional lower bounds for various optimization problems in the edit metric, but now with optimal dependency on the dimension.
We complement our results by showing that no isometric embedding $\varphi:\{0, 1\}^n \to \{0, 1\}^N$ can have rate greater than 15/32 for all positive integers $n$. En route to proving this upper bound, we uncover fundamental structural properties necessary for every Hamming-to-edit isometric embedding. We also prove similar upper and lower bounds for embeddings over larger alphabets.
Finally, we consider embeddings $\varphi:Σ_{\text{in}}^n\to Σ_{\text{out}}^N$ between different input and output alphabets, where the rate is given by $\frac{n\log|Σ_{\text{in}}|}{N\log|Σ_{\text{out}}|}$. In this setting, we show that the rate can be made arbitrarily close to 1.
Answering Related Questions
We introduce the meta-problem Sidestep$(Π, \mathsf{dist}, d)$ for a problem $Π$, a metric $\mathsf{dist}$ over its inputs, and a map $d: \mathbb N \to \mathbb R_+ \cup \{\infty\}$. A solution to Sidestep$(Π, \mathsf{dist}, d)$ on an input $I$ of $Π$ is a pair $(J, Π(J))$ such that $\mathsf{dist}(I,J) \leqslant d(|I|)$ and $Π(J)$ is a correct answer to $Π$ on input $J$. This formalizes the notion of answering a related question (or sidestepping the question), for which we give some practical and theoretical motivations, and compare it to the neighboring concepts of smoothed analysis, planted problems, and edition problems. Informally, we call hardness radius the ``largest'' $d$ such that Sidestep$(Π, \mathsf{dist}, d)$ is NP-hard. This framework calls for establishing the hardness radius of problems $Π$ of interest for the relevant distances $\mathsf{dist}$.
We exemplify it with graph problems and two distances $\mathsf{dist}_Δ$ and $\mathsf{dist}_e$ (the edge edit distance) such that $\mathsf{dist}_Δ(G,H)$ (resp. $\mathsf{dist}_e(G,H)$) is the maximum degree (resp. number of edges) of the symmetric difference of $G$ and $H$ if these graphs are on the same vertex set, and $+\infty$ otherwise. We show that the decision problems Independent Set, Clique, Vertex Cover, Coloring, Clique Cover have hardness radius $n^{\frac{1}{2}-o(1)}$ for $\mathsf{dist}_Δ$, and $n^{\frac{4}{3}-o(1)}$ for $\mathsf{dist}_e$, that Hamiltonian Cycle has hardness radius 0 for $\mathsf{dist}_Δ$, and somewhere between $n^{\frac{1}{2}-o(1)}$ and $n/3$ for $\mathsf{dist}_e$, and that Dominating Set has hardness radius $n^{1-o(1)}$ for $\mathsf{dist}_e$. We leave several open questions.
Edit distance in substitution systems
Let $σ$ be a primitive substitution on an alphabet $\mathcal{A}$, and let $\mathcal{W}_n$ be the set of words of length $n$ determined by $σ$ (i.e., $w \in \mathcal{W}_n$ if $w$ is a subword of $σ^k(a)$ for some $a \in \mathcal{A}$ and $k \geq 1$). It is known that the corresponding substitution dynamical system is loosely Kronecker (also known as zero-entropy loosely Bernoulli), so the diameter of $\mathcal{W}_n$ in the edit distance is $o(n)$. We improve this upper bound to $O(n/\sqrt{\log n})$. The main challenge is handling the case where $σ$ is non-uniform; a better bound is available for the uniform case. Finally, we show that for the Thue--Morse substitution, the diameter of $\mathcal{W}_n$ is at least $\sqrt {n/6} - 1$.
Combinatorial alphabet-dependent bounds for insdel codes
Error-correcting codes resilient to synchronization errors such as insertions and deletions are known as insdel codes. Due to their important applications in DNA storage and computational biology, insdel codes have recently become a focal point of research in coding theory.
In this paper, we present several new combinatorial upper and lower bounds on the maximum size of $q$-ary insdel codes. Our main upper bound is a sphere-packing bound obtained by solving a linear programming (LP) problem. It improves upon previous results for cases when the distance $d$ or the alphabet size $q$ is large. Our first lower bound is derived from a connection between insdel codes and matchings in special hypergraphs. This lower bound, together with our upper bound, shows that for fixed block length $n$ and edit distance $d$, when $q$ is sufficiently large, the maximum size of insdel codes is $ \frac{q^{n-\frac{d}{2}+1}}{{n\choose \frac{d}{2}-1}}(1 \pm o(1))$. The second lower bound refines Alon et al.'s recent logarithmic improvement on Levenshtein's GV-type bound and extends its applicability to large $q$ and $d$.
Upper bounds on the average edit distance between two random strings
We study the average edit distance between two random strings. More precisely, we adapt a technique introduced by Lueker in the context of the average longest common subsequence of two random strings to improve the known upper bound on the average edit distance. We improve all the known upper bounds for small alphabets. We also provide a new implementation of Lueker technique to improve the lower bound on the average length of the longest common subsequence of two random strings for all small alphabets of size other than $2$ and $4$.
Stability of transversal Hamilton cycles and paths
Given graphs $G_1,\ldots,G_s$ all on a common vertex set and a graph $H$ with $e(H) = s$, a copy of $H$ is \emph{transversal} or \emph{rainbow} if it contains one edge from each $G_i$. We establish a stability result for transversal Hamilton cycles: the minimum degree required to guarantee a transversal Hamilton cycle can be lowered as long as the graph collection $G_1,\ldots,G_n$ is far in edit distance from several extremal cases. We obtain an analogous result for Hamilton paths. The proof is a combination of our newly developed regularity-blow-up method for transversals, along with the absorption method.
Quantifying analogy of concepts via ologs and wiring diagrams
We build on the theory of ontology logs (ologs) created by Spivak and Kent, and define a notion of wiring diagrams. In this article, a wiring diagram is a finite directed labelled graph. The labels correspond to types in an olog; they can also be interpreted as readings of sensors in an autonomous system. As such, wiring diagrams can be used as a framework for an autonomous system to form abstract concepts. We show that the graphs underlying skeleton wiring diagrams form a category. This allows skeleton wiring diagrams to be compared and manipulated using techniques from both graph theory and category theory. We also extend the usual definition of graph edit distance to the case of wiring diagrams by using operations only available to wiring diagrams, leading to a metric on the set of all skeleton wiring diagrams. In the end, we give an extended example on calculating the distance between two concepts represented by wiring diagrams, and explain how to apply our framework to any application domain.
Removing induced powers of cycles from a graph via fewest edits
What is the minimum proportion of edges which must be added to or removed from a graph of density $p$ to eliminate all induced cycles of length $h$? The maximum of this quantity over all graphs of density $p$ is measured by the edit distance function, $\text{ed}_{\text{Forb}(C_h)}(p)$, a function which provides a natural metric between graphs and hereditary properties.
Martin determined $\text{ed}_{\text{Forb}(C_h)}(p)$ for all $p \in [0,1]$ when $h \in \{3, \ldots, 9\}$ and determined $\text{ed}_{\text{Forb}(C_{10})}(p)$ for $p \in [1/7, 1]$. Peck determined $\text{ed}_{\text{Forb}(C_h)}(p)$ for all $p \in [0,1]$ for odd cycles, and for $p \in [ 1/\lceil h/3 \rceil, 1]$ for even cycles. In this paper, we fully determine the edit distance function for $C_{10}$ and $C_{12}$. Furthermore, we improve on the result of Peck for even cycles, by determining $\text{ed}_{\text{Forb}(C_h)}(p)$ for all $p \in [p_0, 1/\lceil h/3 \rceil ]$, where $p_0 \leq c/h^2$ for a constant $c$. More generally, if $C_h^t$ is the $t$-th power of the cycle $C_h$, we determine $\text{ed}_{\text{Forb}(C_h^t)}(p)$ for all $p \geq p_0$ in the case when $(t+1) \mid h$, thus improving on earlier work of Berikkyzy, Martin and Peck.
Testing versus estimation of graph properties, revisited
A distance estimator for a graph property $\mathcal{P}$ is an algorithm that given $G$ and $α, \varepsilon >0$ distinguishes between the case that $G$ is $(α-\varepsilon)$-close to $\mathcal{P}$ and the case that $G$ is $α$-far from $\mathcal{P}$ (in edit distance). We say that $\mathcal{P}$ is estimable if it has a distance estimator whose query complexity depends only on $\varepsilon$.
Every estimable property is also testable, since testing corresponds to estimating with $α=\varepsilon$. A central result in the area of property testing, the Fischer--Newman theorem, gives an inverse statement: every testable property is in fact estimable. The proof of Fischer and Newman was highly ineffective, since it incurred a tower-type loss when transforming a testing algorithm for $\mathcal{P}$ into a distance estimator. This raised the natural problem, studied recently by Fiat--Ron and by Hoppen--Kohayakawa--Lang--Lefmann--Stagni, whether one can find a transformation with a polynomial loss. We obtain the following results.
1. If $\mathcal{P}$ is hereditary, then one can turn a tester for $\mathcal{P}$ into a distance estimator with an exponential loss. This is an exponential improvement over the result of Hoppen et. al., who obtained a transformation with a double exponential loss.
2. For every $\mathcal{P}$, one can turn a testing algorithm for $\mathcal{P}$ into a distance estimator with a double exponential loss. This improves over the transformation of Fischer--Newman that incurred a tower-type loss. Our main conceptual contribution in this work is that we manage to turn the approach of Fischer--Newman, which was inherently ineffective, into an efficient one. On the technical level, our main contribution is in establishing certain properties of Frieze--Kannan Weak Regular partitions that are of independent interest.
A Persistence-Driven Edit Distance for Trees with Abstract Weights
In this work we define a novel edit distance for trees considered with some abstract weights on the edges. The metric is driven by the idea of considering trees as topological summaries in the context of persistence and topological data analysis. Several examples related to persistent sets are presented. The metric can be computed with a dynamical binary linear programming approach. This framework is applied and further studied in other works focused on merge trees, where the problems of stability and merge trees estimation are also assessed.
Stability from graph symmetrization arguments in generalized Turán problems
Given graphs $H$ and $F$, $\mathrm{ex}(n,H,F)$ denotes the largest number of copies of $H$ in $F$-free $n$-vertex graphs. Let $χ(H)<χ(F)=r+1$. We say that $H$ is $F$-Turán-stable if the following holds. For any $\varepsilon>0$ there exists $δ>0$ such that if an $n$-vertex $F$-free graph $G$ contains at least $\mathrm{ex}(n,H,F)-δn^{|V(H)|}$ copies of $H$, then the edit distance of $G$ and the $r$-partite Turán graph is at most $\varepsilon n^2$. We say that $H$ is weakly $F$-Turán-stable if the same holds with the Turán graph replaced by any complete $r$-partite graph $T$. It is known that such stability implies exact results in several cases. We show that complete multipartite graphs with chromatic number at most $r$ are weakly $K_{r+1}$-Turán-stable. Answering a question of Morrison, Nir, Norin, Rzażewski and Wesolek positively, we show that for every graph $H$, if $r$ is large enough, then $H$ is $K_{r+1}$-Turán-stable. Finally, we prove that if $H$ is bipartite, then it is weakly $C_{2k+1}$-Turán-stable for $k$ large enough.
Hypercubes and Isometric Words based on Swap and Mismatch Distance
The hypercube of dimension n is the graph whose vertices are the 2^n binary words of length n, and there is an edge between two of them if they have Hamming distance 1. We consider an edit distance based on swaps and mismatches, to which we refer as tilde-distance, and define the tilde-hypercube with edges linking words at tilde-distance 1. Then, we introduce and study some isometric subgraphs of the tilde-hypercube obtained by using special words called tilde-isometric words. The subgraphs keep only the vertices that avoid a given tilde-isometric word as a factor. In the case of word 11, the subgraph is called tilde-Fibonacci cube, as a generalization of the classical Fibonacci cube. The tilde-hypercube and the tilde-Fibonacci cube can be recursively defined; the same holds for the number of their edges. This allows an asymptotic estimation of the number of edges in the tilde-Fibonacci cube, in comparison to the total number in the tilde-hypercube.