Papers by Gennian Ge
86 paper(s) by this author
· All BibTeX
On the Fixed-Length-Burst Levenshtein Ball with Unit Radius
Consider a length-$n$ sequence $\bm{x}$ over a $q$-ary alphabet. The \emph{fixed-length Levenshtein ball} $\mathcal{L}_t(\bm{x})$ of radius $t$ encompasses all length-$n$ $q$-ary sequences that can be derived from $\bm{x}$ by performing $t$ deletions followed by $t$ insertions. Analyzing the size and structure of these balls presents significant challenges in combinatorial coding theory. Recent studies have successfully characterized fixed-length Levenshtein balls in the context of a single deletion and a single insertion. These works have derived explicit formulas for various key metrics, including the exact size of the balls, extremal bounds (minimum and maximum sizes), as well as expected sizes and their concentration properties. However, the general case involving an arbitrary number of $t$ deletions and $t$ insertions $(t>1)$ remains largely uninvestigated. This work systematically examines fixed-length Levenshtein balls with multiple deletions and insertions, focusing specifically on \emph{fixed-length burst Levenshtein balls}, where deletions occur consecutively, as do insertions. We provide comprehensive solutions for explicit cardinality formulas, extremal bounds (minimum and maximum sizes), expected size, and concentration properties surrounding the expected value.
Optimal redundancy of function-correcting codes
Function-correcting codes, introduced by Lenz, Bitar, Wachter-Zeh, and Yaakobi, protect specific function values of a message rather than the entire message. A central challenge is determining the optimal redundancy -- the minimum additional information required to recover function values amid errors. This redundancy depends on both the number of correctable errors $t$ and the structure of message vectors yielding identical function values. While prior works established bounds, key questions remain, such as the optimal redundancy for functions like Hamming weight and Hamming weight distribution, along with efficient code constructions. In this paper, we make the following contributions:
(1) For the Hamming weight function, we improve the lower bound on optimal redundancy from $\frac{10(t-1)}{3}$ to $4t - \frac{4}{3}\sqrt{6t+2} + 2$. On the other hand, we provide a systematical approach to constructing explicit FCCs via a novel connection with Gray codes, which also improve the previous upper bound from $\frac{4t-2}{1 - 2\sqrt{\log{2t}/(2t)}}$ to $4t - \log{t}$. Consequently, we almost determine the optimal redundancy for Hamming weight function.
(2) The Hamming weight distribution function is defined by the value of Hamming weight divided by a given positive integer $T$. Previous work established that the optimal redundancy is $2t$ when $T > 2t$, while the case $T \le 2t$ remained unclear. We show that the optimal redundancy remains $2t$ when $T \ge t+1$. However, in the surprising regime where $T = o(t)$, we achieve near-optimal redundancy of $4t - o(t)$. Our results reveal a significant distinction in behavior of redundancy for distinct choices of $T$.
The Frankl-Pach upper bound is not tight for any uniformity
Published in J. Combin. Theory Ser. A 217 (2026), Paper No. 106078, 9pp
• View Publication
• BIB
For any positive integers $n\ge d+1\ge 3$, what is the maximum size of a $(d+1)$-uniform set system in $[n]$ with VC-dimension at most $d$? In 1984, Frankl and Pach initiated the study of this fundamental problem and provided an upper bound $\binom{n}{d}$ via an elegant algebraic proof. Surprisingly, in 2007, Mubayi and Zhao showed that when $n$ is sufficiently large and $d$ is a prime power, the Frankl-Pach upper bound is not tight. They also remarked that their method requires $d$ to be a prime power, and asked for new ideas to improve the Frankl-Pach upper bound without extra assumptions on $n$ and $d$.
In this paper, we provide an improvement for any $d\ge 2$ and $n\ge 2d+2$, which demonstrates that the long-standing Frankl-Pach upper bound $\binom{n}{d}$ is not tight for any uniformity. Our proof combines a simple yet powerful polynomial method and structural analysis.
Algebraic approach to stability results for Erdős-Ko-Rado theorem
Celebrated results often unfold like episodes in a long-running series. In the field of extremal set thoery, Erdős, Ko, and Rado in 1961 established that any $k$-uniform intersecting family on $[n]$ has a maximum size of $\binom{n-1}{k-1}$, with the unique extremal structure being a star. In 1967, Hilton and Milner followed up with a pivotal result, showing that if such a family is not a star, its size is at most $\binom{n-1}{k-1} - \binom{n-k-1}{k-1} + 1$, and they identified the corresponding extremal structures. In recent years, Han and Kohayakawa, Kostochka and Mubayi, and Huang and Peng have provided the second and third levels of stability results in this line of research.
In this paper, we provide a unified approach to proving the stability result for the Erdős-Ko-Rado theorem at any level. Our framework primarily relies on a robust linear algebra method, which leverages appropriate non-shadows to effectively handle the structural complexities of these intersecting families.
Constrained coding upper bounds via Goulden-Jackson cluster theorem
Motivated by applications in DNA-based data storage, constrained codes have attracted a considerable amount of attention from both academia and industry. We study the maximum cardinality of constrained codes for which the constraints can be characterized by a set of forbidden substrings, where by a substring we mean some consecutive coordinates in a string.
For finite-type constrained codes (for which the set of forbidden substrings is finite), one can compute their capacity (code rate) by the ``spectral method'', i.e., by applying the Perron-Frobenious theorem to the de Brujin graph defined by the code. However, there was no systematic method to compute the exact cardinality of these codes.
We show that there is a surprisingly powerful method arising from enumerative combinatorics, which is based on the Goulden-Jackson cluster theorem (previously not known to the coding community), that can be used to compute not only the capacity, but also the exact formula for the cardinality of these codes, for each fixed code length. Moreover, this can be done by solving a system of linear equations of size equal to the number of constraints.
We also show that the spectral method and the cluster method are inherently related by establishing a direct connection between the spectral radius of the de Brujin graph used in the first method and the convergence radius of the generating function used in the second method.
Lastly, to demonstrate the flexibility of the new method, we use it to give an explicit upper bound on the maximum cardinality of variable-length non-overlapping codes, which are a class of constrained codes defined by an infinite number of forbidden substrings.
On set systems with strongly restricted intersections
Set systems with strongly restricted intersections, called $α$-intersecting families for a vector $α$, were introduced recently as a generalization of several well-studied intersecting families including the classical oddtown and eventown. Given a binary vector $α=(a_1, \ldots, a_k)$, a collection $\mathcal F$ of subsets over an $n$ element set is an $α$-intersecting family modulo $2$ if for each $i=1,2,\ldots,k$, all $i$-wise intersections of distinct members in $\mathcal F$ have sizes with the same parity as $a_i$. Let $f_α(n)$ denote the maximum size of such a family. In this paper, we study the asymptotic behavior of $f_α(n)$ when $n$ goes to infinity. We show that if $t$ is the maximum integer such that $a_t=1$ and $2t\leq k$, then $f_{α(n)} \sim {(t! n)}^{\frac 1 t}$. More importantly, we show that for any constant $c$, as the length $k$ goes larger, $f_α(n)$ is upper bounded by $O (n^c)$ for almost all $α$. Equivalently, no matter what $k$ is, there are only finitely many $α$ satisfying $f_α(n)=Ω(n^c)$. This answers an open problem raised by Johnston and O'Neill in 2023. All of our results can be generalized to modulo $p$ setting for any prime $p$ smoothly.
New results on sparse representations in unions of orthonormal bases
The problem of sparse representation has significant applications in signal processing. The spark of a dictionary plays a crucial role in the study of sparse representation. Donoho and Elad initially explored the spark, and they provided a general lower bound. When the dictionary is a union of several orthonormal bases, Gribonval and Nielsen presented an improved lower bound for spark. In this paper, we introduce a new construction of dictionary, achieving the spark bound given by Gribonval and Nielsen. Our result extends Shen et al.' s findings [IEEE Trans. Inform. Theory, vol. 68, pp. 4230--4243, 2022].
Swap-Robust and Almost Supermagic Complete Graphs for Dynamical Distributed Storage
To prevent service time bottlenecks in distributed storage systems, the access balancing problem has been studied by designing almost supermagic edge labelings of certain graphs to balance the access requests to different servers. In this paper, we introduce the concept of robustness of edge labelings under limited-magnitude swaps, which is important for studying the dynamical access balancing problem with respect to changes in data popularity. We provide upper and lower bounds on the robustness ratio for complete graphs with $n$ vertices, and construct $O(n)$-almost supermagic labelings that are asymptotically optimal in terms of the robustness ratio.
Separating hash families with large universe
Separating hash families are useful combinatorial structures which generalize several well-studied objects in cryptography and coding theory. Let $p_t(N, q)$ denote the maximum size of universe for a $t$-perfect hash family of length $N$ over an alphabet of size $q$. In this paper, we show that $q^{2-o(1)}<p_t(t, q)=o(q^2)$ for all $t\geq 3$, which answers an open problem about separating hash families raised by Blackburn et al. in 2008 for certain parameters. Previously, this result was known only for $t=3, 4$. Our proof is obtained by establishing the existence of a large set of integers avoiding nontrivial solutions to a set of correlated linear equations.
A new variant of the Erdős-Gyárfás problem on $K_{5}$
Motivated by an extremal problem on graph-codes that links coding theory and graph theory, Alon recently proposed a question aiming to find the smallest number $t$ such that there is an edge coloring of $K_{n}$ by $t$ colors with no copy of given graph $H$ in which every color appears an even number of times. When $H=K_{4}$, the question of whether $n^{o(1)}$ colors are enough, was initially emphasized by Alon. Through modifications to the coloring functions originally designed by Mubayi, and Conlon, Fox, Lee and Sudakov, the question of $K_{4}$ has already been addressed. Expanding on this line of inquiry, we further study this new variant of the generalized Ramsey problem and provide a conclusively affirmative answer to Alon's question concerning $K_{5}$.
New constructions of signed difference sets
Signed difference sets have interesting applications in communications and coding theory. A $(v,k,λ)$-difference set in a finite group $G$ of order $v$ is a subset $D$ of $G$ with $k$ distinct elements such that the expressions $xy^{-1}$ for all distinct two elements $x,y\in D$, represent each non-identity element in $G$ exactly $λ$ times. A $(v,k,λ)$-signed difference set is a generalization of a $(v,k,λ)$-difference set $D$, which satisfies all properties of $D$, but has a sign for each element in $D$. We will show some new existence results for signed difference sets by using partial difference sets, product methods, and cyclotomic classes.
Reconstruction of Sequences Distorted by Two Insertions
Reconstruction codes are generalizations of error-correcting codes that can correct errors by a given number of noisy reads. The study of such codes was initiated by Levenshtein in 2001 and developed recently due to applications in modern storage devices such as racetrack memories and DNA storage. The central problem on this topic is to design codes with redundancy as small as possible for a given number $N$ of noisy reads. In this paper, the minimum redundancy of such codes for binary channels with exactly two insertions is determined asymptotically for all values of $N\ge 5$. Previously, such codes were studied only for channels with single edit errors or two-deletion errors.
On supersaturation for oddtown and eventown
We study the supersaturation problems of oddtown and eventown. Given a family $\mathcal A$ of subsets of an $n$ element set, let $op(\mathcal A)$ denote the number of distinct pairs $A,B\in \mathcal A$ for which $|A \cap B|$ is odd. We show that if $\mathcal A$ consists of $n+s$ odd-sized subsets, then $op(\mathcal A)\geq s+2$, which is tight when $s\le n-4$. This disproves a conjecture by O'Neill on the supersaturation problem of oddtown. For the supersaturation problem of eventown, we show that for large enough $n$, if $\mathcal A$ consists of $2^{\lfloor \frac n 2\rfloor}+s$ even-sized subsets, then $op(\mathcal A)\ge s\cdot2^{\lfloor \frac n 2\rfloor-1} $ for any positive integer $s\le \frac{2^{\lfloor\frac n 8\rfloor}} n$. This partially proves a conjecture by O'Neill on the supersaturation problem of eventown. Previously, the correctness of this conjecture was only verified for $s=1$ and $2$. We further provide a twice weaker lower bound in this conjecture for eventown, that is $op(\mathcal{A})\ge s\cdot 2^{\lfloor n/2\rfloor-2}$ for general $n$ and $s$ by using discrete Fourier analysis. Finally, some asymptotic results for the lower bounds of $op(\mathcal A)$ are given when $s$ is large for both problems.
Some results on similar configurations in subsets of $\mathbb{F}_q^d$
In this paper, we study problems about the similar configurations in $\mathbb{F}_q^d$. Let $G=(V, E)$ be a graph, where $V=\{1, 2, \ldots, n\}$ and $E\subseteq{V\choose2}$. For a set $\mathcal{E}$ in $\mathbb{F}_q^d$, we say that $\mathcal{E}$ contains a pair of $G$ with dilation ratio $r$ if there exist distinct $\boldsymbol{x}_1, \boldsymbol{x}_2, \ldots, \boldsymbol{x}_n\in\mathcal{E}$ and distinct $\boldsymbol{y}_1, \boldsymbol{y}_2, \ldots, \boldsymbol{y}_n\in\mathcal{E}$ such that $\|\boldsymbol{y}_i-\boldsymbol{y}_{j}\|=r\|\boldsymbol{x}_i-\boldsymbol{x}_j\|\neq0$ whenever $\{i, j\}\in E$, where $\|\boldsymbol{x}\|:=x_1^2+x_2^2+\cdots+x_d^2$ for $\boldsymbol{x}=(x_1, x_2, \ldots, x_d)\in\mathbb{F}_q^d$. We show that if $\mathcal{E}$ has size at least $C_kq^{d/2}$, then $\mathcal{E}$ contains a pair of $k$-stars with dilation ratio $r$, and that if $\mathcal{E}$ has size at least $C\cdot\min\left\{q^{(2d+1)/3}, \max\left\{q^3, q^{d/2}\right\}\right\}$, then $\mathcal{E}$ contains a pair of $4$-paths with dilation ratio $r$. Our method is based on enumerative combinatorics and graph theory.
On lattice tilings of $\mathbb{Z}^{n}$ by limited magnitude error balls $\mathcal{B}(n,2,1,1)$
Limited magnitude error model has applications in flash memory. In this model, a perfect code is equivalent to a tiling of $\mathbb{Z}^n$ by limited magnitude error balls. In this paper, we give a complete classification of lattice tilings of $\mathbb{Z}^n$ by limited magnitude error balls $\mathcal{B}(n,2,1,1)$.
Improved Gilbert-Varshamov bounds for hopping cyclic codes and optical orthogonal codes
Hopping cyclic codes (HCCs) are (non-linear) cyclic codes with the additional property that the $n$ cyclic shifts of every given codeword are all distinct, where $n$ is the code length. Constant weight binary hopping cyclic codes are also known as optical orthogonal codes (OOCs). HCCs and OOCs have various practical applications and have been studied extensively over the years.
The main concern of this paper is to present improved Gilbert-Varshamov type lower bounds for these codes, when the minimum distance is bounded below by a linear factor of the code length. For HCCs, we improve the previously best known lower bound of Niu, Xing, and Yuan by a linear factor of the code length. For OOCs, we improve the previously best known lower bound of Chung, Salehi, and Wei, and Yang and Fuja by a quadratic factor of the code length. As by-products, we also provide improved lower bounds for frequency hopping sequences sets and error-correcting weakly mutually uncorrelated codes. Our proofs are based on tools from probability theory and graph theory, in particular the McDiarmid's inequality on the concentration of Lipschitz functions and the independence number of locally sparse graphs.
On linear diameter perfect Lee codes with diameter 6
Published
• View Publication
• BIB
In 1968, Golomb and Welch conjectured that there is no perfect Lee codes with radius $r\ge2$ and dimension $n\ge3$. A diameter perfect code is a natural generalization of the perfect code. In 2011, Etzion (IEEE Trans. Inform. Theory, 57(11): 7473--7481, 2011) proposed the following problem: Are there diameter perfect Lee (DPL, for short) codes with diameter greater than four besides the $DPL(3,6)$ code? Later, Horak and AlBdaiwi (IEEE Trans. Inform. Theory, 58(8): 5490--5499, 2012) conjectured that there are no $DPL(n,d)$ codes for dimension $n\ge3$ and diameter $d>4$ except for $(n,d)=(3,6)$. In this paper, we give a counterexample to this conjecture. Moreover, we prove that for $n\ge3$, there is a linear $DPL(n,6)$ code if and only if $n=3,11$.
A generic framework for coded caching and distributed computation schemes
Published
• View Publication
• BIB
Several network communication problems are highly related such as coded caching and distributed computation. The centralized coded caching focuses on reducing the network burden in peak times in a wireless network system and the coded distributed computation studies the tradeoff between computation and communication in distributed system. In this paper, motivated by the study of the only rainbow $3$-term arithmetic progressions set, we propose a unified framework for constructing coded caching schemes. This framework builds bridges between coded caching schemes and lots of combinatorial objects due to the freedom of the choices of families and operations. We prove that any scheme based on a placement delivery array (PDA) can be represented by a rainbow scheme under this framework and lots of other known schemes can also be included in this framework. Moreover, we also present a new coded caching scheme with linear subpacketization and near constant rate using the only rainbow $3$-term arithmetic progressions set. Next, we modify the framework to be applicable to the distributed computing problem. We present a new transmission scheme in the shuffle phase and show that in certain cases it could have a lower communication load than the schemes based on PDAs or resolvable designs with the same number of files.
On the lower bound for kissing numbers of $\ell_p$-spheres in high dimensions
In this paper, we give some new lower bounds for the kissing number of $\ell_p$-spheres. These results improve the previous work due to Xu (2007). Our method is based on coding theory.
A note on the Turán number for the traces of hypergraphs
Let $\mathcal{H}$ be an $r$-uniform hypergraph and $F$ be a graph. We say $\mathcal{H}$ contains $F$ as a trace if there exists some set $S\subseteq V(\mathcal{H})$ such that $\mathcal{H}|_{S}:=\{E\cap S: E\in E(\mathcal{H})\}$ contains a subgraph isomorphic to $F.$ Let $ex_r(n,Tr(F))$ denote the maximum number of edges of an $n$-vertex $r$-uniform hypergraph $\mathcal{H}$ which does not contain $F$ as a trace. In this paper, we improve the lower bounds of $ex_r(n,Tr(F))$ when $F$ is a star, and give some optimal cases. We also improve the upper bound for the case when $\mathcal{H}$ is $3$-uniform and $F$ is $K_{2,t}$ when $t$ is small.