Papers by Yiwei Zhang
5 paper(s) by this author
· All BibTeX
Batch Array Codes
Batch codes are a type of codes specifically designed for coded distributed storage systems and private information retrieval protocols. These codes have got much attention in recent years due to their ability to enable efficient and secure storage in distributed systems.
In this paper, we study an array code version of the batch codes, which is called the \emph{batch array code} (BAC). Under the setting of BAC, each node stores a bucket containing multiple code symbols and responds with a locally computed linear combination of the symbols in its bucket during the recovery of a requested symbol. We demonstrate that BACs can support the same type of requests as the original batch codes but with reduced redundancy. Specifically, we establish information theoretic lower bounds on the code lengths and provide several code constructions that confirm the tightness of the lower bounds for certain parameter regimes.
SPIDER-WEB generates coding algorithms with superior error tolerance and real-time information retrieval capacity
DNA has been considered a promising medium for storing digital information. As an essential step in the DNA-based data storage workflow, coding algorithms are responsible to implement functions including bit-to-base transcoding, error correction, etc. In previous studies, these functions are normally realized by introducing multiple algorithms. Here, we report a graph-based architecture, named SPIDER-WEB, providing an all-in-one coding solution by generating customized algorithms automatically. SPIDERWEB is able to correct a maximum of 4% edit errors in the DNA sequences including substitution and insertion/deletion (indel), with only 5.5% redundant symbols. Since no DNA sequence pretreatment is required for the correcting and decoding processes, SPIDER-WEB offers the function of real-time information retrieval, which is 305.08 times faster than the speed of single-molecule sequencing techniques. Our retrieval process can improve 2 orders of magnitude faster compared to the conventional one under megabyte-level data and can be scalable to fit exabyte-level data. Therefore, SPIDER-WEB holds the potential to improve the practicability in large-scale data storage applications.
New Theoretical Bounds and Constructions of Permutation Codes under Block Permutation Metric
Published
• View Publication
• BIB
Permutation codes under different metrics have been extensively studied due to their potentials in various applications. Generalized Cayley metric is introduced to correct generalized transposition errors, including previously studied metrics such as Kendall's $τ$-metric, Ulam metric and Cayley metric as special cases. Since the generalized Cayley distance between two permutations is not easily computable, Yang et al. introduced a related metric of the same order, named the block permutation metric. Given positive integers $n$ and $d$, let $\mathcal{C}_{B}(n,d)$ denote the maximum size of a permutation code in $S_n$ with minimum block permutation distance $d$. In this paper, we focus on the theoretical bounds of $\mathcal{C}_{B}(n,d)$ and the constructions of permutation codes under block permutation metric. Using a graph theoretic approach, we improve the Gilbert-Varshamov type bound by a factor of $Ω(\log{n})$, when $d$ is fixed and $n$ goes into infinity. We also propose a new encoding scheme based on binary constant weight codes. Moreover, an upper bound beating the sphere-packing type bound is given when $d$ is relatively close to $n$.
Invertible binary matrix with maximum number of $2$-by-$2$ invertible submatrices
Published
• View Publication
• BIB
The problem is related to all-or-nothing transforms (AONT) suggested by Rivest as a preprocessing for encrypting data with a block cipher. Since then there have been various applications of AONTs in cryptography and security. D'Arco, Esfahani and Stinson posed the problem on the constructions of binary matrices for which the desired properties of an AONT hold with the maximum probability. That is, for given integers $t\le s$, what is the maximum number of $t$-by-$t$ invertible submatrices in a binary matrix of order $s$? For the case $t=2$, let $R_2(s)$ denote the maximal proportion of 2-by-2 invertible submatrices. D'Arco, Esfahani and Stinson conjectured that the limit is between 0.492 and 0.625. We completely solve the case $t=2$ by showing that $\lim_{s\rightarrow\infty}R_2(s)=0.5$.
On the mixing properties of piecewise expanding maps under composition with permutations
Published
• View Publication
• BIB
We consider the effect on the mixing properties of a piecewise smooth interval map $f$ when its domain is divided into $N$ equal subintervals and $f$ is composed with a permutation of these. The case of the stretch-and-fold map $f(x)=mx \bmod 1$ for integers $m \geq 2$ is examined in detail. We give a combinatorial description of those permutations $σ$ for which $σ\circ f$ is still (topologically) mixing, and show that the proportion of such permutations tends to $1$ as $N \to \infty$. We then investigate the mixing rate of $σ\circ f$ (as measured by the modulus of the second largest eigenvalue of the transfer operator). In contrast to the situation for continuous time diffusive systems, we show that composition with a permutation cannot improve the mixing rate of $f$, but typically makes it worse. Under some mild assumptions on $m$ and $N$, we obtain a precise value for the worst mixing rate as $σ$ ranges through all permutations; this can be made arbitrarily close to $1$ as $N \to \infty$ (with $m$ fixed). We illustrate the geometric distribution of the second largest eigenvalues in the complex plane for small $m$ and $N$, and propose a conjecture concerning their location in general. Finally, we give examples of other interval maps $f$ for which composition with permutations produces different behaviour than that obtained from the stretch-and-fold map.