arXiv++ Combinatorics

Browse math.CO papers from arXiv

Papers by Christian M. Reidys

74 paper(s) by this author · All BibTeX
2011-04-17 v3
The 5'-3' distance of RNA secondary structures
Published • View PublicationBIB
Recently Yoffe {\it et al.} observed that the average distances between $5'-3'$ ends of RNA molecules are very small and largely independent of sequence length. This observation is based on numerical computations as well as theoretical arguments maximizing certain entropy functionals. In this paper we compute the exact distribution of $5'-3'$ distances of RNA secondary structures for any finite $n$. We furthermore compute the limit distribution and show that already for $n=30$ the exact distribution and the limit distribution are very close. Our results show that the distances of random RNA secondary structures are distinctively lower than those of minimum free energy structures of random RNA sequences.
Linear chord diagrams on two intervals
Consider all possible ways of attaching disjoint chords to two ordered and oriented disjoint intervals so as to produce a connected graph. Taking the intervals to lie in the real axis with the induced orientation and the chords to lie in the upper half plane canonically determines a corresponding fatgraph which has some associated genus $g\geq 0$, and we consider the natural generating function ${\bf C}_g^{[2]}(z)=\sum_{n\geq 0} {\bf c}^{[2]}_g(n)z^n$ for the number ${\bf c}^{[2]}_g(n)$ of distinct such chord diagrams of fixed genus $g\geq 0$ with a given number $n\geq 0$ of chords. We prove here the surprising fact that ${\bf C}^{[2]}_g(z)=z^{2g+1} R_g^{[2]}(z)/(1-4z)^{3g+2} $ is a rational function, for $g\geq 0$, where the polynomial $R^{[2]}_g(z)$ with degree at most $g$ has integer coefficients and satisfies $R_g^{[2]}({1\over 4})\neq 0$. Earlier work had already determined that the analogous generating function ${\bf C}_g(z)=z^{2g}R_g(z)/(1-4z)^{3g-{1\over 2}}$ for chords attached to a single interval is algebraic, for $g\geq 1$, where the polynomial $R_g(z)$ with degree at most $g-1$ has integer coefficients and satisfies $R_g(1/4)\neq 0$ in analogy to the generating function ${\bf C}_0(z)$ for the Catalan numbers. The new results here on ${\bf C}_g^{[2]}(z)$ rely on this earlier work, and indeed, we find that $R_g^{[2]}(z)=R_{g+1}(z) -z\sum_{g_1=1}^g R_{g_1}(z) R_{g+1-g_1}(z)$, for $g\geq 1$.
2010-06-21
Combinatorial analysis of interacting RNA molecules
Published • View PublicationBIB
Recently several minimum free energy (MFE) folding algorithms for predicting the joint structure of two interacting RNA molecules have been proposed. Their folding targets are interaction structures, that can be represented as diagrams with two backbones drawn horizontally on top of each other such that (1) intramolecular and intermolecular bonds are noncrossing and (2) there is no "zig-zag" configuration. This paper studies joint structures with arc-length at least four in which both, interior and exterior stack-lengths are at least two (no isolated arcs). The key idea in this paper is to consider a new type of shape, based on which joint structures can be derived via symbolic enumeration. Our results imply simple asymptotic formulas for the number of joint structures with surprisingly small exponential growth rates. They are of interest in the context of designing prediction algorithms for RNA-RNA interactions.
2010-06-15
On the uniform generation of modular diagrams
Published • View PublicationBIB
In this paper we present an algorithm that generates $k$-noncrossing, $σ$-modular diagrams with uniform probability. A diagram is a labeled graph of degree $\le 1$ over $n$ vertices drawn in a horizontal line with arcs $(i,j)$ in the upper half-plane. A $k$-crossing in a diagram is a set of $k$ distinct arcs $(i_1, j_1), (i_2, j_2),\ldots,(i_k, j_k)$ with the property $i_1 < i_2 < \ldots < i_k < j_1 < j_2 < \ldots< j_k$. A diagram without any $k$-crossings is called a $k$-noncrossing diagram and a stack of length $σ$ is a maximal sequence $((i,j),(i+1,j-1),\dots,(i+(σ-1),j-(σ-1)))$. A diagram is $σ$-modular if any arc is contained in a stack of length at least $σ$. Our algorithm generates after $O(n^k)$ preprocessing time, $k$-noncrossing, $σ$-modular diagrams in $O(n)$ time and space complexity.
2010-06-15
Combinatorics of RNA-RNA interaction
Published • View PublicationBIB
RNA-RNA binding is an important phenomenon observed for many classes of non-coding RNAs and plays a crucial role in a number of regulatory processes. Recently several MFE folding algorithms for predicting the joint structure of two interacting RNA molecules have been proposed. Here joint structure means that in a diagram representation the intramolecular bonds of each partner are pseudoknot-free, that the intermolecular binding pairs are noncrossing, and that there is no so-called ``zig-zag'' configuration. This paper presents the combinatorics of RNA interaction structures including their generating function, singularity analysis as well as explicit recurrence relations. In particular, our results imply simple asymptotic formulas for the number of joint structures.
2010-03-13 v2
Modular, $k$-noncrossing diagrams
Published • View PublicationBIB
In this paper we compute the generating function of modular, $k$-noncrossing diagrams. A $k$-noncrossing diagram is called modular if it does not contains any isolated arcs and any arc has length at least four. Modular diagrams represent the deformation retracts of RNA pseudoknot structures \cite{Stadler:99,Reidys:07pseu,Reidys:07lego} and their properties reflect basic features of these bio-molecules. The particular case of modular noncrossing diagrams has been extensively studied \cite{Waterman:78b, Waterman:79,Waterman:93, Schuster:98}. Let ${\sf Q}_k(n)$ denote the number of modular $k$-noncrossing diagrams over $n$ vertices. We derive exact enumeration results as well as the asymptotic formula ${\sf Q}_k(n)\sim c_k n^{-(k-1)^2-\frac{k-1}{2}}γ_{k}^{-n}$ for $k=3,..., 9$ and derive a new proof of the formula ${\sf Q}_2(n)\sim 1.4848\, n^{-3/2}\,1.8489^{-n}$ \cite{Schuster:98}.
Inverse Folding of RNA Pseudoknot Structures
Published • View PublicationBIB
Background: RNA exhibits a variety of structural configurations. Here we consider a structure to be tantamount to the noncrossing Watson-Crick and \pairGU-base pairings (secondary structure) and additional cross-serial base pairs. These interactions are called pseudoknots and are observed across the whole spectrum of RNA functionalities. In the context of studying natural RNA structures, searching for new ribozymes and designing artificial RNA, it is of interest to find RNA sequences folding into a specific structure and to analyze their induced neutral networks. Since the established inverse folding algorithms, {\tt RNAinverse}, {\tt RNA-SSD} as well as {\tt INFO-RNA} are limited to RNA secondary structures, we present in this paper the inverse folding algorithm {\tt Inv} which can deal with 3-noncrossing, canonical pseudoknot structures. Results: In this paper we present the inverse folding algorithm {\tt Inv}. We give a detailed analysis of {\tt Inv}, including pseudocodes. We show that {\tt Inv} allows to design in particular 3-noncrossing nonplanar RNA pseudoknot 3-noncrossing RNA structures--a class which is difficult to construct via dynamic programming routines. {\tt Inv} is freely available at \url{http://www.combinatorics.cn/cbpc/inv.html}. Conclusions: The algorithm {\tt Inv} extends inverse folding capabilities to RNA pseudoknot structures. In comparison with {\tt RNAinverse} it uses new ideas, for instance by considering sets of competing structures. As a result, {\tt Inv} is not only able to find novel sequences even for RNA secondary structures, it does so in the context of competing structures that potentially exhibit cross-serial interactions.
2010-03-03
The evolution of random reversal graph
Published • View PublicationBIB
The random reversal graph offers new perspectives, allowing to study the connectivity of genomes as well as their most likely distance as a function of the reversal rate. Our main result shows that the structure of the random reversal graph changes dramatically at $λ_n=1/\binom{n+1}{2}$. For $λ_n=(1-ε)/\binom{n+1}{2}$, the random graph consists of components of size at most $O(n\ln(n))$ a.s. and for $(1+ε)/\binom{n+1}{2}$, there emerges a unique largest component of size $\sim \wp(ε) \cdot 2^n\cdot n$!$ a.s.. This "giant" component is furthermore dense in the reversal graph.
Loops in canonical RNA pseudoknot structures
Published • View PublicationBIB
In this paper we compute the limit distributions of the numbers of hairpin-loops, interior-loops and bulges in k-noncrossing RNA structures. The latter are coarse grained RNA structures allowing for cross-serial interactions, subject to the constraint that there are at most k-1 mutually crossing arcs in the diagram representation of the molecule. We prove central limit theorems by means of studying the corresponding bivariate generating functions. These generating functions are obtained by symbolic inflation of Ik5-shapes.
2009-11-16
Random $k$-noncrossing partitions
In this paper, we introduce polynomial time algorithms that generate random $k$-noncrossing partitions and 2-regular, $k$-noncrossing partitions with uniform probability. A $k$-noncrossing partition does not contain any $k$ mutually crossing arcs in its canonical representation and is 2-regular if the latter does not contain arcs of the form $(i,i+1)$. Using a bijection of Chen {\it et al.} \cite{Chen,Reidys:08tan}, we interpret $k$-noncrossing partitions and 2-regular, $k$-noncrossing partitions as restricted generalized vacillating tableaux. Furthermore, we interpret the tableaux as sampling paths of a Markov-processes over shapes and derive their transition probabilities.
2009-10-14
Random 3-noncrossing partitions
In this paper, we introduce polynomial time algorithms that generate random 3-noncrossing partitions and 2-regular, 3-noncrossing partitions with uniform probability. A 3-noncrossing partition does not contain any three mutually crossing arcs in its canonical representation and is 2-regular if the latter does not contain arcs of the form $(i,i+1)$. Using a bijection of Chen {\it et al.} \cite{Chen,Reidys:08tan}, we interpret 3-noncrossing partitions and 2-regular, 3-noncrossing partitions as restricted generalized vacillating tableaux. Furthermore, we interpret the tableaux as sampling paths of Markov-processes over shapes and derive their transition probabilities.
2009-09-22 v2
Random induced subgraphs of Cayley graphs induced by transpositions
Published • View PublicationBIB
In this paper we study random induced subgraphs of Cayley graphs of the symmetric group induced by an arbitrary minimal generating set of transpositions. A random induced subgraph of this Cayley graph is obtained by selecting permutations with independent probability, $λ_n$. Our main result is that for any minimal generating set of transpositions, for probabilities $λ_n=\frac{1+ε_n}{n-1}$ where $n^{-{1/3}+δ}\le ε_n<1$ and $δ>0$, a random induced subgraph has a.s. a unique largest component of size $\wp(ε_n)\frac{1+ε_n}{n-1}n!$, where $\wp(ε_n)$ is the survival probability of a specific branching process.
Random $k$-noncrossing RNA Structures
Published • View PublicationBIB
In this paper we derive polynomial time algorithms that generate random $k$-noncrossing matchings and $k$-noncrossing RNA structures with uniform probability. Our approach employs the bijection between $k$-noncrossing matchings and oscillating tableaux and the $P$-recursiveness of the cardinalities of $k$-noncrossing matchings. The main idea is to consider the tableaux sequences as paths of stochastic processes over shapes and to derive their transition probabilities.
2009-06-22 v2
Shapes of RNA pseudoknot structures
Published • View PublicationBIB
In this paper we study abstract shapes of $k$-noncrossing, $σ$-canonical RNA pseudoknot structures. We consider ${\sf lv}_k^{\sf 1}$- and ${\sf lv}_k^{\sf 5}$-shapes, which represent a generalization of the abstract $π'$- and $π$-shapes of RNA secondary structures introduced by \citet{Giegerich:04ashape}. Using a novel approach we compute the generating functions of ${\sf lv}_k^{\sf 1}$- and ${\sf lv}_k^{\sf 5}$-shapes as well as the generating functions of all ${\sf lv}_k^{\sf 1}$- and ${\sf lv}_k^{\sf 5}$-shapes induced by all $k$-noncrossing, $σ$-canonical RNA structures for fixed $n$. By means of singularity analysis of the generating functions, we derive explicit asymptotic expressions.
2009-06-18
A generalization of the brauer algebra
We study two variations of the Brauer algebra $B_n(x)$. The first is the algebra $A_n(x)$, which generalizes the Brauer algebra by considering loops. The second is the algebra $L_n(x)$, the $A_n(x)$-subalgebra generated by diagrams without horizontal arcs. $A_n(x)$ and $L_n(x)$ have for $x \neq 0$ an hereditary-chain indexed by all integers. Following the ideas of Martin in the context of the partition algebra, and Doran et al. for the Brauer algebra, we study semisimplicity of $A_n(x)$ using restriction and induction in $A_n(x)$ and $L_n(x)$. Our main result is that $A_n(x)$ is semisimple if $x \notin Z$ and that $L_n(x)$ is semisimple if $x \neq 0$.
Inverse folding of RNA pseudoknot structures
Published • View PublicationBIB
Background: RNA exhibits a variety of structural configurations. Here we consider a structure to be tantamount to the noncrossing Watson-Crick and \pairGU-base pairings (secondary structure) and additional cross-serial base pairs. These interactions are called pseudoknots and are observed across the whole spectrum of RNA functionalities. In the context of studying natural RNA structures, searching for new ribozymes and designing artificial RNA, it is of interest to find RNA sequences folding into a specific structure and to analyze their induced neutral networks. Since the established inverse folding algorithms, {\tt RNAinverse}, {\tt RNA-SSD} as well as {\tt INFO-RNA} are limited to RNA secondary structures, we present in this paper the inverse folding algorithm {\tt Inv} which can deal with 3-noncrossing, canonical pseudoknot structures. Results: In this paper we present the inverse folding algorithm {\tt Inv}. We give a detailed analysis of {\tt Inv}, including pseudocodes. We show that {\tt Inv} allows to design in particular 3-noncrossing nonplanar RNA pseudoknot 3-noncrossing RNA structures-a class which is difficult to construct via dynamic programming routines. {\tt Inv} is freely available at \url{http://www.combinatorics.cn/cbpc/inv.html}. Conclusions: The algorithm {\tt Inv} extends inverse folding capabilities to RNA pseudoknot structures. In comparison with {\tt RNAinverse} it uses new ideas, for instance by considering sets of competing structures. As a result, {\tt Inv} is not only able to find novel sequences even for RNA secondary structures, it does so in the context of competing structures that potentially exhibit cross-serial interactions.
RNA-RNA interaction prediction: partition function and base pair pairing probabilities
Published • View PublicationBIB
In this paper, we study the interaction of an antisense RNA and its target mRNA, based on the model introduced by Alkan {\it et al.} (Alkan {\it et al.}, J. Comput. Biol., Vol:267--282, 2006). Our main results are the derivation of the partition function \cite{Backhofen} (Chitsaz {\it et al.}, Bioinformatics, to appear, 2009), based on the concept of tight-structure and the computation of the base pairing probabilities. This paper contains the folding algorithm {\sf rip} which computes the partition function as well as the base pairing probabilities in $O(N^4M^2)+O(N^2M^4)$ time and $O(N^2M^2)$ space, where $N,M$ denote the lengths of the interacting sequences.
2009-02-22
On the decomposition of $k$-noncrossing RNA structures
Published • View PublicationBIB
An $k$-noncrossing RNA structure can be identified with an $k$-noncrossing diagram over $[n]$, which in turn corresponds to a vacillating tableaux having at most $(k-1)$ rows. In this paper we derive the limit distribution of irreducible substructures via studying their corresponding vacillating tableaux. Our main result proves, that the limit distribution of the numbers of irreducible substructures in $k$-noncrossing, $σ$-canonical RNA structures is determined by the density function of a $Γ(-\lnτ_k,2)$-distribution for some $τ_k<1$.
2009-02-22
Irreducibility in RNA structures
Published • View PublicationBIB
In this paper we study irreducibility in RNA structures. By RNA structure we mean RNA secondary as well as RNA pseudoknot structures. In our analysis we shall contrast random and minimum free energy (mfe) configurations. We compute various distributions: of the numbers of irreducible substructures, their locations and sizes, parameterized in terms of the maximal number of mutually crossing arcs, $k-1$, and the minimal size of stacks $σ$. In particular, we analyze the size of the largest irreducible substructure for random and mfe structures, which is the key factor for the folding time of mfe configurations.
Folding 3-noncrossing RNA pseudoknot structures
Published • View PublicationBIB
In this paper we present a selfcontained analysis and description of the novel {\it ab initio} folding algorithm {\sf cross}, which generates the minimum free energy (mfe), 3-noncrossing, $σ$-canonical RNA structure. Here an RNA structure is 3-noncrossing if it does not contain more than three mutually crossing arcs and $σ$-canonical, if each of its stacks has size greater or equal than $σ$. Our notion of mfe-structure is based on a specific concept of pseudoknots and respective loop-based energy parameters. The algorithm decomposes into three parts: the first is the inductive construction of motifs and shadows, the second is the generation of the skeleta-trees rooted in irreducible shadows and the third is the saturation of skeleta via context dependent dynamic programming routines.