Info pre-processing and normalization == Our normalization procedure modifies for array-level variation inside the distribution in probe features due to, for instance , differences in laser power or background intensity. growing repertoire of ucRBPs. Keywords: RNAcompete, RNA-binding protein, DNA microarray, binding site, NUDT21, CNBP == 1 . Introduction == Hundreds of thousands of annotated RBPs are encoded by the 742 sequenced eukaryotic genomes. The vast majority of these RBPs, including those from well-studied organisms, have unknown RNA-binding preferences. For example , in human, recent evidence indicates that there are approximately 1, 2001, 500 human proteins that associate with RNA, representing approximately 68% of the annotated human proteome [1, 2]; however , less than half of human RBPs that contain canonical RBDs (see below) have an established RNA-binding motif [3, 4]. This poor characterization of RBP sequence-binding preferences presents a significant barrier in the analysis of post-transcriptional gene regulation. Most known sequence-specific RBPs contain canonical RBDs that mediate RNA-binding through protein-RNA interactions [5]. The most common sequence-specific eukaryotic RBDs are the RNA recognition motif (RRM) (~246 in human), the Cys-Cys-Cys-His (CCCH-zf) type zinc finger domain (~60 in human), and the HNRNP K homology (KH) domain (~38 in human) [6]. Among these, the ~90 amino acid RRM domain is the most extensively studied [5]. Here, RNA recognition typically occurs on the -sheet surface and is often mediated by three exposed aromatic H-1152 residues [7]. Eukaryotic KH domains span ~70 amino acids and contain an RNA-binding cleft formed by a conserved GXXG motif flanked by two -helices, a variable loop, and a -strand [8]. Lastly, CCCH-zf domains are 1230 amino acids long and bind RNA through stacking interactions, and hydrogen bonding between the protein backbone and Watson-Crick edges of bases [9]. Individual RBDs tend to bind RNA in the micromolar range [5]; however RNA-binding affinity and specificity is significantly increased by combinations of RBDs or by additional RNA contacting residues located in regions outside of the RBD [5]. Not all sequence-specific RBPs contain a canonical RBD. For example , several well-characterized proteins, including the pre-mRNA 3 cleavage and polyadenylation specificity factor 5, NUDT21 (Nudix hydrolase domain and N-terminal region: [10]), histone stem-loop binding protein, SLBP (SLBP domain: [11, 12]), andDrosophilabrain tumor protein, BRAT (NHL domain: [13, 14]), use unconventional RBDsindicated in parenthesesfor sequence-specific RNA recognition. In a genome-wide context, proteomic analyses of proteins UV cross-linked to mRNA in human HeLa, HEK293, and HuH7 cell lines have identified over 800 that associate with mRNA, but do not contain canonical RBDs [2, 15, 16]. We refer to MAPK3 these potentially sequence-specific RBPs asunconventional RBPs (ucRBPs). A similar analysis inDrosophilaidentified ~300 ucRBPs, one of which (CG3800) was shown by CLIP-seq to bind RNA containing specific 5-mer sequences enriched for G and A residues [17]. It was not shown, however , whether CG3800 binds to these sequences autonomously. RNAcompete is anin vitromethod (seeFigure 1) that we developed [18] and have applied to hundreds of RBPs [3]. It recapitulates RNA-binding motifs for a diverse set of well-studied RBPs previously identifiedin vitro(e. g. SELEX experiments: Systematic Evolution of Ligands by Exponential Enrichment [19]) andin vivo(e. g. CLIP experiments: UV Cross-Linking and Immuno-Precipitation [20]). In RNAcompete experiments, purified epitope-tagged RBPs select RNA sequences from a designed (non-randomized) RNA pool. Bound RNAs are identified using microarray hybridizations and analyzed computationally to determine RBP-specific 7-mer RNA-binding profiles. Severalin vitromethods have been H-1152 described since the inception of RNAcompete, including RNA Bind-n-Seq (RNBS) [21], SEQRS [22], RNA-MaP [23], HiTS-RAP [24], and RNA-MITOMI [25]. Of these, only RNA Bind-n-seq has been used in a large-scale study (unpublished ENCODE online data). The RNAcompete methodology has several attractive features including: i) no antibodies are required; ii) no iterative selection or library preparation is necessary; iii) does not require RBP-specific optimizations; iv) is relatively inexpensive H-1152 at scale; v) is amenable to large-scale studies [3]; and, vi) has an established and validated uniform computational analysis pipeline. == Figure 1 . == Schematic of the RNAcompete assay. A GST-tagged RBP (RBP is orange oval, GST-tag is blue crescent), is incubated with a 75-fold excess of a non-random, custom designed RNA pool (multicoloured lines). RNA selectively bound (purple line) to an RBP during a GST-pulldown assay (GST bead is represented as a beige oval) is eluted, directly labeled with either Cy3 or Cy5 (green circles), and hybridized to a custom Agilent 244K microarray. Microarray data is analyzed computationally to generate RNA-binding motifs represented as logos. In this report, we present a detailed protocol for the experimental and computational components of the RNAcompete system. We also highlight the utility of RNAcompete H-1152 via analysis of.