R Bioinformatics Cookbook - Second Edition

By : Dan MacLean

R Bioinformatics Cookbook - Second Edition

By: Dan MacLean

Overview of this book

The updated second edition of R Bioinformatics Cookbook takes a recipe-based approach to show you how to conduct practical research and analysis in computational biology with R. You’ll learn how to create a useful and modular R working environment, along with loading, cleaning, and analyzing data using the most up-to-date Bioconductor, ggplot2, and tidyverse tools. This book will walk you through the Bioconductor tools necessary for you to understand and carry out protocols in RNA-seq and ChIP-seq, phylogenetics, genomics, gene search, gene annotation, statistical analysis, and sequence analysis. As you advance, you'll find out how to use Quarto to create data-rich reports, presentations, and websites, as well as get a clear understanding of how machine learning techniques can be applied in the bioinformatics domain. The concluding chapters will help you develop proficiency in key skills, such as gene annotation analysis and functional programming in purrr and base R. Finally, you'll discover how to use the latest AI tools, including ChatGPT, to generate, edit, and understand R code and draft workflows for complex analyses. By the end of this book, you'll have gained a solid understanding of the skills and techniques needed to become a bioinformatics specialist and efficiently work with large and complex bioinformatics datasets.

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Conventions used

Get in touch

Share Your Thoughts

Download a free PDF copy of this book

Chapter 1: Setting Up Your R Bioinformatics Working Environment

Technical requirements

Setting up an R project in a directory

Using the here package to simplify working with paths

Using the devtools package to work with the latest non-CRAN packages

Setting up your machine for the compilation of source packages

Using the renv package to create a project-specific set of packages

Installing and managing different versions of Bioconductor packages in environments

Using bioconda to install external tools

Free Chapter

Chapter 2: Loading, Tidying, and Cleaning Data in the tidyverse

Technical requirements

Loading data from files with readr

Tidying a wide format table into a tidy table with tidyr

Tidying a long format table into a tidy table with tidyr

Combining tables using join functions

Reformatting and extracting existing data into new columns using stringr

Computing new data columns from existing ones and applying arbitrary functions using mutate()

Using dplyr to summarize data in large tables

Using datapasta to create R objects from cut-and-paste data

Chapter 3: ggplot2 and Extensions for Publication Quality Plots

Technical requirements

Combining many plot types in ggplot2

Comparing changes in distributions with ggridges

Customizing plots with ggeasy

Highlighting selected values in busy plots with gghighlight

Plotting variability and confidence intervals better with ggdist

Making interactive plots with plotly

Clarifying label placement with ggrepel

Zooming and making callouts from selected plot sections with facetzoom

Chapter 4: Using Quarto to Make Data-Rich Reports, Presentations, and Websites

Technical requirements

Using Markdown and Quarto for literate computation

Creating different document formats from the same source

Creating data-rich presentations from code

Creating websites from collections of Quarto documents

Adding interactivity with Shiny

Chapter 5: Easily Performing Statistical Tests Using Linear Models

Technical requirements

Modeling data with a linear model

Using a linear model to compare the mean of two groups

Using a linear model and ANOVA to compare multiple groups in a single variable

Using linear models and ANOVA to compare multiple groups in multiple variables

Testing and accounting for interactions between variables in linear models

Doing tests for differences in data in two categorical variables

Making predictions using linear models

Chapter 6: Performing Quantitative RNA-seq

Technical requirements

Estimating differential expression with edgeR

Estimating differential expression with DESeq2

Estimating differential expression with Kallisto and Sleuth

Using Sleuth to analyze time course experiments

Analyzing splice variants with SGSeq

Performing power analysis with powsimR

Finding unannotated transcribed regions

Finding regions showing high expression ab initio using bumphunter

Differential peak analysis

Estimating batch effects with SVA

Finding allele-specific expression with AllelicImbalance

Presenting RNA-Seq data using ComplexHeatmap

Chapter 7: Finding Genetic Variants with HTS Data

Technical requirements

Finding SNPs and INDELs from sequence data using VariantTools

Getting ready

Predicting open reading frames in long reference sequences

Plotting features on genetic maps with karyoploteR

Selecting and classifying variants with VariantAnnotation

Extracting information in genomic regions of interest

Finding phenotype and genotype associations with GWAS

Estimating the copy number at a locus of interest

Chapter 8: Searching Gene and Protein Sequences for Domains and Motifs

Technical requirements

Finding DNA motifs with universalmotif

Finding protein domains with PFAM and bio3d

Finding InterPro domains

Finding transmembrane domains with tmhmm and pureseqTM

Creating figures of protein domains using drawProteins

Performing multiple alignments of proteins or genes

Aligning genomic length sequences with DECIPHER

Novel feature detection in proteins

3D structure protein alignment in bio3d

Chapter 9: Phylogenetic Analysis and Visualization

Technical requirements

Reading and writing varied tree formats with ape and treeio

Visualizing trees of many genes quickly with ggtree

Quantifying and estimating the differences between trees with treespace

Extracting and working with subtrees using ape

Creating dot plots for alignment visualizations

Reconstructing trees from alignments using phangorn

Finding orthologue candidates using reciprocal BLASTs

Chapter 10: Analyzing Gene Annotations

Technical requirements

Retrieving gene and genome annotations from BioMart

Getting Gene Ontology information for functional analysis from appropriate databases

Using AnnoDB packages for genome annotation

Using ClusterProfiler for determining GO enrichment in clusters

Finding GO enrichment in an Ontology Conditional way with topGO

Finding enriched KEGG pathways

Retrieving and working with SNPs

Chapter 11: Machine Learning with mlr3

Technical requirements

Defining a task and learner to implement k-nearest neighbors (k-NNs) in mlr3

Testing the fit of the model using cross-validation

Using logistic regression to classify the relative likelihood of two outcomes

Classifying using random forest and interpreting it with iml

Dimension reduction with PCA in mlr3 pipelines

Creating a tSNE and UMAP embedding

Clustering with k-means and hierarchical clustering

Chapter 12: Functional Programming with purrr and base R

Technical requirements

Making base R objects “tidy”

Using nested dataframes for functional programming

Using the apply family of functions

Using the map family of functions in purrr

Working with lists in purrr

Chapter 13: Turbo-Charging Development in R with ChatGPT

Technical requirements

Interpreting complicated code with ChatGPT assistance

Debugging and improving code with ChatGPT

Generating code with ChatGPT

Getting ready

Writing documentation for R functions with ChatGPT

Writing unit tests for R functions with ChatGPT

Finding R packages to build a workflow with ChatGPT

Index

Why subscribe?

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Download a free PDF copy of this book

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Loading data from files with readr

The readr R package is a package that provides functions for reading and writing tabular data in a variety of formats, including comma-separated values (CSV), tab-separated values (TSV), and delimiter-separated files. It is designed to be flexible and stop helpfully when data changes or unexpected items appear in the input. The two main advantages over base R functions include consistency in interface and output and the ability to be explicit about types and inspect those types.

This latter advantage can help to avoid errors when reading data, as well as make data cleaning and manipulation easier. readr functions can also automatically infer the data types of each column, which can be useful for a preliminary inspection of large datasets or when the data types are not known.

Consistency in interface and output is one of the main advantages of readr functions. readr functions provide a consistent interface for reading different types of data, which can make it easier to work with multiple types of files. For example, the read_csv() function can be used to read CSV files, while the read_tsv() function can be used to read TSV files. Additionally, readr functions return a tibble, a modern version of a data frame that is more consistent in its output and easier to read than the base R data frame.

Getting ready

For this recipe, we’ll need the readr and rbioinfcookbook packages. The latter contains a census_2021.csv file that carries UK census data from 2021, from the UK Office for National Statistics (https://www.ons.gov.uk/). You will need to inspect it, especially its header, to understand the process in this recipe. The first step shows you how to find where the file is on your filesystem.

Note the delimiters in the file are commas (,) and that the first seven lines contain metadata that isn’t part of the main data. Also, look at the messy column headings and note that the numbers themselves are internally delimited by commas.

How to do it…

We begin by getting a filename for the sample in the package:

Load the package and get a filename:

library(readr)filename <- fs::path_package("extdata",                              "census_2021.csv",                              package="rbioinfcookbook"                             )

Specify a vector of new names for columns:

col_names = c(            c("area_code", "country", "region",               "area_name", "all_persons", "under_4"),            paste0(seq(5, 85, by = 5),"_to_",seq(9, 89, by =5)),c("over_90"))

Set the column types based on new names and contents:

col_types = cols(  area_code = col_character(),  country = col_factor(levels = c("England", "Wales")),  region = col_factor(levels = c("North East",                                 "Yorkshire Humber",                                 "East Midlands",                                 "West Midlands",                                 "East of England",                                          "London",                                 "South East",                                 "South West",                                 "Wales"), ordered = TRUE                      ),  area_name = col_character(),  .default = col_number())

Put it together and read the file:

df <- read_csv(filename,               skip = 8,               col_names = col_names,               col_types = col_types )

And with that, we’ve loaded in a file with careful checking of the data types.

How it works…

Step 1 loads the library(readr) package. This package contains functions for reading and writing tabular data in a variety of formats, including CSV. The fs::path_package("extdata", "census_2021.csv", package="rbioinfcookbook") function is used to create a file path to the census_2021.csv file. It simply finds the place where the rbioinfcookbook package was installed and looks inside the extdata directory for the file, then it returns the full file path that leads to the file. Quite often, we would see the system.file() function used for this purpose. system.file() is a fine choice when everything works, but when it can’t find the file, it returns a blank string, which can be hard to debug. fs::path_package() is nicer to work with and will return an error when it can’t find the file.

In step 2, a vector of new column names is specified by the code. The vector contains several strings for the first few column names, and then a sequence of strings is created and a long list of age-related columns is created by concatenating two sequences of numbers. The resulting vector is stored in the col_names variable.

In step 3, we specify the R type we want each column to be. The categorical columns are set explicitly to factors, with the region being ordered explicitly in a rough geographical northeast to southwest way. The area_name column contains over 300 names, so we won’t make them an explicit factor and stick with it as a general text containing character type. The rest of the columns contain numeric data, so we make that the default with .default.

Finally, the read_csv() function is used to read the file specified in step 1 and create a data frame. The skip argument is used to skip the first eight rows, which include the metadata in the file and the messy header, the col_names argument is used to specify the new column names stored in col_names, and the col_types argument is used to specify the column types stored in col_types.

There’s more…

We used the read_csv() function for comma-separated data, but many more functions are available for different delimiters:

Function	Delimiter
`read_csv()`	CSV
`read_tsv()`	TSV
`read_delim()`	User-specified delimited files
`read_fwf()`	Fixed-width files
`read_table()`	Whitespace-separated files
`read_log()`	Web log files

Table 2.1 – Parser functions and the type of input file delimiter they work on in readr

For different local conventions on—for example—decimal separators and grouping marks, you can use the locale functions.

R Bioinformatics Cookbook - Second Edition

By : Dan MacLean

R Bioinformatics Cookbook - Second Edition

By: Dan MacLean

Overview of this book

Related Content you might be interested in

Current Title:

R Bioinformatics Cookbook - Second Edition

Bioinformatics with Python Cookbook

Deep Learning for Genomics

R Programming Fundamentals

Loading data from files with readr

Getting ready

How to do it…

How it works…

There’s more…

See also