Bioinformatics with Python Cookbook

Bioinformatics with Python Cookbook - Second Edition

By : Tiago Antao

Buy this Book

Bioinformatics with Python Cookbook - Second Edition

By: Tiago Antao

Buy this Book

Overview of this book

Bioinformatics is an active research field that uses a range of simple-to-advanced computations to extract valuable information from biological data. This book covers next-generation sequencing, genomics, metagenomics, population genetics, phylogenetics, and proteomics. You'll learn modern programming techniques to analyze large amounts of biological data. With the help of real-world examples, you'll convert, analyze, and visualize datasets using various Python tools and libraries. This book will help you get a better understanding of working with a Galaxy server, which is the most widely used bioinformatics web-based pipeline system. This updated edition also includes advanced next-generation sequencing filtering techniques. You'll also explore topics such as SNP discovery using statistical approaches under high-performance computing frameworks such as Dask and Spark. By the end of this book, you'll be able to use and implement modern programming techniques and frameworks to deal with the ever-increasing deluge of bioinformatics data.

Title Page

About Packt

Contributors

Preface

Free Chapter

Python and the Surrounding Software Ecology

Introduction

Installing the required software with Anaconda

Installing the required software with Docker

Interfacing with R via rpy2

Performing R magic with Jupyter Notebook

Next-Generation Sequencing

Introduction

Accessing GenBank and moving around NCBI databases

Performing basic sequence analysis

Working with modern sequence formats

Working with alignment data

Analyzing data in VCF

Studying genome accessibility and filtering SNP data

Processing NGS data with HTSeq

Working with Genomes

Introduction

Working with high-quality reference genomes

Dealing with low-quality genome references

Traversing genome annotations

Extracting genes from a reference using annotations

Finding orthologues with the Ensembl REST API

Retrieving gene ontology information from Ensembl

Population Genetics

Introduction

Managing datasets with PLINK

Introducing the Genepop format

Exploring a dataset with Bio.PopGen

Computing F-statistics

Performing Principal Components Analysis

Investigating population structure with admixture

Population Genetics Simulation

Introduction

Introducing forward-time simulations

Simulating selection

Simulating population structure using island and stepping-stone models

Modeling complex demographic scenarios

Phylogenetics

Introduction

Preparing a dataset for phylogenetic analysis

Aligning genetic and genomic data

Comparing sequences

Reconstructing phylogenetic trees

Playing recursively with trees

Visualizing phylogenetic data

Using the Protein Data Bank

Introduction

Finding a protein in multiple databases

Introducing Bio.PDB

Extracting more information from a PDB file

Computing molecular distances on a PDB file

Performing geometric operations

Animating with PyMOL

Parsing mmCIF files using Biopython

Bioinformatics Pipelines

Introduction

Introducing Galaxy servers

Accessing Galaxy using the API

Developing a Galaxy tool

Using generic pipelines with bioinformatics data

Deploying a variant analysis pipeline with Airflow

Python for Big Genomics Datasets

Introduction

Using high-performance data formats – HDF5

Doing parallel computing with Dask

Using high-performance data formats – Parquet

Computing sequencing statistics using Spark

Optimizing code with Cython and Numba

Working with modern sequence formats

Here, we will work with FASTQ files, the standard format output used by modern sequencers. You will learn how to work with quality scores per base and also consider the variations in output coming from different sequencing machines and databases. This is the first recipe that will use real data (big data) from the Human 1,000 Genomes Project. We will start with a brief description of the project.

Getting ready

The Human 1,000 Genomes Project aims to catalog worldwide human genetic variation and takes advantage of modern sequencing technology to do WGS. This project makes all data publicly available, which includes output from sequencers, sequence alignments, and SNP calls, among many other artifacts. The name "1,000 Genomes" is actually a misnomer, because it currently includes more than 2,500 samples. These samples are divided into 26 populations, spanning the whole planet. We will mostly use data from four populations: African Yorubans (YRI), Utah Residents...

Bioinformatics with Python Cookbook - Second Edition

By : Tiago Antao

Bioinformatics with Python Cookbook - Second Edition

By: Tiago Antao

Overview of this book

Related Content you might be interested in

Current Title:

Bioinformatics with Python Cookbook - Second Edition

Working with modern sequence formats

Getting ready