Bioinformatics with Python Cookbook

Bioinformatics with Python Cookbook - Second Edition

By : Tiago Antao

Buy this Book

Bioinformatics with Python Cookbook - Second Edition

By: Tiago Antao

Buy this Book

Overview of this book

Bioinformatics is an active research field that uses a range of simple-to-advanced computations to extract valuable information from biological data. This book covers next-generation sequencing, genomics, metagenomics, population genetics, phylogenetics, and proteomics. You'll learn modern programming techniques to analyze large amounts of biological data. With the help of real-world examples, you'll convert, analyze, and visualize datasets using various Python tools and libraries. This book will help you get a better understanding of working with a Galaxy server, which is the most widely used bioinformatics web-based pipeline system. This updated edition also includes advanced next-generation sequencing filtering techniques. You'll also explore topics such as SNP discovery using statistical approaches under high-performance computing frameworks such as Dask and Spark. By the end of this book, you'll be able to use and implement modern programming techniques and frameworks to deal with the ever-increasing deluge of bioinformatics data.

Title Page

About Packt

Contributors

Preface

Free Chapter

Python and the Surrounding Software Ecology

Introduction

Installing the required software with Anaconda

Installing the required software with Docker

Interfacing with R via rpy2

Performing R magic with Jupyter Notebook

Next-Generation Sequencing

Introduction

Accessing GenBank and moving around NCBI databases

Performing basic sequence analysis

Working with modern sequence formats

Working with alignment data

Analyzing data in VCF

Studying genome accessibility and filtering SNP data

Processing NGS data with HTSeq

Working with Genomes

Introduction

Working with high-quality reference genomes

Dealing with low-quality genome references

Traversing genome annotations

Extracting genes from a reference using annotations

Finding orthologues with the Ensembl REST API

Retrieving gene ontology information from Ensembl

Population Genetics

Introduction

Managing datasets with PLINK

Introducing the Genepop format

Exploring a dataset with Bio.PopGen

Computing F-statistics

Performing Principal Components Analysis

Investigating population structure with admixture

Population Genetics Simulation

Introduction

Introducing forward-time simulations

Simulating selection

Simulating population structure using island and stepping-stone models

Modeling complex demographic scenarios

Phylogenetics

Introduction

Preparing a dataset for phylogenetic analysis

Aligning genetic and genomic data

Comparing sequences

Reconstructing phylogenetic trees

Playing recursively with trees

Visualizing phylogenetic data

Using the Protein Data Bank

Introduction

Finding a protein in multiple databases

Introducing Bio.PDB

Extracting more information from a PDB file

Computing molecular distances on a PDB file

Performing geometric operations

Animating with PyMOL

Parsing mmCIF files using Biopython

Bioinformatics Pipelines

Introduction

Introducing Galaxy servers

Accessing Galaxy using the API

Developing a Galaxy tool

Using generic pipelines with bioinformatics data

Deploying a variant analysis pipeline with Airflow

Python for Big Genomics Datasets

Introduction

Using high-performance data formats – HDF5

Doing parallel computing with Dask

Using high-performance data formats – Parquet

Computing sequencing statistics using Spark

Optimizing code with Cython and Numba

Visualizing phylogenetic data

In this recipe, we will discuss how to visualize phylogenetic trees. DendroPy has only simple visualization mechanisms based on drawing textual ASCII trees, but Biopython has quite a rich infrastructure, which we will leverage here.

Getting ready

This will require you to have completed all of the previous recipes. Remember that we have the files for the whole genus Ebola virus, including the RAxML tree. Furthermore, a simplified genus version will have been produced in the previous recipe. As usual, you can find this content in the Chapter06/Visualization.ipynb Notebook file.

How to do it...

Take a look at the following steps:

Let's load all the phylogenetic data:

from copy import deepcopy
from Bio import Phylo
ebola_tree = Phylo.read('my_ebola.nex', 'nexus')
ebola_tree.name = 'Ebolavirus tree'
ebola_simple_tree = Phylo.read('ebola_simple.nex', 'nexus')
ebola_simple_tree.name = 'Ebolavirus simplified tree'

For all of the trees that we read, we will change the name...

Bioinformatics with Python Cookbook - Second Edition

By : Tiago Antao

Bioinformatics with Python Cookbook - Second Edition

By: Tiago Antao

Overview of this book

Related Content you might be interested in

Current Title:

Bioinformatics with Python Cookbook - Second Edition

Visualizing phylogenetic data

Getting ready

How to do it...