Bioinformatics with Python Cookbook

Bioinformatics with Python Cookbook - Second Edition

By : Tiago Antao

Buy this Book

Bioinformatics with Python Cookbook - Second Edition

By: Tiago Antao

Buy this Book

Overview of this book

Bioinformatics is an active research field that uses a range of simple-to-advanced computations to extract valuable information from biological data. This book covers next-generation sequencing, genomics, metagenomics, population genetics, phylogenetics, and proteomics. You'll learn modern programming techniques to analyze large amounts of biological data. With the help of real-world examples, you'll convert, analyze, and visualize datasets using various Python tools and libraries. This book will help you get a better understanding of working with a Galaxy server, which is the most widely used bioinformatics web-based pipeline system. This updated edition also includes advanced next-generation sequencing filtering techniques. You'll also explore topics such as SNP discovery using statistical approaches under high-performance computing frameworks such as Dask and Spark. By the end of this book, you'll be able to use and implement modern programming techniques and frameworks to deal with the ever-increasing deluge of bioinformatics data.

Title Page

About Packt

Contributors

Preface

Free Chapter

Python and the Surrounding Software Ecology

Introduction

Installing the required software with Anaconda

Installing the required software with Docker

Interfacing with R via rpy2

Performing R magic with Jupyter Notebook

Next-Generation Sequencing

Introduction

Accessing GenBank and moving around NCBI databases

Performing basic sequence analysis

Working with modern sequence formats

Working with alignment data

Analyzing data in VCF

Studying genome accessibility and filtering SNP data

Processing NGS data with HTSeq

Working with Genomes

Introduction

Working with high-quality reference genomes

Dealing with low-quality genome references

Traversing genome annotations

Extracting genes from a reference using annotations

Finding orthologues with the Ensembl REST API

Retrieving gene ontology information from Ensembl

Population Genetics

Introduction

Managing datasets with PLINK

Introducing the Genepop format

Exploring a dataset with Bio.PopGen

Computing F-statistics

Performing Principal Components Analysis

Investigating population structure with admixture

Population Genetics Simulation

Introduction

Introducing forward-time simulations

Simulating selection

Simulating population structure using island and stepping-stone models

Modeling complex demographic scenarios

Phylogenetics

Introduction

Preparing a dataset for phylogenetic analysis

Aligning genetic and genomic data

Comparing sequences

Reconstructing phylogenetic trees

Playing recursively with trees

Visualizing phylogenetic data

Using the Protein Data Bank

Introduction

Finding a protein in multiple databases

Introducing Bio.PDB

Extracting more information from a PDB file

Computing molecular distances on a PDB file

Performing geometric operations

Animating with PyMOL

Parsing mmCIF files using Biopython

Bioinformatics Pipelines

Introduction

Introducing Galaxy servers

Accessing Galaxy using the API

Developing a Galaxy tool

Using generic pipelines with bioinformatics data

Deploying a variant analysis pipeline with Airflow

Python for Big Genomics Datasets

Introduction

Using high-performance data formats – HDF5

Doing parallel computing with Dask

Using high-performance data formats – Parquet

Computing sequencing statistics using Spark

Optimizing code with Cython and Numba

Parsing mmCIF files using Biopython

The mmCIF file format is probably the future. Biopython doesn't have full functionality to work with it yet, but we will take a look at what is here now.

Getting ready

As Bio.PDB is not able to automatically download mmCIF files, you need to get your protein file and rename it to 1tup.cif. This can be found at https://github.com/PacktPublishing/Bioinformatics-with-Python-Cookbook-Second-Edition/blob/master/Datasets.ipynb under the 1TUP.cif name.

You can find this content in the Chapter07/mmCIF.ipynb Notebook file.

How to do it...

Take a look at the following steps:

Let's parse the file. We just use the MMCIF parser instead of the PDB parser:

from Bio import PDB
parser = PDB.MMCIFParser()
p53_1tup = parser.get_structure('P53', '1tup.cif')

Let's inspect the following chains:

def describe_model(name, pdb):
    print()
    for model in p53_1tup:
        for chain in model:
            print('%s - Chain: %s. Number of residues: %d. Number of atoms: %d.' %
         ...

Bioinformatics with Python Cookbook - Second Edition

By : Tiago Antao

Bioinformatics with Python Cookbook - Second Edition

By: Tiago Antao

Overview of this book

Related Content you might be interested in

Current Title:

Bioinformatics with Python Cookbook - Second Edition

Parsing mmCIF files using Biopython

Getting ready

How to do it...