Book Image

Geospatial Data Analytics on AWS

By : Scott Bateman, Janahan Gnanachandran, Jeff DeMuth

Book Image

Geospatial Data Analytics on AWS

By: Scott Bateman, Janahan Gnanachandran, Jeff DeMuth

Overview of this book

Managing geospatial data and building location-based applications in the cloud can be a daunting task. This comprehensive guide helps you overcome this challenge by presenting the concept of working with geospatial data in the cloud in an easy-to-understand way, along with teaching you how to design and build data lake architecture in AWS for geospatial data. You’ll begin by exploring the use of AWS databases like Redshift and Aurora PostgreSQL for storing and analyzing geospatial data. Next, you’ll leverage services such as DynamoDB and Athena, which offer powerful built-in geospatial functions for indexing and querying geospatial data. The book is filled with practical examples to illustrate the benefits of managing geospatial data in the cloud. As you advance, you’ll discover how to analyze and visualize data using Python and R, and utilize QuickSight to share derived insights. The concluding chapters explore the integration of commonly used platforms like Open Data on AWS, OpenStreetMap, and ArcGIS with AWS to enable you to optimize efficiency and provide a supportive community for continuous learning. By the end of this book, you’ll have the necessary tools and expertise to build and manage your own geospatial data lake on AWS, along with the knowledge needed to tackle geospatial data management challenges and make the most of AWS services.

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Conventions used

Share Your Thoughts

Download a free PDF copy of this book

Part 1: Introduction to the Geospatial Data Ecosystem

Part 1: Introduction to the Geospatial Data Ecosystem

Free Chapter

Chapter 1: Introduction to Geospatial Data in the Cloud

Chapter 1: Introduction to Geospatial Data in the Cloud

Introduction to cloud computing and AWS

Storing geospatial data in the cloud

Building your geospatial data strategy

Geospatial data management best practices

Cost management in the cloud

Chapter 2: Quality and Temporal Geospatial Data Concepts

Chapter 2: Quality and Temporal Geospatial Data Concepts

Quality impact on geospatial data

Transmission methods

Understanding file formats

Normalizing data

Considering temporal dimensions

Part 2: Geospatial Data Lakes using Modern Data Architecture

Part 2: Geospatial Data Lakes using Modern Data Architecture

Chapter 3: Geospatial Data Lake Architecture

Chapter 3: Geospatial Data Lake Architecture

Modern data architecture overview

The AWS modern data architecture pillars

Geospatial Data Lake

Designing a geospatial data lake using modern data architecture

Chapter 4: Using Geospatial Data with Amazon Redshift

Chapter 4: Using Geospatial Data with Amazon Redshift

What is Redshift?

Understanding Redshift partitioning

Redshift Spectrum

Redshift geohashing support

Redshift geospatial support

Launching a Redshift cluster and running a geospatial query

Chapter 5: Using Geospatial Data with Amazon Aurora PostgreSQL

Chapter 5: Using Geospatial Data with Amazon Aurora PostgreSQL

Lab prerequisites

Setting up the database

Connecting to the database

Geospatial data loading

Queries and transformations

Architectural considerations

Chapter 6: Serverless Options for Geospatial

Chapter 6: Serverless Options for Geospatial

What is serverless?

Geospatial applications and S3 web hosting

Python with Lambda and API Gateway

Deploying your first serverless geospatial application

Chapter 7: Querying Geospatial Data with Amazon Athena

Chapter 7: Querying Geospatial Data with Amazon Athena

Setting up and configuring Athena

Geospatial data formats

Spatial query structure

Spatial functions

AWS service integration

Architectural considerations

Part 3: Analyzing and Visualizing Geospatial Data in AWS

Part 3: Analyzing and Visualizing Geospatial Data in AWS

Chapter 8: Geospatial Containers on AWS

Chapter 8: Geospatial Containers on AWS

Understanding containers

Deploying containers

Chapter 9: Using Geospatial Data with Amazon EMR

Chapter 9: Using Geospatial Data with Amazon EMR

Introducing Hadoop

Common Hadoop frameworks

Geospatial with EMR

Chapter 10: Geospatial Data Analysis Using R on AWS

Chapter 10: Geospatial Data Analysis Using R on AWS

Introduction to the R geospatial data analysis ecosystem

Setting up R and RStudio on EC2

RStudio on Amazon SageMaker

Analyzing and visualizing geospatial data using RStudio

Chapter 11: Geospatial Machine Learning with SageMaker

Chapter 11: Geospatial Machine Learning with SageMaker

AWS ML background

Common libraries and algorithms

Introducing Geospatial ML with SageMaker

Deploying a SageMaker Geospatial example

Architectural considerations

Chapter 12: Using Amazon QuickSight to Visualize Geospatial Data

Chapter 12: Using Amazon QuickSight to Visualize Geospatial Data

Geospatial visualization background

Amazon QuickSight overview

Connecting to your data source

Visualization layout

Putting it all together

Reports and collaboration

Part 4: Accessing Open Source and Commercial Platforms and Services

Part 4: Accessing Open Source and Commercial Platforms and Services

Chapter 13: Open Data on AWS

Chapter 13: Open Data on AWS

What is open data?

The Registry of Open Data on AWS

Analyzing open data

Federated queries with Athena

Open Data on AWS benefits

Chapter 14: Leveraging OpenStreetMap on AWS

Chapter 14: Leveraging OpenStreetMap on AWS

What is OpenStreetMap?

Accessing OSM from AWS

Application – ski lift scout

The OSM community

Architectural considerations

Chapter 15: Feature Servers and Map Servers on AWS

Chapter 15: Feature Servers and Map Servers on AWS

Types of servers and deployment options

Capabilities and cloud integrations

Deploying a container on AWS with ECR and EC2

Further reading

Chapter 16: Satellite and Aerial Imagery on AWS

Chapter 16: Satellite and Aerial Imagery on AWS

Imagery options

Architectural considerations

Demonstrating satellite imagery using AWS

Index

Other Books You May Enjoy

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Download a free PDF copy of this book

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

Storing geospatial data in the cloud

As you learn about the possibilities for storing geospatial data in the cloud, it may seem daunting due to the number of options available. Many AWS customers experiment with Amazon Simple Storage Service (S3) for geospatial data storage as their first project. Relational databases, NoSQL databases, and caching options commonly follow in the evolution of geospatial technical architectures. General GIS data storage best practices still apply to the cloud, so much of the knowledge that practitioners have gained over the years directly applies to geospatial data management on AWS. Familiar GIS file formats that work well in S3 include the following:

Shapefiles (.shp, .shx, .dbf, .prj, and others)
File geodatabases (.gdb)
Keyhole Markup Language (.kml)
Comma-Separated Values (.csv)
Geospatial JavaScript Object Notation (.geojson)
Geostationary Earth Orbit Tagged Image File Format (.tiff)

The physical location of data is still important for latency-sensitive workloads. Formats and organization of data can usually remain unchanged when moving to S3 to limit the impact of migrations. Spatial indexes and use-based access patterns will dramatically improve the performance and ability of your system to deliver the desired capabilities to your users.

Relational databases have long been the cornerstone of most enterprise GIS environments. This is especially true for vector datasets. AWS offers the most comprehensive set of relational database options with flexible sizing and architecture to meet your specific requirements. For customers looking to migrate geodatabases to the cloud with the least amount of environmental change, Amazon Elastic Compute Cloud (EC2) virtual machine instances provide a similar capability to what is commonly used in on-premises data centers. Each database server can be instantiated on the specific operating system that is used by the source server. Using EC2 with Amazon Elastic Block Store (EBS) network-attached storage provides the highest level of control and flexibility. Each server is created by specifying the amount of CPU, memory, and network throughput desired. Relational database management system (RDBMS) software can be manually installed on the EC2 instance, or an Amazon Machine Image (AMI) for the particular use case can be selected from the AWS catalog to remove manual steps from the process. While this option provides the highest degree of flexibility, it also requires the most database configuration and administration knowledge.

Many customers find it useful to leverage Amazon Relational Database Service (RDS) to establish database clusters and instances for their GIS environments. RDS can be leveraged by creating full-featured database Microsoft SQL Server, Oracle, PostgreSQL, MySQL, or MariaDB clusters. AWS allows the selection of specific instance types to focus on memory or compute optimization in a variety of configurations. Multiple Availability Zone (AZ)-enabled databases can be created to establish fault tolerance or improve performance. Using RDS dramatically simplifies database administration, and decreases the time required to select, provision, and configure your geospatial database using the specific technical parameters to meet the business requirements.

Amazon Aurora provides an open source path to highly capable and performant relational databases. PostgreSQL or MySQL environments can be created with specific settings for the desired capabilities. Although this may mean converting data from a source format, such as Microsoft SQL Server or Oracle, the overall cost savings and simplified management make this an attractive option to modernize and right-size any geospatial database.

In addition to standard relational database options, AWS provides other services to manage and use geospatial data. Amazon Redshift is the fastest and most widely used cloud data warehouse and supports geospatial data through the geometry data type. Users can query spatial data in Redshift’s built-in SQL functions to find the distance between two points, interrogate polygon relationships, and provide other location insights into their data. Amazon DynamoDB is a fully managed, key-value NoSQL database with an SLA of up to 99.999% availability. For organizations leveraging MongoDB, Amazon DocumentDB provides a fully managed option for simplified instantiation and management. Finally, AWS offers the Amazon OpenSearch Service for petabyte-scale data storage, search, and visualization.

The best part is that you don’t have to choose a single option for your geospatial environment. Often, companies find that different workloads benefit from having the ability to choose the most appropriate data landscape. Combining Infrastructure as a Service (IaaS) workloads with fully managed databases and modern databases is not only possible but a signature of a well-architected geospatial environment. Transactional systems may benefit from relational geodatabases, while mobile applications may be more aligned with NoSQL data stores. When you operate in a world of consumption-based resources, there is no downside to using the most appropriate data store for each workload. Having familiarity with the cloud options for storing geospatial data is crucial in strategic planning, which we will cover in the next topic.