Python is a programming language heavily used by data analysts and data scientists these days. There are numerous scientific and statistical data processing libraries available, as well as charting and plotting libraries, that can be used in Python programs. Python is also widely used as a programming language to develop data processing applications in Spark. This brings in a great flexibility to have a unified data processing and data analysis framework with Spark, Python and Python libraries, enabling us to do scientific and statistical processing, and charting and plotting. There are numerous such libraries that work with Python. Out of all those, theNumPy and SciPy libraries are being used here to do numerical, statistical, and scientific data processing. The matplotlib library is being used here to do the charting and plotting that produces 2D images.
Apache Spark 2 for Beginners
By :
Apache Spark 2 for Beginners
By:
Overview of this book
<p>Spark is one of the most widely-used large-scale data processing engines and runs extremely fast. It is a framework that has tools that are equally useful for application developers as well as data scientists.</p>
<p>This book starts with the fundamentals of Spark 2 and covers the core data processing framework and API, installation, and application development setup. Then the Spark programming model is introduced through real-world examples followed by Spark SQL programming with DataFrames. An introduction to SparkR is covered next. Later, we cover the charting and plotting features of Python in conjunction with Spark data processing. After that, we take a look at Spark's stream processing, machine learning, and graph processing libraries. The last chapter combines all the skills you learned from the preceding chapters to develop a real-world Spark application.</p>
<p>By the end of this book, you will have all the knowledge you need to develop efficient large-scale applications using Apache Spark.</p>
Table of Contents (15 chapters)
Apache Spark 2 for Beginners
Credits
About the Author
About the Reviewer
www.PacktPub.com
Preface
Free Chapter
Spark Fundamentals
Spark Programming Model
Spark SQL
Spark Programming with R
Spark Data Analysis with Python
Spark Stream Processing
Spark Machine Learning
Spark Graph Processing
Designing Spark Applications
Customer Reviews