
Hands-On Big Data Analytics with PySpark
Analyze large datasets and discover techniques for testing, immunizing, and parallelizing Spark jobs
Publisher:Packt Publishing Limited
By: Rudy Lai and Bartłomiej Potaczek
Paid access
|Sep 2024Table of Contents
- Installing Pyspark and Setting up Your Development Environment
- Getting Your Big Data into the Spark Environment Using RDDs
- Big Data Cleaning and Wrangling with Spark Notebooks
- Aggregating and Summarizing Data into Useful Reports
- Powerful Exploratory Data Analysis with MLlib
- Putting Structure on Your Big Data with SparkSQL
- Transformations and Actions
- Immutable Design
- Avoiding Shuffle and Reducing Operational Expenses
- Saving Data in the Correct Format
- Working with the Spark Key/Value API
- Testing Apache Spark Jobs
- Leveraging the Spark GraphX API
PDF ISBN: 978-1-83864-883-1
Publisher: Packt Publishing Limited
Copyright owner: © 2019 Packt Publishing Limited
Publication date: 2024
Language: English
Pages: 182
Related subjects:
