-
Book Overview & Buying
-
Table Of Contents
Exploratory Data Analysis with Python Cookbook
By :
Replacing values in rows or columns is a common practice when working with tabular data. There are many reasons why we may need to replace specific values within a dataset. Python provides the flexibility to replace single values or multiple values within our dataset. We can use the replace method to achieve this.
We will work with the Marketing Campaign data again for this recipe.
We will remove duplicate data using the pandas library:
pandas library:import pandas as pd
.csv file into a dataframe using read_csv. Then, subset the dataframe to include only relevant columns:marketing_data = pd.read_csv("data/marketing_campaign.csv")marketing_data = marketing_data[['ID', 'Year_Birth', 'Kidhome', 'Teenhome']]
ID Year_Birth Kidhome Teenhome
0 5524 1957 0 0
1 2174 1954 1 1
2 4141 1965 0 0
3 6182 1984 1 0
4 5324 1981 1 0
marketing_data.shape
(2240, 4)
Teenhome with has teen and has no teen:marketing_data['Teenhome_replaced'] = marketing_data['Teenhome'].replace([0,1,2],['has no teen','has teen','has teen'])
marketing_data[['Teenhome','Teenhome_replaced']].head()
Teenhome Teenhome_replaced
0 0 has no teen
1 1 has teen
2 0 has no teen
3 0 has no teen
4 0 has no teen
Great! We just replaced values in our dataset.
We refer to pandas as pd in step 1. In step 2, we use read_csv to load the .csv file into a pandas dataframe and call it marketing_data. We also subset the dataframe to include only four relevant columns. In step 3, we inspect the dataset using head() to see the first five rows in the dataset. Using the shape method, we get a sense of the number of rows and columns.
In step 4, we use the replace method to replace values within the Teenhome column. The first argument of the method is a list of the existing values that we want to replace, while the second argument contains a list of the values we want to replace it with. It is important to note that the lists for both arguments must be the same length.
In step 5, we inspect the result.
In some cases, we may need to replace a group of values that have complex patterns that cannot be explicitly stated. An example could be certain phone numbers or email addresses. In such cases, the replace method gives us the ability to use regex for pattern matching and replacement. Regex is short for regular expressions, and it is used for pattern matching.
pandas: https://datatofish.com/replace-values-pandas-dataframe/replace method in pandas: https://www.geeksforgeeks.org/replace-values-in-pandas-dataframe-using-regex/