From Transformations to Cleaned Data¶
This module focuses on the crucial step of transforming raw survey responses into clean, analytically meaningful variables. You will learn systematic approaches to data management that prepare your dataset for rigorous analysis.
Clean Data?
Well, if you ever see clean data... run away.
Theory¶
Where We Are in the Course¶
- Preparing for Milestone 4: Analysis
- Moving from basic exploration to more sophisticated variable transformation
- Data cleaning principles vs. data wrangling techniques
- Understanding when to filter vs. when to recode
- Preserving data while improving analytical quality
Key Concepts in Data Transformation¶
-
Recoding vs. Filtering
- Filtering: Removes observations from analysis
- Recoding: Transforms values while preserving cases
- Trade-offs: Data preservation vs. analytical clarity
-
Variable Creation Strategies
- Binary variables: From categorical responses
- Categorical bins: From continuous measurements
- Composite scales: Combining related items
Application¶
In-Class Exercise¶
Working together on data transformation challenges using ANES 2020 data. We'll explore systematic approaches to:
- Handle missing and invalid responses appropriately
- Create analytically useful variables from raw survey items
- Document transformation decisions for research transparency
Workshop Session¶
Interactive coding session where we apply transformation techniques to prepare data for hypothesis testing.
Code: Live Demo & Hands-on Lab¶
Student Group Live Demo (Group 1)¶
- Topic: Survey Data Recoding & Variable Creation (Handling
-9/-8/DK codes toNaN, masks vs..replace(), binary indicators). - Review the Live Demo Guidelines & Schedule.
Data Management & Variable Creation¶
Tip
To run Python scripts in VS Code, follow 📘 Running Python Scripts in VS Code.
What You'll Practice¶
Using ANES 2020 data, you will master:
- Recoding variables using both mask-based and dictionary methods
- Creating new variables through binary transformations
- Binning continuous data using
pd.cut()for categorical analysis - Comparing filtering vs. recoding approaches
- Preparing cleaned datasets for analysis
Key Techniques Covered¶
1. Mask-Based Recoding
# Create masks for specific values
mask = df['variable'] == target_value
df.loc[mask, 'variable'] = new_value
2. List-Based Replacement
# Using lists as a recoding tool
old_labels = [1, 2, 3]
new_labels = ["Low", "Medium", "High"]
df['variable'] = df['variable'].replace(old_labels, new_labels)
3. Variable Creation
# Binary variables
df['binary_var'] = (df['categorical_var'] == target_value).astype(int)
# Categorical bins
df['age_groups'] = pd.cut(df['age'], bins=[0, 35, 50, 65, 100],
labels=["Young", "Middle-aged", "Senior", "Elder"])
4. Additive Scale Creation
# Combine related survey items into an additive scale
scale_items = ['V242417', 'V242418', 'V242419'] # trust-related items
df[scale_items] = df[scale_items].replace([-9, -8, -7, -6, -5], pd.NA) # recode missing
df[scale_items] = 5 - df[scale_items] # Reverse values so higher = more trust
df['trust_scale'] = df[scale_items].sum(axis=1)/12 # create new scale variable
Additive scales combine multiple related survey items into a single measure by summing responses across rows.
Remember: The goal is not just to transform data, but to do so systematically and transparently. Every recoding decision should be defensible and documented.
Interactive Python Script¶
- Download and open
06_data_management_and_scales.pyfrom the course materials in VS Code.
Practice Problems¶
-
Vote Intention Transformation
- Clean the voting intention variable
- Create binary variables for major candidates
- Compare filtering vs. recoding approaches
-
Feeling Thermometer Analysis
- Recode COVID approval responses
- Create age categorical variables
- Generate cross-tabulations for analysis
-
Data Quality Assessment
- Compare filtered vs. recoded datasets
- Assess impact on analytical conclusions
- Document transformation decisions
Get Ready for Next Session: Think. Explore. Practice.¶
Think¶
- How do your key outcome variables vary across sociodemographic subgroups (e.g., age cohorts, education levels, party identification)?
- How will cross-tabulations and subgroup mean comparisons help substantiate the bivariate relationships required in Milestone 4?
- Suggested Reading: Chetty, R., Jackson, M. O., Kuchler, T., Stroebel, J., Hendren, N., Fluegge, R. B., ... & Wernerfelt, N. (2022). Social capital I: measurement and associations with economic mobility. Nature, 608(7921), 108-121. - Demonstrates how large-scale subgroup aggregation and cross-sectional comparisons uncover deep socioeconomic and political patterns.
Explore¶
- Live Demo (Group 2): Subgroup Analysis & Cross-Tabulations. Group 2 prepares a 10–15 min demonstration and shares the handout on WhatsApp before class (groups can book a meeting with the instructor to get direction).
- The class should review the Pandas GroupBy Guide to follow along and lead the peer discussion.
Practice¶
- Finalize your cleaned dataset and export your bivariate figures into your Typst manuscript for Milestone 4.
- Complete Milestone 4 - Analysis (Due Friday, Feb 05 at 08:00 before class via email).