Q.Find the average displacement of the car given the number of cylinders. (Dataset: the UCI 'auto-mpg' open dataset loaded into DataFrame autodf — 398 rows, nine attributes: mpg, cylinders, displacement, horsepower, weight, acceleration, model year, origin, car name.)
You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
Start your 14-day free trial to unlock the full solution →To find the average displacement grouped by the number of cylinders, use the Pandas groupby() method on the cylinders column, select the displacement column, and then apply the mean() aggregation function.
When you need to calculate a summary statistic (like average, sum, count, min, max) for different categories within your data, the groupby() operation is the most efficient and idiomatic way to do it in Pandas. The question asks for the "average displacement... given the number of cylinders," which directly translates to grouping the data by the cylinders column and then computing the mean of the displacement column for each group.
This process follows the "split-apply-combine" paradigm:
- Split: The DataFrame
autodfis logically split into smaller DataFrames, one for each unique value in thecylinderscolumn. - Apply: For each of these smaller DataFrames, the
mean()function is applied specifically to thedisplacementcolumn. - Combine: The results from each group's
mean()calculation are combined into a single Pandas Series, where the index represents the uniquecylindersvalues and the values are their corresponding average displacements.
Here's how you can achieve this:
import pandas as pd
# Assume autodf is already loaded as described in the problem statement.
# For demonstration purposes, a sample autodf is created below,
# mimicking the structure of the UCI 'auto-mpg' dataset.
# In a real scenario, you would use the pre-loaded autodf directly.
data = {
'mpg': [18.0, 15.0, 18.0, 16.0, 17.0, 21.0, 21.0, 22.0, 20.0, 19.0, 28.0, 31.0],
'cylinders': [8, 8, 8, 8, 8, 6, 6, 4, 4, 6, 4, 4],
'displacement': [307.0, 350.0, 318.0, 304.0, 302.0, 199.0, 220.0, 140.0, 140.0, 150.0, 112.0, 119.0],
'horsepower': [130, 165, 150, 150, 140, 97, 95, 90, 95, 88, 85, 92],
'weight': [3504, 3693, 3436, 3449, 3449, 2774, 2833, 2648, 2648, 2500, 2372, 2434],
'acceleration': [12.0, 11.5, 11.0, 12.0, 10.5, 15.5, 15.5, 16.0, 16.0, 17.0, 15.0, 15.0],
'model year': [70, 70, 70, 70, 70, 70, 70, 70, 70, 70, 79, 79],
'origin': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1],
'car name': [
'chevrolet chevelle malibu', 'buick skylark 320', 'plymouth satellite',
'amc rebel sst', 'ford torino', 'ford galaxie 500', 'chevrolet impala',
'plymouth fury iii', 'pontiac catalina', 'amc ambassador dpl',
'honda civic 1500 gl', 'toyota corolla'
]
}
autodf = pd.DataFrame(data)
# Calculate the average displacement for each number of cylinders
average_displacement_by_cylinders = autodf.groupby('cylinders')['displacement'].mean()
# Display the result
print(average_displacement_by_cylinders)
Explanation of Key Lines:
autodf.groupby('cylinders'):- This is the core of the operation. It groups the
autodfDataFrame by the unique values found in thecylinderscolumn. …
- This is the core of the operation. It groups the
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.