Skip to content
Case Study · Q6

Q.Find the average displacement of the car given the number of cylinders. (Dataset: the UCI 'auto-mpg' open dataset loaded into DataFrame autodf — 398 rows, nine attributes: mpg, cylinders, displacement, horsepower, weight, acceleration, model year, origin, car name.)

CBSENCERTSubjective· 2mImportance★★★★★est
79% · 42/53 Questions
🔒 Locked · start free trial →

You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.

Start your 14-day free trial to unlock the full solution →

To find the average displacement grouped by the number of cylinders, use the Pandas groupby() method on the cylinders column, select the displacement column, and then apply the mean() aggregation function.

When you need to calculate a summary statistic (like average, sum, count, min, max) for different categories within your data, the groupby() operation is the most efficient and idiomatic way to do it in Pandas. The question asks for the "average displacement... given the number of cylinders," which directly translates to grouping the data by the cylinders column and then computing the mean of the displacement column for each group.

This process follows the "split-apply-combine" paradigm:

  1. Split: The DataFrame autodf is logically split into smaller DataFrames, one for each unique value in the cylinders column.
  2. Apply: For each of these smaller DataFrames, the mean() function is applied specifically to the displacement column.
  3. Combine: The results from each group's mean() calculation are combined into a single Pandas Series, where the index represents the unique cylinders values and the values are their corresponding average displacements.

Here's how you can achieve this:

import pandas as pd

# Assume autodf is already loaded as described in the problem statement.
# For demonstration purposes, a sample autodf is created below,
# mimicking the structure of the UCI 'auto-mpg' dataset.
# In a real scenario, you would use the pre-loaded autodf directly.
data = {
    'mpg': [18.0, 15.0, 18.0, 16.0, 17.0, 21.0, 21.0, 22.0, 20.0, 19.0, 28.0, 31.0],
    'cylinders': [8, 8, 8, 8, 8, 6, 6, 4, 4, 6, 4, 4],
    'displacement': [307.0, 350.0, 318.0, 304.0, 302.0, 199.0, 220.0, 140.0, 140.0, 150.0, 112.0, 119.0],
    'horsepower': [130, 165, 150, 150, 140, 97, 95, 90, 95, 88, 85, 92],
    'weight': [3504, 3693, 3436, 3449, 3449, 2774, 2833, 2648, 2648, 2500, 2372, 2434],
    'acceleration': [12.0, 11.5, 11.0, 12.0, 10.5, 15.5, 15.5, 16.0, 16.0, 17.0, 15.0, 15.0],
    'model year': [70, 70, 70, 70, 70, 70, 70, 70, 70, 70, 79, 79],
    'origin': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1],
    'car name': [
        'chevrolet chevelle malibu', 'buick skylark 320', 'plymouth satellite',
        'amc rebel sst', 'ford torino', 'ford galaxie 500', 'chevrolet impala',
        'plymouth fury iii', 'pontiac catalina', 'amc ambassador dpl',
        'honda civic 1500 gl', 'toyota corolla'
    ]
}
autodf = pd.DataFrame(data)

# Calculate the average displacement for each number of cylinders
average_displacement_by_cylinders = autodf.groupby('cylinders')['displacement'].mean()

# Display the result
print(average_displacement_by_cylinders)

Explanation of Key Lines:

  1. autodf.groupby('cylinders'):
    • This is the core of the operation. It groups the autodf DataFrame by the unique values found in the cylinders column. …

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.