Skip to content
Case Study · Q2

Q.Give description of the generated DataFrame autodf. (Dataset: the UCI 'auto-mpg' open dataset — 398 rows, nine attributes: mpg, cylinders, displacement, horsepower, weight, acceleration, model year, origin, car name.)

Odisha ChseTextbookSubjective· 2mImportance★★★★★est
70% · 37/53 Questions
✓ Free question

The autodf DataFrame is a tabular data structure in Pandas representing the UCI 'auto-mpg' dataset: 398 rows (car records) and 9 columns (mpg, cylinders, displacement, horsepower, weight, acceleration, model year, origin, car name), with a mix of numeric and string dtypes, and exactly 6 missing horsepower values.

A Pandas DataFrame is a two-dimensional, size-mutable, potentially heterogeneous tabular data structure with labeled axes (rows and columns) — think of it as a spreadsheet or a SQL table, where each column can hold its own data type.

The 'auto-mpg' dataset is a classic dataset used for regression tasks, containing fuel-efficiency and engineering data for cars mostly from the 1970s and early 1980s. Loading it produces autodf, which organizes this information into a structured, queryable form.

Structure of autodf

  1. Shape: autodf has 398 rows and 9 columns — one row per car record, one column per attribute.

  2. Index: Pandas assigns a default RangeIndex, an integer index running from 0 to 397.

  3. Columns and their types:

    ColumnDescriptionTypical dtype
    mpgfuel efficiency, miles per gallonfloat64
    cylindersnumber of engine cylinders (categorical, e.g. 3/4/5/6/8)int64
    displacementengine displacementfloat64
    horsepowerengine horsepower — has missing valuesfloat64
    weightvehicle weight (lbs)int64/float64
    acceleration0–60 mph timefloat64
    model_yearmodel year (categorical, e.g. 70 for 1970)int64
    origincountry of origin code (categorical: 1 = USA, 2 = Europe, 3 = Japan)int64
    car_nameunique car name/model stringobject

Data types at a glance

  • float64: mpg, displacement, horsepower, acceleration
  • int64: cylinders, weight, model_year, origin
  • object (string): car_name

Missing values

The only attribute with missing data is horsepower, which has 6 missing values out of 398 rows (originally marked ? in the raw file, converted to NaN on load with na_values='?'). All other 8 columns are complete.

autodf.info()

Output:

<class 'pandas.core.frame.DataFrame'>
RangeIndex: 398 entries, 0 to 397
Data columns (total 9 columns):
 #   Column        Non-Null Count  Dtype  
---  ------        --------------  -----  
 0   mpg           398 non-null    float64
 1   cylinders     398 non-null    int64  
 2   displacement  398 non-null    float64
 3   horsepower    392 non-null    float64
 4   weight        398 non-null    int64  
 5   acceleration  398 non-null    float64
 6   model_year    398 non-null    int64  
 7   origin        398 non-null    int64  
 8   car_name      398 non-null    object 
dtypes: float64(4), int64(4), object(1)
memory usage: 28.1+ KB
autodf.head()

Output:

    mpg  cylinders  displacement  horsepower  weight  acceleration  model_year  origin                   car_name
0  18.0          8         307.0       130.0  3504.0          12.0          70       1  chevrolet chevelle malibu
1  15.0          8         350.0       165.0  3693.0          11.5          70       1          buick skylark 320
2  18.0          8         318.0       150.0  3436.0          11.0          70       1         plymouth satellite
3  16.0          8         304.0       150.0  3433.0          12.0          70       1              amc rebel sst
4  17.0          8         302.0       140.0  3449.0          10.5          70       1                ford torino
Important

horsepower's Non-Null Count (392) being less than the total 398 entries is exactly how .info() surfaces missing data — the gap of 6 rows is the NaNs left behind after na_values='?' converted the file's ? markers. This is the same figure used in Exercise 4 (finding and handling missing values): 398 total rows, 392 complete, 6 with a missing horsepower.

Why "description" means more than just printing the table

Describing a DataFrame is not the same as displaying it. A good description covers: its shape (how much data), its columns and their meaning (what each attribute represents), its dtypes (whether pandas can do arithmetic on a column or only compare it as text), and data quality (which columns have gaps). autodf.describe() adds a fifth dimension — summary statistics (mean, std, min, quartiles, max) for every numeric column — which is useful once you've confirmed the structure above is correct.

✓Final answer

autodf is a Pandas DataFrame with 398 rows (one per car) and 9 columns — mpg, cylinders, displacement, horsepower, weight, acceleration, model_year, origin, car_name — using a default 0–397 integer index. Eight columns are fully populated; horsepower has 6 missing values (392 non-null), loaded as NaN from the file's ? markers. Data types are a mix of float64 (mpg, displacement, horsepower, acceleration), int64 (cylinders, weight, model_year, origin), and object/string (car_name).

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.