Q.Give description of the generated DataFrame autodf. (Dataset: the UCI 'auto-mpg' open dataset — 398 rows, nine attributes: mpg, cylinders, displacement, horsepower, weight, acceleration, model year, origin, car name.)
The autodf DataFrame is a tabular data structure in Pandas representing the UCI 'auto-mpg' dataset: 398 rows (car records) and 9 columns (mpg, cylinders, displacement, horsepower, weight, acceleration, model year, origin, car name), with a mix of numeric and string dtypes, and exactly 6 missing horsepower values.
A Pandas DataFrame is a two-dimensional, size-mutable, potentially heterogeneous tabular data structure with labeled axes (rows and columns) — think of it as a spreadsheet or a SQL table, where each column can hold its own data type.
The 'auto-mpg' dataset is a classic dataset used for regression tasks, containing fuel-efficiency and engineering data for cars mostly from the 1970s and early 1980s. Loading it produces autodf, which organizes this information into a structured, queryable form.
Structure of autodf
-
Shape:
autodfhas 398 rows and 9 columns — one row per car record, one column per attribute. -
Index: Pandas assigns a default
RangeIndex, an integer index running from 0 to 397. -
Columns and their types:
Column Description Typical dtype mpgfuel efficiency, miles per gallon float64cylindersnumber of engine cylinders (categorical, e.g. 3/4/5/6/8) int64displacementengine displacement float64horsepowerengine horsepower — has missing values float64weightvehicle weight (lbs) int64/float64acceleration0–60 mph time float64model_yearmodel year (categorical, e.g. 70 for 1970) int64origincountry of origin code (categorical: 1 = USA, 2 = Europe, 3 = Japan) int64car_nameunique car name/model string object
Data types at a glance
float64:mpg,displacement,horsepower,accelerationint64:cylinders,weight,model_year,originobject(string):car_name
Missing values
The only attribute with missing data is horsepower, which has 6 missing values out of 398 rows (originally marked ? in the raw file, converted to NaN on load with na_values='?'). All other 8 columns are complete.
autodf.info()
Output:
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 398 entries, 0 to 397
Data columns (total 9 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 mpg 398 non-null float64
1 cylinders 398 non-null int64
2 displacement 398 non-null float64
3 horsepower 392 non-null float64
4 weight 398 non-null int64
5 acceleration 398 non-null float64
6 model_year 398 non-null int64
7 origin 398 non-null int64
8 car_name 398 non-null object
dtypes: float64(4), int64(4), object(1)
memory usage: 28.1+ KB
autodf.head()
Output:
mpg cylinders displacement horsepower weight acceleration model_year origin car_name
0 18.0 8 307.0 130.0 3504.0 12.0 70 1 chevrolet chevelle malibu
1 15.0 8 350.0 165.0 3693.0 11.5 70 1 buick skylark 320
2 18.0 8 318.0 150.0 3436.0 11.0 70 1 plymouth satellite
3 16.0 8 304.0 150.0 3433.0 12.0 70 1 amc rebel sst
4 17.0 8 302.0 140.0 3449.0 10.5 70 1 ford torino
horsepower's Non-Null Count (392) being less than the total 398 entries is exactly how .info() surfaces missing data — the gap of 6 rows is the NaNs left behind after na_values='?' converted the file's ? markers. This is the same figure used in Exercise 4 (finding and handling missing values): 398 total rows, 392 complete, 6 with a missing horsepower.
Why "description" means more than just printing the table
Describing a DataFrame is not the same as displaying it. A good description covers: its shape (how much data), its columns and their meaning (what each attribute represents), its dtypes (whether pandas can do arithmetic on a column or only compare it as text), and data quality (which columns have gaps). autodf.describe() adds a fifth dimension — summary statistics (mean, std, min, quartiles, max) for every numeric column — which is useful once you've confirmed the structure above is correct.
autodf is a Pandas DataFrame with 398 rows (one per car) and 9 columns — mpg, cylinders, displacement, horsepower, weight, acceleration, model_year, origin, car_name — using a default 0–397 integer index. Eight columns are fully populated; horsepower has 6 missing values (392 non-null), loaded as NaN from the file's ? markers. Data types are a mix of float64 (mpg, displacement, horsepower, acceleration), int64 (cylinders, weight, model_year, origin), and object/string (car_name).
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.