Q.Display the first 10 rows of the DataFrame autodf. (Dataset: the UCI 'auto-mpg' open dataset — 398 rows, nine attributes: mpg, cylinders, displacement, horsepower, weight, acceleration, model year, origin, car name.)
autodf.head(10) returns the first 10 rows of the DataFrame — the standard way to sample the top of a dataset during initial exploration, before running any real analysis.
Why look at the first rows at all?
Before you compute a single statistic, you inspect the data: are the column names what you expect, do the values look like the right kind of number, is anything obviously broken (wrong units, text where you expected numbers, all-zero columns)? head(n) is the fastest way to do that — it prints the top n rows without scanning or summarizing the whole DataFrame, so it's cheap even on a dataset far bigger than 398 rows.
import pandas as pd
# Assuming autodf is already loaded
autodf.head(10)
Output:
| mpg | cylinders | displacement | horsepower | weight | acceleration | model_year | origin | car_name | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 18.0 | 8 | 307.0 | 130.0 | 3504 | 12.0 | 70 | 1 | chevrolet chevelle malibu |
| 1 | 15.0 | 8 | 350.0 | 165.0 | 3693 | 11.5 | 70 | 1 | buick skylark 320 |
| 2 | 18.0 | 8 | 318.0 | 150.0 | 3436 | 11.0 | 70 | 1 | plymouth satellite |
| 3 | 16.0 | 8 | 304.0 | 150.0 | 3433 | 12.0 | 70 | 1 | amc rebel sst |
| 4 | 17.0 | 8 | 302.0 | 140.0 | 3449 | 10.5 | 70 | 1 | ford torino |
| 5 | 15.0 | 8 | 429.0 | 198.0 | 4341 | 10.0 | 70 | 1 | ford galaxie 500 |
| 6 | 14.0 | 8 | 454.0 | 220.0 | 4354 | 9.0 | 70 | 1 | chevrolet impala |
| 7 | 14.0 | 8 | 440.0 | 215.0 | 4312 | 8.5 | 70 | 1 | plymouth fury iii |
| 8 | 14.0 | 8 | 455.0 | 225.0 | 4425 | 10.0 | 70 | 1 | pontiac catalina |
| 9 | 15.0 | 8 | 390.0 | 190.0 | 3850 | 8.5 | 70 | 1 | amc ambassador dpl |
Reading what the output tells you
Even before running any statistics, this preview already tells a story:
model_yearis 70 for all ten rows — the dataset is loaded in chronological order, so its earliest records are all from 1970 (stored as the two-digit code70, not1970).cylindersis 8 andoriginis 1 for every one of these rows — every early car is a large-engine, American-made (origin == 1) model. This is a real, visible pattern in the printed rows, not a coincidence of a random sample: if you only ever looked athead(), you would wrongly conclude the whole dataset is 8-cylinder American cars. Fuel-efficient 4-cylinder and non-US cars (origin2 = Europe, 3 = Japan) appear later, as model years progress — which is exactly whyhead()is a sanity check, not a substitute for looking at the full distribution (e.g.autodf['cylinders'].value_counts()orautodf['origin'].value_counts()).mpgin this early slice ranges from 14.0 to 18.0 — comparatively low, consistent with the large-displacement engines shown in the same rows. That relationship (bigger engine → lower mpg) is one of the first things this dataset is normally used to illustrate.
head() vs. related methods
| Method | What it returns |
|---|---|
autodf.head(10) | first 10 rows, in file/index order |
autodf.tail(10) | last 10 rows, in file/index order |
autodf.sample(10) | 10 randomly chosen rows (different every call unless you set random_state) |
autodf[:10] | equivalent to head(10) via positional slicing |
head()/tail() are deterministic and index-order-based, which is why they're the right tool for "does this look like it loaded correctly?" — sample() is the better tool once you want an unbiased look at the dataset as a whole, precisely because (as above) the first 10 rows alone are not representative of the full 398.
If you want to see the last rows instead, use autodf.tail(n). Both methods return a new DataFrame slice without modifying the original — they never mutate autodf.
autodf.head(10) displays the first 10 rows of the DataFrame — all nine columns (mpg, cylinders, displacement, horsepower, weight, acceleration, model_year, origin, car_name). In this dataset those first 10 records are all 1970-model, 8-cylinder, American-made (origin=1) cars with correspondingly low mpg — illustrating why head() is useful for a quick sanity check, but not a substitute for examining the dataset's full range with methods like .describe() or .value_counts().
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.