Q.What would happen if we use df.plot(kind=’scatter’) instead of plt.scatter() in Program 4-13?
[Context: Program 4-13 is:
import numpy as np
import matplotlib.pyplot as plt
discount= np.array([10,20,30,40,50])
saleInRs=np.array([40000,45000,48000,50000,100000])
size=discount10
plt.scatter(x=discount,y=saleInRs,s=size,color='red',linewidth=3,marker='',edgecolor='blue')
plt.title('Sales Vs Discount')
plt.xlabel('Discount offered')
plt.ylabel('Sales in Rs')
plt.show()]
You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
Start your 14-day free trial to unlock the full solution →Using df.plot(kind='scatter') instead of plt.scatter() would fail because discount and saleInRs are NumPy arrays, not a pandas DataFrame; the syntax and capabilities also differ.
Why the two approaches are fundamentally different
The question asks what happens when we swap a Matplotlib function for a pandas plotting method. The answer hinges on understanding that these tools operate on different data structures and have different interfaces.
Program 4-13 works directly with NumPy arrays (discount and saleInRs). The plt.scatter() function from Matplotlib is designed to accept raw array-like data — lists, NumPy arrays, or sequences — and plot them directly. You pass x and y coordinates explicitly, along with optional parameters like marker size (s), color, marker shape, edge color, and line width.
The pandas method df.plot(kind='scatter'), on the other hand, is a DataFrame method. It expects to be called on a pandas DataFrame object, and you specify column names (as strings) for the x and y axes. The syntax would look like:
df.plot(kind='scatter', x='column_name_for_x', y='column_name_for_y')
What would actually happen
If you tried to write:
df.plot(kind='scatter')
in Program 4-13, Python would raise an error because there is no DataFrame object named df in the code. The variables discount and saleInRs are NumPy arrays, not DataFrames, and NumPy arrays do not have a .plot() method.
Even if you first converted the arrays into a DataFrame:
import pandas as pd
df = pd.DataFrame({'discount': discount, 'saleInRs': saleInRs})
df.plot(kind='scatter', x='discount', y='saleInRs')
you would encounter functional limitations:
| Feature | plt.scatter() | df.plot(kind='scatter') |
|---|---|---|
Marker size (s) | Accepts array for variable sizes | Accepts array via s= parameter |
| Marker shape | marker='*' works | marker='*' works |
| Edge color | edgecolor='blue' | edgecolors='blue' (note the 's') |
| Line width | linewidth=3 | linewidths=3 (note the 's') |
| Direct control | Full Matplotlib control | Wrapper around Matplotlib |
The pandas wrapper is convenient when your data is already in a DataFrame and you want quick exploratory plots, but it abstracts away some of Matplotlib's fine-grained control. Notice the parameter name differences: pandas uses edgecolors and linewidths (plural), while the raw Matplotlib function uses edgecolor and linewidth (singular) — though both accept the same values. …
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.