Write the python statement for the following question on the basis of given dataset:
[Table: the printed dataset —
| Name | Degree | Score | |
|---|---|---|---|
| 0 | Aparna | MBA | 90.0 |
| 1 | Pankaj | BCA | NaN |
| 2 | Ram | M.Tech | 80.0 |
| 3 | Ramesh | MBA | 98.0 |
| 4 | Naveen | NaN | 97.0 |
| 5 | Krrishnav | BCA | 78.0 |
| 6 | Bhawna | MBA | 89.0 |
| ] |
a) To create the above DataFrame.
b) To print the Degree and maximum marks in each stream.
c) To fill the NaN with 76.
d) To set the index to Name.
e) To display the name and degree wise average marks of each student.
f) To count the number of students in MBA.
g) To print the mode marks BCA.
You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
Start your 14-day free trial to unlock the full solution →This solution demonstrates how to perform common data manipulation tasks using the Pandas library in Python, including DataFrame creation, grouping and aggregation, handling missing values, setting indices, and filtering data.
Working with DataFrames in Pandas is fundamental for data analysis in Python. A DataFrame is a two-dimensional, size-mutable, and potentially heterogeneous tabular data structure with labeled axes (rows and columns). It's essentially like a spreadsheet or a SQL table. The power of Pandas lies in its ability to efficiently perform operations like filtering, sorting, grouping, and aggregating data, which are crucial for extracting insights from raw datasets.
Let's address each part of the question using the provided dataset.
(a) To create the above DataFrame.
To create a Pandas DataFrame, we typically use pd.DataFrame(). A common and clear way is to provide a dictionary where keys are column names and values are lists of data for each column. This allows us to explicitly define the structure and content of our DataFrame.
import pandas as pd
import numpy as np # For NaN values
# Data for the DataFrame
data = {
'Name': ['Aparna', 'Pankaj', 'Ram', 'Ramesh', 'Naveen', 'Krrishnav', 'Bhawna'],
'Degree': ['MBA', 'BCA', 'M.Tech', 'MBA', np.nan, 'BCA', 'MBA'],
'Score': [90.0, np.nan, 80.0, 98.0, 97.0, 78.0, 89.0]
}
# Create the DataFrame
df = pd.DataFrame(data)
print("DataFrame created:")
print(df)
DataFrame created:
Name Degree Score
0 Aparna MBA 90.0
1 Pankaj BCA NaN
2 Ram M.Tech 80.0
3 Ramesh MBA 98.0
4 Naveen NaN 97.0
5 Krrishnav BCA 78.0
6 Bhawna MBA 89.0
Explanation:
import pandas as pd: This line imports the Pandas library, aliasing it aspdfor convenience.import numpy as np: We import NumPy asnpbecause Pandas usesnp.nanto represent missing values (Not a Number).data = {...}: A dictionary is created where each key corresponds to a column name ('Name','Degree','Score') and its value is a list containing the data for that column.df = pd.DataFrame(data): This is the core command to construct the DataFrame from thedatadictionary. Pandas automatically assigns a default integer index (0, 1, 2, ...) if not specified.
(b) To print the Degree and maximum marks in each stream.
To find the maximum marks for each degree stream, we need to group the DataFrame by the 'Degree' column and then apply an aggregation function (maximum) to the 'Score' column.
# Group by 'Degree' and find the maximum 'Score'
max_marks_per_degree = df.groupby('Degree')['Score'].max()
print("\nDegree and maximum marks in each stream:")
print(max_marks_per_degree)
Degree and maximum marks in each stream:
Degree
BCA 78.0
M.Tech 80.0
MBA 98.0
Name: Score, dtype: float64
Explanation:
df.groupby('Degree'): This groups the DataFrame rows based on unique values in the 'Degree' column. It creates aDataFrameGroupByobject.['Score']: After grouping, we select the 'Score' column because that's where we want to perform our aggregation..max(): This is the aggregation function applied to the 'Score' column within each group. It calculates the maximum score for each unique degree. The result is a Pandas Series where the index is 'Degree' and the values are the maximum scores.
(c) To fill the NaN with 76.
Missing data (represented as NaN) can interfere with calculations and analysis. Pandas provides the fillna() method to replace these missing values with a specified value.
# Fill NaN values in the entire DataFrame with 76
df_filled = df.fillna(76)
print("\nDataFrame after filling NaN with 76:")
print(df_filled)
DataFrame after filling NaN with 76:
Name Degree Score
0 Aparna MBA 90.0
1 Pankaj BCA 76.0
2 Ram M.Tech 80.0
3 Ramesh MBA 98.0
4 Naveen 76 97.0
5 Krrishnav BCA 78.0
6 Bhawna MBA 89.0
Explanation:
df.fillna(76): This method replaces allNaNvalues in the DataFramedfwith the integer76.- It's important to note that
fillna()returns a new DataFrame with theNaNvalues replaced. It does not modify the original DataFramedfunlessinplace=Trueis specified. Here, we store the result indf_filled.
(d) To set the index to Name.
The index of a DataFrame provides labels for rows. By default, it's a range of integers. Setting a meaningful column (like 'Name' in this case) as the index can make data retrieval and alignment more intuitive.
# Set 'Name' column as the index
df_indexed = df_filled.set_index('Name')
print("\nDataFrame with 'Name' as index:")
print(df_indexed)
DataFrame with 'Name' as index:
Degree Score
Name
Aparna MBA 90.0
Pankaj BCA 76.0
Ram M.Tech 80.0
Ramesh MBA 98.0
Naveen 76 97.0
Krrishnav BCA 78.0
Bhawna MBA 89.0
Explanation:
df_filled.set_index('Name'): This method takes the 'Name' column and uses its values as the new row labels (index) for the DataFrame.- Similar to
fillna(),set_index()returns a new DataFrame by default. We store this new DataFrame indf_indexed. If you wanted to modifydf_filleddirectly, you would usedf_filled.set_index('Name', inplace=True).
(e) To display the name and degree wise average marks of each student.
This phrasing implies grouping by both 'Name' and 'Degree' and then calculating the average of 'Score'. Since each student (Name) has a unique entry with a specific Degree and Score in this dataset, grouping by both 'Name' and 'Degree' and then taking the mean of 'Score' will effectively return each student's individual score, as there's only one score per (Name, Degree) combination to average.
# Group by 'Name' and 'Degree' and calculate the average 'Score'
# We use the DataFrame after filling NaN values for consistent results.
avg_marks_name_degree = df_filled.groupby(['Name', 'Degree'])['Score'].mean()
print("\nName and degree wise average marks of each student:")
print(avg_marks_name_degree)
Name and degree wise average marks of each student:
Name Degree
Aparna MBA 90.0
Bhawna MBA 89.0
Krrishnav BCA 78.0
Naveen 76 97.0
Pankaj BCA 76.0
Ram M.Tech 80.0
Ramesh MBA 98.0
Name: Score, dtype: float64
Explanation:
df_filled.groupby(['Name', 'Degree']): This groups the DataFrame by unique combinations of values in both the 'Name' and 'Degree' columns. This creates a MultiIndex for the resulting Series.['Score'].mean(): For each unique (Name, Degree) group, the average of the 'Score' column is calculated. As discussed, with unique student entries, this effectively lists each student's score.
(f) To count the number of students in MBA.
To count students in a specific degree, we first need to filter the DataFrame to include only those students whose 'Degree' is 'MBA'. Then, we can count the number of rows in this filtered subset.
# Filter for students with 'MBA' degree
mba_students = df_filled[df_filled['Degree'] == 'MBA']
# Count the number of students in the filtered DataFrame
num_mba_students = len(mba_students)
print(f"\nNumber of students in MBA: {num_mba_students}")
Number of students in MBA: 3
Explanation: …
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.