Skip to content
Exercises · Q14
Q.

Write the python statement for the following question on the basis of given dataset:

[Table: the printed dataset —

NameDegreeScore
0AparnaMBA90.0
1PankajBCANaN
2RamM.Tech80.0
3RameshMBA98.0
4NaveenNaN97.0
5KrrishnavBCA78.0
6BhawnaMBA89.0
]

a) To create the above DataFrame.

b) To print the Degree and maximum marks in each stream.

c) To fill the NaN with 76.

d) To set the index to Name.

e) To display the name and degree wise average marks of each student.

f) To count the number of students in MBA.

g) To print the mode marks BCA.

Telangana TsbieTextbookSubjective· 5mImportance★★★★★
64% · 34/53 Questions
🔒 Locked · start free trial →

You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.

Start your 14-day free trial to unlock the full solution →

This solution demonstrates how to perform common data manipulation tasks using the Pandas library in Python, including DataFrame creation, grouping and aggregation, handling missing values, setting indices, and filtering data.

Working with DataFrames in Pandas is fundamental for data analysis in Python. A DataFrame is a two-dimensional, size-mutable, and potentially heterogeneous tabular data structure with labeled axes (rows and columns). It's essentially like a spreadsheet or a SQL table. The power of Pandas lies in its ability to efficiently perform operations like filtering, sorting, grouping, and aggregating data, which are crucial for extracting insights from raw datasets.

Let's address each part of the question using the provided dataset.


(a) To create the above DataFrame.

To create a Pandas DataFrame, we typically use pd.DataFrame(). A common and clear way is to provide a dictionary where keys are column names and values are lists of data for each column. This allows us to explicitly define the structure and content of our DataFrame.

import pandas as pd
import numpy as np # For NaN values

# Data for the DataFrame
data = {
    'Name': ['Aparna', 'Pankaj', 'Ram', 'Ramesh', 'Naveen', 'Krrishnav', 'Bhawna'],
    'Degree': ['MBA', 'BCA', 'M.Tech', 'MBA', np.nan, 'BCA', 'MBA'],
    'Score': [90.0, np.nan, 80.0, 98.0, 97.0, 78.0, 89.0]
}

# Create the DataFrame
df = pd.DataFrame(data)

print("DataFrame created:")
print(df)
DataFrame created:
        Name  Degree  Score
0     Aparna     MBA   90.0
1     Pankaj     BCA    NaN
2        Ram  M.Tech   80.0
3     Ramesh     MBA   98.0
4     Naveen     NaN   97.0
5  Krrishnav     BCA   78.0
6     Bhawna     MBA   89.0

Explanation:

  • import pandas as pd: This line imports the Pandas library, aliasing it as pd for convenience.
  • import numpy as np: We import NumPy as np because Pandas uses np.nan to represent missing values (Not a Number).
  • data = {...}: A dictionary is created where each key corresponds to a column name ('Name', 'Degree', 'Score') and its value is a list containing the data for that column.
  • df = pd.DataFrame(data): This is the core command to construct the DataFrame from the data dictionary. Pandas automatically assigns a default integer index (0, 1, 2, ...) if not specified.

(b) To print the Degree and maximum marks in each stream.

To find the maximum marks for each degree stream, we need to group the DataFrame by the 'Degree' column and then apply an aggregation function (maximum) to the 'Score' column.

# Group by 'Degree' and find the maximum 'Score'
max_marks_per_degree = df.groupby('Degree')['Score'].max()

print("\nDegree and maximum marks in each stream:")
print(max_marks_per_degree)
Degree and maximum marks in each stream:
Degree
BCA       78.0
M.Tech    80.0
MBA       98.0
Name: Score, dtype: float64

Explanation:

  • df.groupby('Degree'): This groups the DataFrame rows based on unique values in the 'Degree' column. It creates a DataFrameGroupBy object.
  • ['Score']: After grouping, we select the 'Score' column because that's where we want to perform our aggregation.
  • .max(): This is the aggregation function applied to the 'Score' column within each group. It calculates the maximum score for each unique degree. The result is a Pandas Series where the index is 'Degree' and the values are the maximum scores.

(c) To fill the NaN with 76.

Missing data (represented as NaN) can interfere with calculations and analysis. Pandas provides the fillna() method to replace these missing values with a specified value.

# Fill NaN values in the entire DataFrame with 76
df_filled = df.fillna(76)

print("\nDataFrame after filling NaN with 76:")
print(df_filled)
DataFrame after filling NaN with 76:
        Name  Degree  Score
0     Aparna     MBA   90.0
1     Pankaj     BCA   76.0
2        Ram  M.Tech   80.0
3     Ramesh     MBA   98.0
4     Naveen      76   97.0
5  Krrishnav     BCA   78.0
6     Bhawna     MBA   89.0

Explanation:

  • df.fillna(76): This method replaces all NaN values in the DataFrame df with the integer 76.
  • It's important to note that fillna() returns a new DataFrame with the NaN values replaced. It does not modify the original DataFrame df unless inplace=True is specified. Here, we store the result in df_filled.

(d) To set the index to Name.

The index of a DataFrame provides labels for rows. By default, it's a range of integers. Setting a meaningful column (like 'Name' in this case) as the index can make data retrieval and alignment more intuitive.

# Set 'Name' column as the index
df_indexed = df_filled.set_index('Name')

print("\nDataFrame with 'Name' as index:")
print(df_indexed)
DataFrame with 'Name' as index:
           Degree  Score
Name                    
Aparna        MBA   90.0
Pankaj        BCA   76.0
Ram        M.Tech   80.0
Ramesh        MBA   98.0
Naveen         76   97.0
Krrishnav     BCA   78.0
Bhawna        MBA   89.0

Explanation:

  • df_filled.set_index('Name'): This method takes the 'Name' column and uses its values as the new row labels (index) for the DataFrame.
  • Similar to fillna(), set_index() returns a new DataFrame by default. We store this new DataFrame in df_indexed. If you wanted to modify df_filled directly, you would use df_filled.set_index('Name', inplace=True).

(e) To display the name and degree wise average marks of each student.

This phrasing implies grouping by both 'Name' and 'Degree' and then calculating the average of 'Score'. Since each student (Name) has a unique entry with a specific Degree and Score in this dataset, grouping by both 'Name' and 'Degree' and then taking the mean of 'Score' will effectively return each student's individual score, as there's only one score per (Name, Degree) combination to average.

# Group by 'Name' and 'Degree' and calculate the average 'Score'
# We use the DataFrame after filling NaN values for consistent results.
avg_marks_name_degree = df_filled.groupby(['Name', 'Degree'])['Score'].mean()

print("\nName and degree wise average marks of each student:")
print(avg_marks_name_degree)
Name and degree wise average marks of each student:
Name       Degree
Aparna     MBA       90.0
Bhawna     MBA       89.0
Krrishnav  BCA       78.0
Naveen     76        97.0
Pankaj     BCA       76.0
Ram        M.Tech    80.0
Ramesh     MBA       98.0
Name: Score, dtype: float64

Explanation:

  • df_filled.groupby(['Name', 'Degree']): This groups the DataFrame by unique combinations of values in both the 'Name' and 'Degree' columns. This creates a MultiIndex for the resulting Series.
  • ['Score'].mean(): For each unique (Name, Degree) group, the average of the 'Score' column is calculated. As discussed, with unique student entries, this effectively lists each student's score.

(f) To count the number of students in MBA.

To count students in a specific degree, we first need to filter the DataFrame to include only those students whose 'Degree' is 'MBA'. Then, we can count the number of rows in this filtered subset.

# Filter for students with 'MBA' degree
mba_students = df_filled[df_filled['Degree'] == 'MBA']

# Count the number of students in the filtered DataFrame
num_mba_students = len(mba_students)

print(f"\nNumber of students in MBA: {num_mba_students}")
Number of students in MBA: 3

Explanation: …

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.