Q.(a) List any two characteristics of Pandas library in Python.
🔒You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
🔒 Start your 14-day free trial to unlock the full solution →Part (a)Concept understanding — Series vs Data Structures
Series vs Data Structures — A First Look
Imagine you walk into a library. You see books arranged on shelves. That arrangement — the shelves, the order, the categories — is a data structure. It is a way of organising information so you can find it, add to it, or remove it without chaos. Now imagine you pick up a single book from that shelf. That book has a title, an author, and a year of publication. That single, self-contained record is like a series — one item, with its own internal order and meaning.
In the world of information management, these two ideas are foundational. A data structure is the container, the system, the framework. A series is one coherent unit of content inside that framework. You cannot have a meaningful series without a data structure to hold it, and a data structure without series would be an empty shelf.
The Intuition: A Filing Cabinet
Think of a filing cabinet in an office. The cabinet itself, with its drawers and labelled dividers, is a data structure. It tells you where things go, how they are grouped, and how to retrieve them. Inside one drawer, you might have a folder labelled "Quarterly Reports 2024." That folder is a series — a set of related documents that belong together, arranged in a logical order (say, by month). The series has a beginning and an end; it is a complete unit of information.
Now, if you pull out that folder and look inside, you see individual reports. Each report is a record or an item. The series is the collection of those items, bound by a common theme and a consistent structure. The data structure is the cabinet that keeps the series safe and findable.
The Precise Meaning
In formal terms, a data structure is a specialised format for organising, processing, retrieving, and storing data. It defines the relationships between data elements and the operations that can be performed on them. Examples include arrays, linked lists, stacks, queues, trees, and graphs. Each has its own rules about how data is added, removed, or accessed.
A series (in the context of data management) is a sequence of related data points or records that share a common schema or purpose. It is often time-ordered or logically ordered. In a database, a series might be a table of monthly sales figures. In a library catalogue, it might be a set of books published under the same title across different years.
The key distinction is this: a data structure is about how data is organised and stored. A series is about what data belongs together as a meaningful unit. The data structure is the architecture; the series is the content.
Why It Matters
For a commerce or humanities student, this distinction is not just technical jargon. It affects how you think about information in the real world.
- In business, a company's customer database is a data structure. Each customer's purchase history is a series. If the data structure is poorly designed, you cannot find the series you need. If the series is incomplete, you cannot analyse customer behaviour.
- In history, an archive of letters is a data structure. The letters from a particular decade form a series. The structure determines whether you can trace a narrative across time. …
Part (b)Concept understanding — DataFrame Querying
DataFrame Querying: Finding What You Need in a Table
Think of a DataFrame as a digital spreadsheet — rows of records (like each customer, each student, each transaction) and columns of attributes (name, age, city, purchase amount). Querying is simply the act of asking questions of that table: "Show me only the rows where the city is Delhi" or "Give me all students who scored above 80 in Economics."
You already do this in everyday life. When you open your contacts app and type "R" to see only names starting with R, you are querying. When you filter your email to show only unread messages, you are querying. A DataFrame query is the same idea, but more powerful and precise.
The Precise Meaning
In the context of a DataFrame (a two-dimensional, labelled data structure), querying means selecting a subset of rows and/or columns based on some condition or set of conditions. The original DataFrame remains unchanged; you get a new, smaller view of the data that answers your specific question.
The key parts of any query are:
- Which rows? You specify a condition — for example, rows where
age > 18orcity == "Mumbai". - Which columns? You specify which attributes you want to see — for example, only
nameandscore, or all columns.
A query can combine multiple conditions using logical operators like AND, OR, and NOT. For instance: "Show me all rows where the city is Bangalore AND the purchase amount is greater than 5000."
Why It Matters
DataFrame querying is the foundation of data analysis. Without it, you are stuck looking at the entire table — thousands or millions of rows — and trying to spot patterns by eye. That is impossible for any real-world dataset.
Querying lets you:
- Find specific records — "Which customers have not paid their bill?"
- Compare groups — "How do sales in the North zone differ from the South zone?"
- Prepare data for further analysis — "Extract only the records from the last financial year."
- Check for errors — "Are there any rows where the age is negative?"
A query does not change the original data. It creates a new, filtered view. This is crucial: you can run many different queries on the same DataFrame without ever altering the source. The original table stays safe.
A Simple Example in Words
Imagine a DataFrame called Students with columns: Name, Stream, Percentage, City.
A query like "Show me all Commerce students" would return a new table containing only those rows where the Stream column has the value "Commerce". All columns — Name, Stream, Percentage, City — would still appear, but only for Commerce students.
A more specific query: "Show me the names and percentages of Commerce students from Delhi who scored above 75." This query:
- Filters rows to those where
Streamis "Commerce" ANDCityis "Delhi" ANDPercentageis greater than 75. - Selects only the
NameandPercentagecolumns.
The result is a small, focused table — exactly the information you need, nothing more. …
Part (a)
Two characteristics of the Pandas library:
- It provides two powerful labelled data structures — Series (1-D) and DataFrame (2-D) — for handling tabular data with row/column labels. …
Part (a): two Pandas characteristics — labelled Series/DataFrame structures, and automatic alignment with missing-data handling (fast vectorised ops, easy file I/O).
Part (b): Boolean indexing filters rows of a DataFrame using a True/False condition, e.g. df[df['Score'] > 80].
Part (a)
Pandas is Python's core library for data analysis. Two of its defining characteristics:
- Labelled data structures — Series and DataFrame. A Series is a one-dimensional labelled array (like one spreadsheet column) and a DataFrame is a two-dimensional labelled table (like a whole sheet or SQL table). Because rows and columns carry labels, you can access data by name rather than by position. …
- CBSE 2026Set 90/1/11 markMCQQ.Aditya is working on a DataFrame named df. He has written the statement : print(df.loc['S2']) What will the above statement do ? (A) Display the data of the row having label 'S2'. (B) Display the columns of the DataFrame. (C) Display the data type of the column having label 'S2'. (D) Display the index numbers of the DataFrame.
›Reveal solutionSolution
The statement
print(df.loc['S2'])will display all the data contained within the row of the DataFramedfthat has 'S2' as its label.When working with DataFrames in Pandas, a fundamental task is to access specific parts of your data. DataFrames are essentially tabular data structures, much like a spreadsheet, with rows and columns. Each row and each column can have a label, which acts as its identifier. Pandas provides powerful accessors like
.locand.ilocto retrieve data based on these labels or integer positions, respectively.The
.locaccessor is specifically designed for label-based indexing. This means you use the actual labels (names) of your rows and columns to select data, rather than their numerical positions. When you usedf.loc['S2'], you are instructing Pandas to look for a row in the DataFramedfwhose index label (or row label) is exactly 'S2'.ImportantThe
.locaccessor always uses labels. If you provide a single label, like 'S2' in this case, it is interpreted as a row label.Upon finding the row with the label 'S2', the statement
df.loc['S2']will retrieve all the data points (values) present in that particular row across all its columns. The result of this operation will be a Pandas Series, where the index of the Series will correspond to the column labels of the original DataFrame, and the values of the Series will be the data from the 'S2' row for each respective column.Finally, wrapping this with
print()simply outputs this resulting Series to the console, making the data of that specific row visible.Let's consider the given options:
- (A) Display the data of the row having label 'S2'. This aligns perfectly with our understanding of how
df.loc['S2']functions. It targets a specific row by its label and retrieves all its associated data. …
- (A) Display the data of the row having label 'S2'. This aligns perfectly with our understanding of how
- CBSE 2026Set 90/1/11 markMCQQ.In the context of creating a Pandas Series from a dictionary, which of the following statement is correct ? (A) The values of the dictionary become the indices of the Series. (B) The keys of the dictionary become the values of the Series. (C) The keys of the dictionary become the indices of the Series. (D) The Series will have default integer indices starting from 0, ignoring the dictionary keys.
›Reveal solutionSolution
When creating a Pandas Series from a dictionary, the dictionary's keys become the Series' indices, and the dictionary's values become the Series' values.
To understand how a Pandas Series is created from a dictionary, we first need to grasp the fundamental nature of both these data structures. A Pandas Series is essentially a one-dimensional array-like object capable of holding any data type, but with a crucial difference: it has an associated array of data labels, called its index. This index allows for efficient data retrieval and manipulation using meaningful labels rather than just numerical positions.
A Python dictionary, on the other hand, is a collection of key-value pairs. Each key in a dictionary must be unique and immutable, serving as a distinct identifier for its corresponding value. The values can be of any data type and do not need to be unique. Dictionaries are designed for fast lookups based on these unique keys.
When you use a dictionary to construct a Pandas Series, the design philosophy of Pandas aims to preserve the inherent logical mapping present in the dictionary. The unique, descriptive keys of the dictionary are perfectly suited to serve as the labels for the Series, which are its indices. Consequently, the data associated with these keys in the dictionary naturally become the actual data points, or values, within the Series.
For example, if you have a dictionary like
{'apple': 10, 'banana': 20, 'cherry': 30}, and you create a Series from it, 'apple', 'banana', and 'cherry' will become the indices of the Series, while 10, 20, and 30 will be the corresponding values. This mapping ensures that the Series retains the meaningful associations established in the original dictionary.ImportantThis behavior is a core design choice in Pandas, leveraging the key-value structure of dictionaries to provide meaningful, custom labels (indices) for the Series data.
Let's consider the given options in light of this understanding:
- (A) The values of the dictionary become the indices of the Series. This is incorrect. If dictionary values became indices, the original keys would be lost as labels, and the values themselves might not always be unique or suitable as indices. …
- CBSE 2026Set 90/1/11 markMCQQ.Q. 20 and Q. 21 are Assertion (A) and Reason (R) Type questions. Choose the correct option as : (A) Both (A) and (R) are True, and (R) correctly explains (A). (B) Both (A) and (R) are True, but (R) does not correctly explain (A). (C) (A) is True, but (R) is False. (D) (A) is False, but (R) is True. Assertion (A) : Pandas DataFrame can store values of multiple data types in multiple Columns. Reason (R) : DataFrames are implemented using 2D arrays, which allows only numeric values.
›Reveal solutionSolution
Pandas DataFrames can indeed hold diverse data types across their columns, but they are not implemented as a single 2D array that restricts values to only numeric types.
Let's break down the nature of Pandas DataFrames and their underlying structure to understand this assertion and reason.
Pandas is a powerful library in Python, widely used for data manipulation and analysis. Its two primary data structures are the Series and the DataFrame. Understanding the distinction between these is key to grasping how DataFrames handle different data types.
A Pandas Series is essentially a one-dimensional labeled array capable of holding data of any single type (integer, float, string, boolean, etc.). Think of it like a single column in a spreadsheet. All elements within one Series must generally conform to the same data type. If you try to mix types, Pandas will often "upcast" them to a common, more general type (like
objectfor mixed strings and numbers).A Pandas DataFrame, on the other hand, is a two-dimensional labeled data structure with columns of potentially different types. You can visualize it as a table, much like a spreadsheet or a SQL table. Each column in a DataFrame is, in fact, a Pandas Series. This is a crucial point: because each column is an independent Series, and each Series can hold a specific data type, a DataFrame can naturally accommodate columns with varying data types.
NoteThis ability to have columns of different data types (e.g., one column for names as strings, another for ages as integers, and a third for salaries as floats) is one of the most powerful features of DataFrames, making them incredibly flexible for real-world datasets.
Now, let's evaluate the given statements:
Assertion (A): Pandas DataFrame can store values of multiple data types in multiple Columns.
This statement is True. As explained, a DataFrame is a collection of Series, and each Series (column) can have its own data type. Therefore, a DataFrame can easily have a column of integers, another of strings, another of floating-point numbers, and so on.
Reason (R): DataFrames are implemented using 2D arrays, which allows only numeric values.
This statement is False on two counts. …
- CBSE 2025Set 90/1/11 markMCQQ.Which of the following data structures is used for storing one-dimensional labelled data in Python Pandas? (A) Integer (B) Dictionary (C) Series (D) DataFrame
›Reveal solutionSolution
The Pandas Series is the data structure specifically designed for storing one-dimensional labelled data in Python.
When we work with data in Python, especially for analysis, we often need more specialized tools than the basic lists or dictionaries that come with the language. This is where libraries like Pandas come in. Pandas provides powerful, flexible, and easy-to-use data structures that are built on top of Python, making data manipulation and analysis much more efficient. The question asks about a specific type of data: "one-dimensional labelled data." Let's break down what that means and then see which Pandas structure fits.
"One-dimensional" refers to data that can be thought of as a single sequence or a list of items, like a single column of numbers, a list of names, or a series of temperatures recorded over time. It doesn't have multiple rows and multiple columns simultaneously, like a spreadsheet. "Labelled data" means that each item in this sequence isn't just accessed by its numerical position (like
list[0]), but also by a meaningful label or index. Think of it like a dictionary where each value has a key, or a spreadsheet column where each row has a descriptive label.Now, let's consider the given options in the context of Pandas:
-
(A) Integer: An integer is a single numerical value (e.g.,
5,100). It is a basic data type, not a data structure designed to store a collection of items, let alone one-dimensional labelled data. So, this option is incorrect. -
(B) Dictionary: A Python dictionary (
dict) does store labelled data, where each value is associated with a unique key. For example,{'apple': 10, 'banana': 20}. While dictionaries are fundamental to Python and can be used to create Pandas data structures, a dictionary itself is a core Python data type, not a Pandas-specific data structure designed for advanced data analysis with features like vectorized operations or handling missing data in the way Pandas does. -
(C) Series: This is precisely the data structure in Pandas designed for one-dimensional labelled data. A Pandas Series can be thought of as a single column of data, where each element has an associated label, called an "index." This index can be numerical (like
0, 1, 2...) or custom (like['Jan', 'Feb', 'Mar']). All elements within a Series are typically of the same data type (homogeneous), which makes it very efficient for operations. For example, you could have a Series storing the population of different cities, where the city names are the labels (index) and the population figures are the data. …
-
- CBSE 2025Set 90/1/11 markMCQQ.In Python Pandas, DataFrame.______[] is used for label indexing with DataFrames. (A) label (B) index (C) labindex (D) loc
›Reveal solutionSolution
The
DataFrame.loc[]accessor is used for label-based indexing in Pandas DataFrames, allowing selection of data by row and column labels. The correct option is (D).Concept and Intuition
In Pandas, a DataFrame is a two-dimensional, size-mutable, potentially heterogeneous tabular data structure with labeled axes (rows and columns). When you work with DataFrames, you often need to select specific subsets of your data. There are two primary ways to do this: by label and by integer position.
Label-based indexing means you are referring to the actual names of your rows (the index) and columns. For example, if you have a DataFrame with rows labeled 'Alice', 'Bob', 'Charlie' and columns 'Age', 'City', you would use these labels to retrieve data. This is intuitive because it directly maps to how you might think about your data in a spreadsheet. Pandas provides a dedicated accessor for this purpose, ensuring clarity and preventing ambiguity, especially when row labels might coincidentally be integers.
Step-by-step Solution
-
Understanding Indexing in Pandas:
Pandas DataFrames offer powerful ways to select data. The primary methods for selection are
locandiloc. These are not functions, but accessors that allow you to use square bracket notation[]immediately after them to specify your selection criteria. -
Label-based Indexing with
loc:The
locaccessor is specifically designed for label-based indexing. This means you provide the actual labels (names) of the rows and columns you want to select.Consider a simple DataFrame:
import pandas as pd data = {'Name': ['Alice', 'Bob', 'Charlie'], 'Age': [25, 30, 35], 'City': ['New York', 'London', 'Paris']} df = pd.DataFrame(data, index=['A', 'B', 'C']) print(df)This DataFrame looks like:
Name Age City A Alice 25 New York B Bob 30 London C Charlie 35 ParisTo select the 'Age' of 'Bob' using labels, you would use
df.loc['B', 'Age'].Tiplocis inclusive when slicing with labels. For example,df.loc['A':'C', 'Name']would include rows 'A', 'B', and 'C'. -
Contrasting with Position-based Indexing (
iloc):While
locuses labels,iloc(integer location) is used for position-based indexing. It treats the DataFrame as a grid, where rows and columns are accessed by their integer positions, starting from 0.For the DataFrame above,
df.iloc[1, 1]would also give you the 'Age' of 'Bob' (row at position 1, column at position 1). It's crucial to understand the difference to avoid errors, especially when your row labels are integers. …
-
- CBSE 2024Set 90/1/11 markMCQQ.Assertion (A) : A Series is a one dimensional array and a DataFrame is a two-dimensional array containing sequence of values of any data type. (int, float, list, string, etc.) Reason (R) : Both Series and DataFrames have by default numeric indexes starting from zero. (A) Both (A) and (R) are true and (R) is the correct explanation for (A). (B) Both (A) and (R) are true and (R) is not the correct explanation for (A). (C) (A) is true and (R) is false. (D) (A) is false but (R) is true.
›Reveal solutionSolution
The Assertion is correct in describing Series as one-dimensional and DataFrame as two-dimensional, but the Reason, while true, does not explain why that dimensional difference exists — so both statements are true but unrelated.
Let’s begin with the Assertion. When the NCERT textbook introduces pandas, it defines a Series as a one-dimensional labelled array capable of holding any data type — integers, floats, strings, even Python lists or other objects. A DataFrame, by contrast, is a two-dimensional labelled data structure, essentially a collection of Series sharing a common index. So the Assertion is spot on: Series is one-dimensional (like a single column), DataFrame is two-dimensional (like a table with rows and columns), and both can hold mixed data types. That part is correct.
Now the Reason: it states that both Series and DataFrames have by default numeric indexes starting from zero. This is also true. When you create a Series or DataFrame without specifying an index, pandas automatically assigns a default integer index: 0, 1, 2, … up to (n-1). So the Reason is factually accurate.
NoteThe default index is indeed numeric and zero-based, but you can override it with custom labels (strings, dates, etc.) — the default is just a convenience. …
- CBSE 2023Set 90/1/11 markMCQQ.Which of the following is a two-dimensional labelled data structure of Python? (A) Relation (B) Data frame (C) Series (D) Square
›Reveal solutionSolution
A DataFrame is the two-dimensional labelled data structure in Python's Pandas library, akin to a spreadsheet or a SQL table.
When working with data in Python, especially for analysis and manipulation, the Pandas library provides highly efficient and flexible data structures. These structures are designed to handle various types of data, making them indispensable tools. The question specifically asks for a two-dimensional labelled data structure. To understand this, we need to differentiate between the primary data structures offered by Pandas: Series and DataFrame.
A Series is a one-dimensional labelled array capable of holding any data type (integers, strings, floating point numbers, Python objects, etc.). Think of it like a single column in a spreadsheet or a list in Python, but with an added feature: each item in the Series has a unique label, called an index. This index allows for easy retrieval and manipulation of data based on these labels, rather than just numerical positions. For example, you might have a Series storing the marks of students, where each mark is associated with a student's roll number or name as its label. Because it's a single sequence of values, it is considered one-dimensional.
NoteThe "labelled" aspect means that each piece of data isn't just at a numerical position (like
list[0]), but can also be accessed by a custom label (likeseries['Roll_No_101']).In contrast, a DataFrame is a two-dimensional labelled data structure with columns of potentially different types. It is essentially a table, much like a spreadsheet or a SQL database table. A DataFrame has both a labelled row index and labelled columns. This means you can access data by referring to both its row label and its column label. For instance, if you have a table of student data, you might have columns for 'Name', 'Roll Number', 'Marks', and 'Grade'. Each row would represent a different student, and each column would represent a different attribute. This tabular arrangement, with both rows and columns, makes it inherently two-dimensional. …
🎓Unlock everything free for 14 days
- ✓Full step-by-step solutions
- ✓Concept-first explanations
- ✓Methods, shortcuts & mistakes
- ✓PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.