Site icon Ampersand Tutorials

Pandas First Steps: Load and Analyze Data (2026)

Quick answer: Pandas is the Python library for working with tables of data. Install it with pip install pandas (or it comes bundled with Anaconda), load a CSV with df = pd.read_csv('data.csv'), and inspect it with df.head(), df.info() and df.describe(). Those three inspection calls are the muscle memory every data analysis starts with. This guide walks through the first full session: load, look, select and filter.

Installing pandas the right way

If you followed our Anaconda installation guide, pandas is already installed — open Jupyter and import it. Otherwise create a clean project setup:

python -m venv venv
venv\Scripts\activate        # Windows  (venv/bin/activate on Mac/Linux)
pip install pandas

The virtual environment keeps this project’s packages separate from your system Python — the habit that prevents the classic “it works on my machine” failures. Verify: python -c "import pandas; print(pandas.__version__)".

Loading and inspecting your first dataset

import pandas as pd

df = pd.read_csv('students.csv')   # DataFrame = your table in memory

df.head()        # first 5 rows - the 'did it load?' check
df.shape         # (rows, columns)
df.info()        # column types + missing-value overview
df.describe()    # count, mean, std, min, max for numeric columns

Read those outputs in a fixed order, every time: does shape match what you expected? Does info() show the right types (dates as objects means parsing needed)? Are there missing values (info() shows non-null counts)? Ninety percent of bad analysis comes from skipping this five-minute look.

Selecting and filtering: the daily operations

df['grade']                     # one column (Series)
df[['name', 'grade']]           # multiple columns
df[df['grade'] == 'A']          # filter rows by condition
df[(df['score'] > 80) & (df['city'] == 'Chennai')]   # AND (&), OR (|)

df.sort_values('score', ascending=False).head(10)    # top 10 by score
df.groupby('city')['score'].mean()                   # average per city

Note the pattern: conditions in parentheses (Python’s &/| bind tighter than comparisons otherwise), and groupby as the split-apply-combine workhorse. With these six operations — select columns, filter rows, sort, group, then aggregate — you can answer most first-pass business questions. Real practice data is one search away: pandas ships with titanic-style sample CSVs all over the web, or load anything from data.gov.

If you want structured, one-to-one coaching on pandas and analytics with your own datasets, Ampersand Academy teaches Python analytics with instructor-reviewed exercises.

Frequently asked questions

How do I install pandas in Python?

Run pip install pandas in your project environment. If you use Anaconda, pandas comes preinstalled. Always install inside a virtual environment to keep project dependencies isolated.

What is a DataFrame in pandas?

A DataFrame is pandas’ table structure: rows and labeled columns, like a spreadsheet in memory. You load CSVs, Excel files or databases into DataFrames and analyze them with selection, filtering and grouping operations.

How do I filter rows in pandas?

Put a condition in square brackets: df[df[‘score’] > 80]. Combine conditions with & and | inside parentheses, one pair per condition.

What is the difference between loc and iloc?

loc selects by label (column names, index values), iloc selects by integer position. df.loc[0, ‘name’] gets the name column value at index label 0; df.iloc[0, 1] gets the first row, second column regardless of labels.

How do I handle missing values in pandas?

Detect with df.isnull().sum() to count per column, then either drop rows with df.dropna() or fill with df.fillna(value). The right choice depends on why the data is missing – inspect before deleting.

Exit mobile version