Quick answer: A data frame is R’s spreadsheet: rows of observations, columns of variables. Create one with df <- data.frame(name = c("Asha", "Rahul"), score = c(91, 84)), inspect it with head(df), str(df) and summary(df), filter rows with df[df$score > 85, ], and compute per-group summaries with aggregate(score ~ city, data = df, FUN = mean). This guide walks through the first session after installing R and RStudio.
Creating and inspecting a data frame
df <- data.frame(
name = c("Asha", "Rahul", "Meera", "Arjun"),
city = c("Chennai", "Delhi", "Chennai", "Mumbai"),
score = c(91, 84, 78, 95)
)
head(df) # first rows
str(df) # structure: column types
summary(df) # statistical overview
nrow(df) # row count
The assignment arrow <- is R’s convention (equals works, but style guides prefer the arrow). str() is the command to run first, always: it reveals whether R read your numbers as numbers (num) or as text (chr) — the root of many mysterious downstream errors.
Everyday operations
df$score # one column
df[df$score > 85, ] # rows where score > 85 (note the comma)
df[df$city == "Chennai", ] # rows for one city
df[order(-df$score), ] # sort descending
mean(df$score) # simple stats
aggregate(score ~ city, data = df, FUN = mean) # mean per city
The bracket notation df[rows, columns] is the key mental model — filter with a condition before the comma, select columns after it, and leave either side blank to mean “all”. The comma matters: df[df$score > 85] without it is a different (and confusing) operation.
The tidyverse upgrade path
install.packages("tidyverse") # once
library(dplyr)
df %>%
filter(score > 80) %>%
group_by(city) %>%
summarise(mean_score = mean(score))
The pipe %>% reads left-to-right like a sentence — filter, then group, then summarize — which is why dplyr became the standard for real R work. Learn base R brackets first (they are everywhere in documentation), then let dplyr make it pleasant.
For one-to-one R training with your own datasets and instructor review, Ampersand Academy teaches R programming and data analysis one-to-one.
Frequently asked questions
What is a data frame in R?
A data frame is R’s table structure: columns of equal length that can hold different types (numbers, text, factors). It is the standard structure for datasets and the output of read.csv.
What does the arrow <- mean in R?
It is the assignment operator: df <- data.frame(…) stores the result in df. It is the community-preferred style; = also works in most contexts.
How do I filter rows in R?
Base R: df[df$score > 85, ] – condition before the comma inside single brackets. With dplyr: df %>% filter(score > 85).
Why is my numeric column a character in R?
Usually stray text, commas as thousand separators or empty strings in the source. Check with str(), clean the values, and convert with as.numeric().
Should I learn base R or tidyverse first?
Learn base R brackets and core functions first – they appear everywhere – then add dplyr’s pipes for day-to-day analysis. Most curricula, including one-to-one training, teach them in that order.

