This guide is for readers who can run R code and want a systematic introduction to the ggplot2 methods used most often in real work. Every example uses a base R or ggplot2 built-in dataset, so no external data download is required. You can copy each code chunk directly into RStudio.
Think about a plot in this order:
The core mental model is ggplot(data, aes(…)) + geom_*() +
scale_*() + facet_*() + coord_*() + theme_*(). Each
+ adds or modifies a layer.
| Dataset | Contents | Typical uses |
|---|---|---|
mpg |
Fuel economy and specifications for 1999 and 2008 models | Scatterplots, bars, facets |
diamonds |
Specifications and prices for about 54,000 diamonds | Large data, distributions, log scales |
economics |
Monthly U.S. economic indicators | Dates and time series |
mtcars |
Performance data for 32 cars | Small demonstrations and modeling |
dim(mpg)
#> [1] 234 11
names(mpg)
#> [1] "manufacturer" "model" "displ" "year" "cyl" "trans"
#> [7] "drv" "cty" "hwy" "fl" "class"
head(mpg, 3)ggplot() creates the canvas, aes() declares
mappings from variables to visual properties, and
geom_point() says to draw one point for each row.
ggplot(data = mpg, mapping = aes(x = displ, y = hwy)) +
geom_point() +
labs(
title = "Larger engines usually have lower highway fuel economy",
x = "Engine displacement (L)",
y = "Highway fuel economy (mpg)"
)Figure 1: Engine displacement versus highway fuel economy.
Data and mappings can also be supplied to individual layers. A global
aes() is inherited by later layers; a local
aes() affects only its own layer.
aes(): map a variable to a
visual property. This creates data-driven variation and usually a
guide.aes(): set one constant value
for every mark. This usually creates no guide.ggplot(mpg, aes(displ, hwy)) +
geom_point(aes(color = class), alpha = 0.72, size = 2.2) +
labs(
title = "Put variables inside aes() and constants outside",
color = "Vehicle class",
x = "Engine displacement (L)",
y = "Highway fuel economy (mpg)"
) +
guides(color = guide_legend(nrow = 2, byrow = TRUE,
override.aes = list(alpha = 1, size = 3)))Figure 2: Color is mapped to vehicle class; alpha and point size are fixed.
One of the most common mistakes is
aes(color = "steelblue"). It treats the string as a
category and creates a legend rather than setting points to blue.
Here, point color is mapped to vehicle class while trend lines are
grouped only by drive type (drv). group
controls which observations belong to one line, and
show.legend = FALSE hides the trend-line guide.
ggplot(mpg, aes(displ, hwy)) +
geom_point(aes(color = class), alpha = 0.55) +
geom_smooth(
aes(group = drv),
method = "lm", formula = y ~ x,
se = FALSE, color = "grey25", linewidth = 0.8,
show.legend = FALSE
) +
labs(
title = "Each layer can have its own mapping",
subtitle = "Points use vehicle class; trends are grouped by drive type",
color = "Vehicle class",
x = "Engine displacement (L)",
y = "Highway fuel economy (mpg)"
)Figure 3: Local mappings let layers use different grouping rules.
When a layer should not inherit global mappings, use
inherit.aes = FALSE and fully specify that layer’s own
data and aes().
Discrete or repeated values often overlap. Common remedies are
smaller points, transparency, or geom_jitter() to add a
small visual displacement.
ggplot(mpg, aes(factor(cyl), cty, color = factor(cyl))) +
geom_jitter(width = 0.16, height = 0, alpha = 0.5, size = 1.8,
show.legend = FALSE) +
labs(
title = "Use jitter to reduce overlap on a discrete x-axis",
x = "Number of cylinders",
y = "City fuel economy (mpg)"
)Figure 4: Jitter reveals discrete observations that would otherwise overlap.
geom_jitter() helps display individual observations, but
the displacement is only a display technique and must not be interpreted
as measurement error. For large data, also consider two-dimensional
bins, contours, or sampling.
The default method used by geom_smooth() depends on data
size. For a clear analytical meaning, specify the method explicitly:
method = "lm": a linear model;method = "loess": a local smooth suitable for smaller
data;se = TRUE/FALSE: show or hide the confidence band;formula = y ~ x: make the model formula explicit.ggplot(mpg, aes(displ, hwy, color = drv)) +
geom_point(alpha = 0.45) +
geom_smooth(method = "lm", formula = y ~ x, se = TRUE, linewidth = 0.9) +
scale_color_brewer(palette = "Dark2") +
labs(
title = "Fit a linear trend within each group",
subtitle = "The color mapping also groups points and trend lines by drv",
color = "Drive type",
x = "Engine displacement (L)",
y = "Highway fuel economy (mpg)"
)Figure 5: A scatterplot with grouped linear regression lines and confidence bands.
A smooth describes association, not automatic causation. Extrapolating beyond the observed x-range is especially risky.
geom_line() sorts by x before joining observations and
is appropriate for time series. geom_path() joins rows in
their existing order and is useful for trajectories. Multiple lines need
the correct group; mapping color or
linetype usually establishes groups automatically.
unemp_median <- median(economics$unemploy, na.rm = TRUE)
ggplot(economics, aes(date, unemploy)) +
geom_line(color = "#2c7fb8", linewidth = 0.65) +
geom_hline(yintercept = unemp_median, linetype = "dashed", color = "#d95f0e") +
scale_x_date(date_breaks = "8 years", date_labels = "%Y") +
scale_y_continuous(labels = scales::label_number(big.mark = ",")) +
labs(
title = "U.S. unemployment over time",
subtitle = "The dashed line is the median over the full period",
x = NULL,
y = "Unemployed people (thousands)",
caption = "Data: ggplot2::economics"
)Figure 6: A time series with a reference line.
A time variable should be a Date or date-time object
rather than an ordinary string. For grouped time series, also verify
that observations are sorted by time within every group.
geom_bar() versus geom_col()These functions look similar but express different data semantics:
geom_bar(): supply x or y only; the default
stat = "count" counts rows in each category;geom_col(): supply both a category and a value; it
draws values that have already been computed.ggplot(mpg, aes(y = class, fill = class)) +
geom_bar(show.legend = FALSE, width = 0.72) +
scale_fill_brewer(palette = "Blues") +
labs(
title = "geom_bar() counts raw rows",
x = "Number of vehicles",
y = NULL
)Figure 7: geom_bar() automatically counts observations in each vehicle class.
class_mean <- aggregate(hwy ~ class, data = mpg, FUN = mean)
class_mean$class <- reorder(class_mean$class, class_mean$hwy)
ggplot(class_mean, aes(hwy, class, fill = hwy)) +
geom_col(width = 0.72, show.legend = FALSE) +
scale_fill_gradient(low = "#bdd7e7", high = "#08519c") +
labs(
title = "geom_col() draws values already computed",
x = "Mean highway fuel economy (mpg)",
y = NULL
)Figure 8: geom_col() draws a precomputed average for each category.
position controls how groups at the same category are
arranged:
| position | Effect | Best question |
|---|---|---|
"stack" |
Stacks groups; the default for bars | What are the total and composition? |
"dodge" |
Places groups side by side | How do absolute group counts compare? |
"fill" |
Normalizes every bar to 100% | How do proportions compare? |
"identity" |
Keeps original positions and overlaps | Are coordinates precomputed or is alpha used? |
mpg_three <- subset(mpg, class %in% c("compact", "midsize", "suv"))
ggplot(mpg_three, aes(class, fill = drv)) +
geom_bar(position = "fill", width = 0.72) +
scale_y_continuous(labels = scales::label_percent()) +
scale_fill_brewer(palette = "Set2") +
labs(
title = "A 100% stacked bar compares composition, not counts",
x = NULL,
y = "Share of vehicles",
fill = "Drive type"
)Figure 9: position = ‘fill’ compares drive-type composition within vehicle classes.
Percentage bars hide sample-size differences. If group totals matter,
label n or also show a count chart.
A histogram’s message can change with bin width. bins
specifies the number of bins; binwidth specifies each bin’s
width. Try several defensible values based on measurement precision and
sample size. Below, the histogram is put on a density scale before a
kernel-density curve is overlaid.
ggplot(mpg, aes(hwy)) +
geom_histogram(
aes(y = after_stat(density)),
binwidth = 2, boundary = 0,
fill = "#9ecae1", color = "white"
) +
geom_density(color = "#d95f0e", linewidth = 1.05, adjust = 1) +
labs(
title = "Choose a bin width before interpreting a distribution",
subtitle = "Histogram binwidth = 2; the curve is a kernel-density estimate",
x = "Highway fuel economy (mpg)",
y = "Density"
)Figure 10: A density histogram and kernel-density curve for highway fuel economy.
after_stat(density) refers to a variable created by the
statistical transformation. It is the modern replacement for the old
..density.. notation.
Boxplots compactly show the median, interquartile range, and outliers; violin plots show distribution shape. When combining them, keep either layer from obscuring the other.
mpg_four <- subset(mpg, class %in% c("compact", "midsize", "pickup", "suv"))
ggplot(mpg_four, aes(class, hwy, fill = class)) +
geom_violin(trim = FALSE, alpha = 0.65, color = NA, show.legend = FALSE) +
geom_boxplot(width = 0.14, outlier.shape = NA, fill = "white",
linewidth = 0.45) +
scale_fill_brewer(palette = "Set2") +
labs(
title = "Distribution shape and robust summaries complement each other",
x = NULL,
y = "Highway fuel economy (mpg)"
)Figure 11: Violins show shape while narrow boxplots show robust summaries.
Violin shapes can be unstable for small groups. Add jittered observations and report group sample sizes when needed. A boxplot outlier is not automatically a data error.
stat_summary(): summarize while plottingstat_summary() can calculate a center and interval by x
group. The point below is a mean and the range is “mean ± one standard
deviation.” This is not a confidence interval, so the caption must
define it accurately.
mean_minus_sd <- function(z) mean(z, na.rm = TRUE) - sd(z, na.rm = TRUE)
mean_plus_sd <- function(z) mean(z, na.rm = TRUE) + sd(z, na.rm = TRUE)
ggplot(mpg, aes(factor(cyl), cty)) +
stat_summary(
fun = mean,
fun.min = mean_minus_sd,
fun.max = mean_plus_sd,
geom = "pointrange",
color = "#2c7fb8",
linewidth = 0.8
) +
labs(
title = "stat_summary() performs grouped summaries while plotting",
subtitle = "Point = mean; range = mean ± 1 standard deviation",
x = "Number of cylinders",
y = "City fuel economy (mpg)"
)Figure 12: Mean city fuel economy and one-standard-deviation ranges by cylinder count.
This is convenient during exploration. For a complex formal analysis,
build an explicit summary data frame first, then draw it with
geom_col(), geom_point(), or
geom_errorbar() so the calculations are easy to audit.
facet_wrap() and facet_grid()facet_wrap(~ var): one faceting variable, wrapped into
rows and columns;facet_grid(rows ~ cols): row and column variables have
distinct meaning;scales = "free_y", and similar options: allow
panel-specific scales, but weaken absolute comparisons across
panels.ggplot(mpg, aes(displ, hwy)) +
geom_point(aes(color = class), alpha = 0.6, show.legend = FALSE) +
geom_smooth(method = "lm", formula = y ~ x, se = FALSE,
color = "grey25", linewidth = 0.7) +
facet_wrap(~ drv, nrow = 1) +
labs(
title = "facet_wrap() keeps common scales for direct comparison",
x = "Engine displacement (L)",
y = "Highway fuel economy (mpg)"
)Figure 13: The displacement–economy relationship faceted by drive type.
facet_data <- subset(mpg, cyl %in% c(4, 6, 8))
ggplot(facet_data, aes(class, fill = class)) +
geom_bar(show.legend = FALSE) +
facet_grid(drv ~ cyl, scales = "free_x", space = "free_x") +
labs(
title = "facet_grid() is useful when rows and columns have meaning",
subtitle = "Rows = drive type; columns = cylinder count",
x = NULL,
y = "Number of vehicles"
) +
theme(axis.text.x = element_text(angle = 45, hjust = 1))Figure 14: Row and column facets encode drive type and cylinder count.
An empty panel can be important information: it says a combination does not occur. Do not remove empty combinations merely for appearance.
Scale functions generally follow
scale_<aesthetic>_<type>(), for example:
scale_x_continuous() and scale_y_log10()
for position;scale_color_brewer() and
scale_fill_viridis_d() for discrete color;scale_color_viridis_c() and
scale_fill_gradient() for continuous color;scale_color_manual() for manually assigned categorical
colors;scale_x_date() for a date axis.color usually controls points, lines, and outlines.
fill controls the interior of bars, areas, and closed
shapes.
diamond_sample <- diamonds[sample.int(nrow(diamonds), 3000), ]
ggplot(diamond_sample, aes(carat, price, color = cut)) +
geom_point(alpha = 0.42, size = 1.2) +
scale_x_log10(
breaks = c(0.3, 0.5, 1, 2, 3),
labels = scales::label_number()
) +
scale_y_log10(labels = scales::label_dollar()) +
scale_color_viridis_d(option = "C", end = 0.9) +
labs(
title = "Log scales reveal multiplicative relationships",
subtitle = "The axes are transformed; the source data are unchanged",
x = "Weight (carats, log scale)",
y = "Price (USD, log scale)",
color = "Cut"
)Figure 15: Diamond weight and price on logarithmic scales.
A log scale cannot display zero or negative values, and it changes the meaning of visual distance. State the transformation in an axis label or caption.
Use a named vector for manual colors so categories are matched by name rather than by an accidental factor order.
drive_colors <- c("4" = "#1b9e77", "f" = "#d95f02", "r" = "#7570b3")
ggplot(mpg, aes(displ, hwy, color = drv)) +
geom_point(alpha = 0.65, size = 2) +
scale_color_manual(
values = drive_colors,
breaks = c("f", "4", "r"),
labels = c("Front-wheel drive", "Four-wheel drive", "Rear-wheel drive")
) +
labs(
title = "A manual scale controls color, order, and labels",
x = "Engine displacement (L)",
y = "Highway fuel economy (mpg)",
color = "Drive type"
)Figure 16: A named vector assigns categorical colors reliably.
Category order affects axes, stacking, and legends. Two useful base R patterns are:
limits versus visual zoomingThis distinction can change a statistical result:
# Deletes out-of-range observations before statistical computation;
# smooths, boxplots, and other summaries may change
p + scale_y_continuous(limits = c(20, 40))
# Changes only the viewing window and keeps all observations
p + coord_cartesian(ylim = c(20, 40))Prefer coord_cartesian() when you only want a closer
view. If filtering is analytically intended, filter explicitly before
plotting so the decision is visible.
coord_cartesian(): zoom without deleting
observations;coord_fixed(ratio = 1): fix the physical ratio of x and
y units;coord_flip(): swap axes in older code; current ggplot2
usually favors swapping x and y directly in aes();coord_polar(): polar coordinates for pie or ring
charts, although angles and areas are hard to compare accurately.Coordinates are applied after scales. When exact comparisons matter, position and length are usually easier to read than angle or area.
labs() manages titles, subtitles, axes, guide titles,
and captions. annotate() adds one-off explanations, while
geom_text() and geom_label() add labels from a
data frame.
hwy_median <- median(mpg$hwy)
ggplot(mpg, aes(displ, hwy)) +
geom_point(color = "#2c7fb8", alpha = 0.55) +
geom_hline(yintercept = hwy_median, linetype = "dashed",
color = "#d95f0e", linewidth = 0.8) +
annotate(
"label", x = 6.2, y = hwy_median + 1.5,
label = paste0("Median = ", hwy_median),
hjust = 1, size = 3.5,
fill = "#fff5eb", color = "#a63603"
) +
annotate(
"curve", x = 3.7, y = 42, xend = 2.4, yend = 43,
curvature = 0.2, arrow = arrow(length = grid::unit(0.12, "inches")),
color = "grey35"
) +
annotate("text", x = 3.8, y = 42, label = "High-efficiency vehicles",
hjust = 0, color = "grey25") +
labs(
title = "Annotations should help the reader see the conclusion",
subtitle = "Avoid labeling every point and creating clutter",
x = "Engine displacement (L)",
y = "Highway fuel economy (mpg)",
caption = "Dashed line: sample median"
) +
coord_cartesian(clip = "off")Figure 17: Reference lines and annotations emphasize a key region.
grid::unit() and arrow() come with R; no
extra plotting extension is required.
A theme controls non-data elements and does not
alter data mappings. Common complete themes include
theme_minimal(), theme_classic(), and
theme_bw(). Use theme() for targeted
changes:
element_text() for text;element_line() for grid and axis lines;element_rect() for backgrounds and borders;element_blank() to remove an element.The setup chunk defines theme_guide(). Wrapping a
project’s visual system in a function is more reliable than copying a
long theme() call onto every plot.
ggplot(mpg, aes(displ, hwy, color = class)) +
geom_point(alpha = 0.65, size = 2) +
scale_color_brewer(palette = "Dark2") +
labs(
title = "Themes control layout; scales control mappings",
subtitle = "The legend is below the plot and minor grid lines are removed",
x = "Engine displacement (L)",
y = "Highway fuel economy (mpg)",
color = "Vehicle class"
) +
guides(color = guide_legend(nrow = 2, byrow = TRUE,
override.aes = list(alpha = 1, size = 3))) +
theme_guide(base_size = 13)Figure 18: A reusable theme unifies type, grid, title, and legend placement.
When multiple layers map to the same variable and use the same scale
name, ggplot2 usually merges their guides. Use guides() to
adjust order, rows, or example marks, and
theme(legend.position = "none") to remove an unnecessary
legend.
When both axes are categorical and a cell represents a count or
value, geom_tile() is a useful choice.
heat <- as.data.frame(table(cyl = mpg$cyl, drv = mpg$drv))
ggplot(heat, aes(cyl, drv, fill = Freq)) +
geom_tile(color = "white", linewidth = 0.8) +
geom_text(aes(label = Freq), color = "white", fontface = "bold") +
scale_fill_viridis_c(option = "B", begin = 0.2, end = 0.9) +
labs(
title = "geom_tile() maps two-dimensional combinations to color",
x = "Number of cylinders",
y = "Drive type",
fill = "Vehicles"
) +
coord_fixed()Figure 19: Heatmap of vehicle counts by cylinder count and drive type.
A continuous palette should have a perceptually consistent progression in lightness. Viridis palettes are generally friendly to common color-vision differences and grayscale printing. If labels disappear on light cells, choose text color conditionally from the fill value.
Do not use the deprecated aes_string(). When a function
receives column names as strings, use .data[[...]]:
make_scatter <- function(data, x, y, color = NULL) {
if (is.null(color)) {
mapping <- aes(x = .data[[x]], y = .data[[y]])
} else {
mapping <- aes(
x = .data[[x]],
y = .data[[y]],
color = .data[[color]]
)
}
ggplot(data, mapping) +
geom_point(alpha = 0.65, size = 2) +
labs(x = x, y = y, color = color) +
theme_guide()
}
make_scatter(mpg, "displ", "hwy", "drv")The function returns a ggplot object, so callers can keep adding layers:
ggsave() saves the last displayed plot by default.
Passing plot explicitly is safer in scripts and functions.
Width, height, and dpi jointly determine raster image
dimensions.
p <- ggplot(mpg, aes(displ, hwy, color = drv)) +
geom_point(alpha = 0.7) +
labs(x = "Engine displacement (L)", y = "Highway MPG") +
theme_guide()
# PNG for documents or the web: 7 by 4.5 inches at 300 dpi
ggsave(
filename = "mpg_scatter.png",
plot = p,
width = 7, height = 4.5, units = "in",
dpi = 300, bg = "white"
)
# Vector PDF for papers and downstream layout
ggsave("mpg_scatter.pdf", p, width = 7, height = 4.5, units = "in")Practical recommendations:
width, height, units, and
bg explicitly for consistent output;| Symptom | Likely cause | Fix |
|---|---|---|
| A strange legend contains a color name | A constant was put inside aes() |
Put color = "red" outside
aes() |
| Bars are gray or their interiors do not change | color and fill were
confused |
Bar and area interiors usually use
fill |
geom_bar() reports a y-related error |
Count bars were used for summarized values | Use geom_col() |
| Separate lines are joined together | Grouping is missing | Map group, color, or
linetype |
| A smooth or boxplot changes after zooming | scale_* (limits=...) removed data |
Use coord_cartesian() for visual
zooming |
| Points form a dark blob | Overplotting | Use smaller points, alpha, jitter, bins, or sampling |
| A histogram conclusion is unstable | Default bins were accepted blindly | Set and compare domain-relevant binwidth
values |
| Category order is illogical | Character/factor defaults are being used | Set factor(levels=...) or use
reorder() |
| Legends repeat or refuse to merge | Layers use different variables or scale names | Align mapping variables and
labs(color/fill=...) |
| “Removed … rows” warning | Missing values, out-of-range values, or scale limits | Inspect is.na(), ranges, and
limits |
Do not add na.rm = TRUE merely to silence a warning.
First determine why values are missing, whether deletion is appropriate,
and whether deletion changes the conclusion.
| Analytical goal | First-choice layer | Common additions |
|---|---|---|
| Relationship between two continuous variables | geom_point() |
alpha, geom_smooth() |
| Continuous value over time | geom_line() |
geom_point(), reference lines |
| Category frequencies | geom_bar() |
position, horizontal y mapping |
| Pre-summarized category values | geom_col() |
Ordering, error bars |
| One-variable distribution | geom_histogram() /
geom_density() |
Explicit binwidth /
adjust |
| Comparing grouped distributions | geom_boxplot() /
geom_violin() |
Jittered points, sample sizes |
| Values for two categorical dimensions | geom_tile() |
Continuous fill scale, labels |
| Repeated comparisons across groups | facet_wrap() /
facet_grid() |
Shared scales |
ggsave(),
and retain the plotting code and session information.This document uses modern ggplot2 syntax: linewidth for
line width, after_stat() for computed variables, and
.data[[...]] in programmatic mappings.
cat("R:", R.version.string, "\n")
#> R: R version 4.6.1 (2026-06-24)
cat("ggplot2:", as.character(packageVersion("ggplot2")), "\n")
#> ggplot2: 4.0.3
cat("knitr:", as.character(packageVersion("knitr")), "\n")
#> knitr: 1.51
cat("rmarkdown:", as.character(packageVersion("rmarkdown")), "\n")
#> rmarkdown: 2.31You have now covered ggplot2’s most common grammar, geoms, statistical transformations, positions, facets, scales, coordinates, themes, annotations, programming interface, and export workflow. In real projects, make the data meaning and comparison task clear before adding visual decoration.