The V Lab
AudienceReaders who can run basic R code
Scope15 sections · 20 rendered charts
FormatFollow in sequence or browse by topic

1 How to use this guide

This guide is for readers who can run R code and want a systematic introduction to the ggplot2 methods used most often in real work. Every example uses a base R or ggplot2 built-in dataset, so no external data download is required. You can copy each code chunk directly into RStudio.

Think about a plot in this order:

  1. Data: what does one row represent, and are the variable types correct?
  2. Mapping: which variables map to x, y, color, fill, size, or group?
  3. Geometric object (geom): should observations be shown as points, lines, bars, boxes, or densities?
  4. Statistical transformation (stat): show raw values, counts, bins, a smooth, or a summary?
  5. Position and coordinates: stack, dodge, jitter, or zoom?
  6. Scale: how should colors, breaks, labels, and transformations appear?
  7. Facets and theme: how should small multiples and non-data styling be organized?

The core mental model is ggplot(data, aes(…)) + geom_*() + scale_*() + facet_*() + coord_*() + theme_*(). Each + adds or modifies a layer.

1.1 Datasets used in this guide

Dataset Contents Typical uses
mpg Fuel economy and specifications for 1999 and 2008 models Scatterplots, bars, facets
diamonds Specifications and prices for about 54,000 diamonds Large data, distributions, log scales
economics Monthly U.S. economic indicators Dates and time series
mtcars Performance data for 32 cars Small demonstrations and modeling
dim(mpg)
#> [1] 234  11
names(mpg)
#>  [1] "manufacturer" "model"        "displ"        "year"         "cyl"          "trans"       
#>  [7] "drv"          "cty"          "hwy"          "fl"           "class"
head(mpg, 3)

2 Grammar of graphics fundamentals

2.1 The smallest useful plot

ggplot() creates the canvas, aes() declares mappings from variables to visual properties, and geom_point() says to draw one point for each row.

ggplot(data = mpg, mapping = aes(x = displ, y = hwy)) +
  geom_point() +
  labs(
    title = "Larger engines usually have lower highway fuel economy",
    x = "Engine displacement (L)",
    y = "Highway fuel economy (mpg)"
  )
Scatterplot showing that vehicles with larger engine displacement generally have lower highway miles per gallon.

Figure 1: Engine displacement versus highway fuel economy.

Data and mappings can also be supplied to individual layers. A global aes() is inherited by later layers; a local aes() affects only its own layer.

2.2 Mapping versus setting

  • Inside aes(): map a variable to a visual property. This creates data-driven variation and usually a guide.
  • Outside aes(): set one constant value for every mark. This usually creates no guide.
ggplot(mpg, aes(displ, hwy)) +
  geom_point(aes(color = class), alpha = 0.72, size = 2.2) +
  labs(
    title = "Put variables inside aes() and constants outside",
    color = "Vehicle class",
    x = "Engine displacement (L)",
    y = "Highway fuel economy (mpg)"
  ) +
  guides(color = guide_legend(nrow = 2, byrow = TRUE,
                              override.aes = list(alpha = 1, size = 3)))
Displacement versus highway fuel economy scatterplot colored by vehicle class.

Figure 2: Color is mapped to vehicle class; alpha and point size are fixed.

One of the most common mistakes is aes(color = "steelblue"). It treats the string as a category and creates a legend rather than setting points to blue.

# Wrong: treats "steelblue" as a data category
ggplot(mpg, aes(displ, hwy)) + geom_point(aes(color = "steelblue"))

# Correct: fixes every point to the same blue
ggplot(mpg, aes(displ, hwy)) + geom_point(color = "steelblue")

2.3 Global mappings, local mappings, and grouping

Here, point color is mapped to vehicle class while trend lines are grouped only by drive type (drv). group controls which observations belong to one line, and show.legend = FALSE hides the trend-line guide.

ggplot(mpg, aes(displ, hwy)) +
  geom_point(aes(color = class), alpha = 0.55) +
  geom_smooth(
    aes(group = drv),
    method = "lm", formula = y ~ x,
    se = FALSE, color = "grey25", linewidth = 0.8,
    show.legend = FALSE
  ) +
  labs(
    title = "Each layer can have its own mapping",
    subtitle = "Points use vehicle class; trends are grouped by drive type",
    color = "Vehicle class",
    x = "Engine displacement (L)",
    y = "Highway fuel economy (mpg)"
  )
Scatterplot colored by vehicle class with separate linear trend lines for drive types.

Figure 3: Local mappings let layers use different grouping rules.

When a layer should not inherit global mappings, use inherit.aes = FALSE and fully specify that layer’s own data and aes().

3 Common geoms

3.1 Scatterplots: overplotting, transparency, and jitter

Discrete or repeated values often overlap. Common remedies are smaller points, transparency, or geom_jitter() to add a small visual displacement.

ggplot(mpg, aes(factor(cyl), cty, color = factor(cyl))) +
  geom_jitter(width = 0.16, height = 0, alpha = 0.5, size = 1.8,
              show.legend = FALSE) +
  labs(
    title = "Use jitter to reduce overlap on a discrete x-axis",
    x = "Number of cylinders",
    y = "City fuel economy (mpg)"
  )
Jittered scatterplot of city fuel economy by cylinder count.

Figure 4: Jitter reveals discrete observations that would otherwise overlap.

geom_jitter() helps display individual observations, but the displacement is only a display technique and must not be interpreted as measurement error. For large data, also consider two-dimensional bins, contours, or sampling.

3.2 Smooths and regression lines

The default method used by geom_smooth() depends on data size. For a clear analytical meaning, specify the method explicitly:

  • method = "lm": a linear model;
  • method = "loess": a local smooth suitable for smaller data;
  • se = TRUE/FALSE: show or hide the confidence band;
  • formula = y ~ x: make the model formula explicit.
ggplot(mpg, aes(displ, hwy, color = drv)) +
  geom_point(alpha = 0.45) +
  geom_smooth(method = "lm", formula = y ~ x, se = TRUE, linewidth = 0.9) +
  scale_color_brewer(palette = "Dark2") +
  labs(
    title = "Fit a linear trend within each group",
    subtitle = "The color mapping also groups points and trend lines by drv",
    color = "Drive type",
    x = "Engine displacement (L)",
    y = "Highway fuel economy (mpg)"
  )
Displacement versus highway fuel economy colored by drive type with a separate fitted linear regression for each group.

Figure 5: A scatterplot with grouped linear regression lines and confidence bands.

A smooth describes association, not automatic causation. Extrapolating beyond the observed x-range is especially risky.

3.3 Lines and time series

geom_line() sorts by x before joining observations and is appropriate for time series. geom_path() joins rows in their existing order and is useful for trajectories. Multiple lines need the correct group; mapping color or linetype usually establishes groups automatically.

unemp_median <- median(economics$unemploy, na.rm = TRUE)

ggplot(economics, aes(date, unemploy)) +
  geom_line(color = "#2c7fb8", linewidth = 0.65) +
  geom_hline(yintercept = unemp_median, linetype = "dashed", color = "#d95f0e") +
  scale_x_date(date_breaks = "8 years", date_labels = "%Y") +
  scale_y_continuous(labels = scales::label_number(big.mark = ",")) +
  labs(
    title = "U.S. unemployment over time",
    subtitle = "The dashed line is the median over the full period",
    x = NULL,
    y = "Unemployed people (thousands)",
    caption = "Data: ggplot2::economics"
  )
Monthly U.S. unemployment count from 1967 to 2015 with a dashed median reference line.

Figure 6: A time series with a reference line.

A time variable should be a Date or date-time object rather than an ordinary string. For grouped time series, also verify that observations are sorted by time within every group.

3.4 Bars: geom_bar() versus geom_col()

These functions look similar but express different data semantics:

  • geom_bar(): supply x or y only; the default stat = "count" counts rows in each category;
  • geom_col(): supply both a category and a value; it draws values that have already been computed.
ggplot(mpg, aes(y = class, fill = class)) +
  geom_bar(show.legend = FALSE, width = 0.72) +
  scale_fill_brewer(palette = "Blues") +
  labs(
    title = "geom_bar() counts raw rows",
    x = "Number of vehicles",
    y = NULL
  )
Horizontal bar chart of vehicle-class frequencies.

Figure 7: geom_bar() automatically counts observations in each vehicle class.

class_mean <- aggregate(hwy ~ class, data = mpg, FUN = mean)
class_mean$class <- reorder(class_mean$class, class_mean$hwy)

ggplot(class_mean, aes(hwy, class, fill = hwy)) +
  geom_col(width = 0.72, show.legend = FALSE) +
  scale_fill_gradient(low = "#bdd7e7", high = "#08519c") +
  labs(
    title = "geom_col() draws values already computed",
    x = "Mean highway fuel economy (mpg)",
    y = NULL
  )
Horizontal bars showing mean highway fuel economy by vehicle class, ordered from low to high.

Figure 8: geom_col() draws a precomputed average for each category.

3.5 Stack, dodge, and fill

position controls how groups at the same category are arranged:

position Effect Best question
"stack" Stacks groups; the default for bars What are the total and composition?
"dodge" Places groups side by side How do absolute group counts compare?
"fill" Normalizes every bar to 100% How do proportions compare?
"identity" Keeps original positions and overlaps Are coordinates precomputed or is alpha used?
mpg_three <- subset(mpg, class %in% c("compact", "midsize", "suv"))

ggplot(mpg_three, aes(class, fill = drv)) +
  geom_bar(position = "fill", width = 0.72) +
  scale_y_continuous(labels = scales::label_percent()) +
  scale_fill_brewer(palette = "Set2") +
  labs(
    title = "A 100% stacked bar compares composition, not counts",
    x = NULL,
    y = "Share of vehicles",
    fill = "Drive type"
  )
One-hundred-percent stacked bars showing drive-type shares for compact, midsize, and SUV vehicles.

Figure 9: position = ‘fill’ compares drive-type composition within vehicle classes.

Percentage bars hide sample-size differences. If group totals matter, label n or also show a count chart.

4 Distributions and statistical transformations

4.1 Histograms and densities

A histogram’s message can change with bin width. bins specifies the number of bins; binwidth specifies each bin’s width. Try several defensible values based on measurement precision and sample size. Below, the histogram is put on a density scale before a kernel-density curve is overlaid.

ggplot(mpg, aes(hwy)) +
  geom_histogram(
    aes(y = after_stat(density)),
    binwidth = 2, boundary = 0,
    fill = "#9ecae1", color = "white"
  ) +
  geom_density(color = "#d95f0e", linewidth = 1.05, adjust = 1) +
  labs(
    title = "Choose a bin width before interpreting a distribution",
    subtitle = "Histogram binwidth = 2; the curve is a kernel-density estimate",
    x = "Highway fuel economy (mpg)",
    y = "Density"
  )
Histogram of highway fuel economy with an orange kernel-density curve overlaid.

Figure 10: A density histogram and kernel-density curve for highway fuel economy.

after_stat(density) refers to a variable created by the statistical transformation. It is the modern replacement for the old ..density.. notation.

4.2 Boxplots and violin plots

Boxplots compactly show the median, interquartile range, and outliers; violin plots show distribution shape. When combining them, keep either layer from obscuring the other.

mpg_four <- subset(mpg, class %in% c("compact", "midsize", "pickup", "suv"))

ggplot(mpg_four, aes(class, hwy, fill = class)) +
  geom_violin(trim = FALSE, alpha = 0.65, color = NA, show.legend = FALSE) +
  geom_boxplot(width = 0.14, outlier.shape = NA, fill = "white",
               linewidth = 0.45) +
  scale_fill_brewer(palette = "Set2") +
  labs(
    title = "Distribution shape and robust summaries complement each other",
    x = NULL,
    y = "Highway fuel economy (mpg)"
  )
Violin and boxplots of highway fuel economy for four vehicle classes.

Figure 11: Violins show shape while narrow boxplots show robust summaries.

Violin shapes can be unstable for small groups. Add jittered observations and report group sample sizes when needed. A boxplot outlier is not automatically a data error.

4.3 stat_summary(): summarize while plotting

stat_summary() can calculate a center and interval by x group. The point below is a mean and the range is “mean ± one standard deviation.” This is not a confidence interval, so the caption must define it accurately.

mean_minus_sd <- function(z) mean(z, na.rm = TRUE) - sd(z, na.rm = TRUE)
mean_plus_sd  <- function(z) mean(z, na.rm = TRUE) + sd(z, na.rm = TRUE)

ggplot(mpg, aes(factor(cyl), cty)) +
  stat_summary(
    fun = mean,
    fun.min = mean_minus_sd,
    fun.max = mean_plus_sd,
    geom = "pointrange",
    color = "#2c7fb8",
    linewidth = 0.8
  ) +
  labs(
    title = "stat_summary() performs grouped summaries while plotting",
    subtitle = "Point = mean; range = mean ± 1 standard deviation",
    x = "Number of cylinders",
    y = "City fuel economy (mpg)"
  )
Point ranges showing mean city fuel economy plus or minus one standard deviation for each cylinder count.

Figure 12: Mean city fuel economy and one-standard-deviation ranges by cylinder count.

This is convenient during exploration. For a complex formal analysis, build an explicit summary data frame first, then draw it with geom_col(), geom_point(), or geom_errorbar() so the calculations are easy to audit.

5 Facets: small multiples for group comparison

5.1 facet_wrap() and facet_grid()

  • facet_wrap(~ var): one faceting variable, wrapped into rows and columns;
  • facet_grid(rows ~ cols): row and column variables have distinct meaning;
  • scales = "free_y", and similar options: allow panel-specific scales, but weaken absolute comparisons across panels.
ggplot(mpg, aes(displ, hwy)) +
  geom_point(aes(color = class), alpha = 0.6, show.legend = FALSE) +
  geom_smooth(method = "lm", formula = y ~ x, se = FALSE,
              color = "grey25", linewidth = 0.7) +
  facet_wrap(~ drv, nrow = 1) +
  labs(
    title = "facet_wrap() keeps common scales for direct comparison",
    x = "Engine displacement (L)",
    y = "Highway fuel economy (mpg)"
  )
Three panels show displacement versus highway fuel economy for front-, four-, and rear-wheel-drive vehicles.

Figure 13: The displacement–economy relationship faceted by drive type.

facet_data <- subset(mpg, cyl %in% c(4, 6, 8))

ggplot(facet_data, aes(class, fill = class)) +
  geom_bar(show.legend = FALSE) +
  facet_grid(drv ~ cyl, scales = "free_x", space = "free_x") +
  labs(
    title = "facet_grid() is useful when rows and columns have meaning",
    subtitle = "Rows = drive type; columns = cylinder count",
    x = NULL,
    y = "Number of vehicles"
  ) +
  theme(axis.text.x = element_text(angle = 45, hjust = 1))
Bar-chart grid with drive type in rows and cylinder count in columns, showing counts by vehicle class.

Figure 14: Row and column facets encode drive type and cylinder count.

An empty panel can be important information: it says a combination does not occur. Do not remove empty combinations merely for appearance.

6 Scales: controlling visual mappings

6.1 Continuous, discrete, and manual scales

Scale functions generally follow scale_<aesthetic>_<type>(), for example:

  • scale_x_continuous() and scale_y_log10() for position;
  • scale_color_brewer() and scale_fill_viridis_d() for discrete color;
  • scale_color_viridis_c() and scale_fill_gradient() for continuous color;
  • scale_color_manual() for manually assigned categorical colors;
  • scale_x_date() for a date axis.

color usually controls points, lines, and outlines. fill controls the interior of bars, areas, and closed shapes.

diamond_sample <- diamonds[sample.int(nrow(diamonds), 3000), ]

ggplot(diamond_sample, aes(carat, price, color = cut)) +
  geom_point(alpha = 0.42, size = 1.2) +
  scale_x_log10(
    breaks = c(0.3, 0.5, 1, 2, 3),
    labels = scales::label_number()
  ) +
  scale_y_log10(labels = scales::label_dollar()) +
  scale_color_viridis_d(option = "C", end = 0.9) +
  labs(
    title = "Log scales reveal multiplicative relationships",
    subtitle = "The axes are transformed; the source data are unchanged",
    x = "Weight (carats, log scale)",
    y = "Price (USD, log scale)",
    color = "Cut"
  )
Double-log scatterplot of carat and price for three thousand diamonds, colored by cut.

Figure 15: Diamond weight and price on logarithmic scales.

A log scale cannot display zero or negative values, and it changes the meaning of visual distance. State the transformation in an axis label or caption.

6.2 Manual colors and category order

Use a named vector for manual colors so categories are matched by name rather than by an accidental factor order.

drive_colors <- c("4" = "#1b9e77", "f" = "#d95f02", "r" = "#7570b3")

ggplot(mpg, aes(displ, hwy, color = drv)) +
  geom_point(alpha = 0.65, size = 2) +
  scale_color_manual(
    values = drive_colors,
    breaks = c("f", "4", "r"),
    labels = c("Front-wheel drive", "Four-wheel drive", "Rear-wheel drive")
  ) +
  labs(
    title = "A manual scale controls color, order, and labels",
    x = "Engine displacement (L)",
    y = "Highway fuel economy (mpg)",
    color = "Drive type"
  )
Scatterplot of displacement and highway fuel economy for three drive types using manually selected colors.

Figure 16: A named vector assigns categorical colors reliably.

Category order affects axes, stacking, and legends. Two useful base R patterns are:

# Set a meaningful business order
dat$level <- factor(dat$level, levels = c("Low", "Medium", "High"))

# Reorder groups by a numerical summary
dat$group <- reorder(dat$group, dat$value, FUN = median)

6.3 limits versus visual zooming

This distinction can change a statistical result:

# Deletes out-of-range observations before statistical computation;
# smooths, boxplots, and other summaries may change
p + scale_y_continuous(limits = c(20, 40))

# Changes only the viewing window and keeps all observations
p + coord_cartesian(ylim = c(20, 40))

Prefer coord_cartesian() when you only want a closer view. If filtering is analytically intended, filter explicitly before plotting so the decision is visible.

7 Coordinate systems, labels, and annotations

7.1 Coordinate systems

  • coord_cartesian(): zoom without deleting observations;
  • coord_fixed(ratio = 1): fix the physical ratio of x and y units;
  • coord_flip(): swap axes in older code; current ggplot2 usually favors swapping x and y directly in aes();
  • coord_polar(): polar coordinates for pie or ring charts, although angles and areas are hard to compare accurately.

Coordinates are applied after scales. When exact comparisons matter, position and length are usually easier to read than angle or area.

7.2 Titles, labels, reference lines, and annotations

labs() manages titles, subtitles, axes, guide titles, and captions. annotate() adds one-off explanations, while geom_text() and geom_label() add labels from a data frame.

hwy_median <- median(mpg$hwy)

ggplot(mpg, aes(displ, hwy)) +
  geom_point(color = "#2c7fb8", alpha = 0.55) +
  geom_hline(yintercept = hwy_median, linetype = "dashed",
             color = "#d95f0e", linewidth = 0.8) +
  annotate(
    "label", x = 6.2, y = hwy_median + 1.5,
    label = paste0("Median = ", hwy_median),
    hjust = 1, size = 3.5,
    fill = "#fff5eb", color = "#a63603"
  ) +
  annotate(
    "curve", x = 3.7, y = 42, xend = 2.4, yend = 43,
    curvature = 0.2, arrow = arrow(length = grid::unit(0.12, "inches")),
    color = "grey35"
  ) +
  annotate("text", x = 3.8, y = 42, label = "High-efficiency vehicles",
           hjust = 0, color = "grey25") +
  labs(
    title = "Annotations should help the reader see the conclusion",
    subtitle = "Avoid labeling every point and creating clutter",
    x = "Engine displacement (L)",
    y = "Highway fuel economy (mpg)",
    caption = "Dashed line: sample median"
  ) +
  coord_cartesian(clip = "off")
Displacement versus highway fuel economy scatterplot with a dashed median line and a label for high-efficiency vehicles.

Figure 17: Reference lines and annotations emphasize a key region.

grid::unit() and arrow() come with R; no extra plotting extension is required.

8 Themes, legends, and reusable styling

8.1 What a theme controls

A theme controls non-data elements and does not alter data mappings. Common complete themes include theme_minimal(), theme_classic(), and theme_bw(). Use theme() for targeted changes:

  • element_text() for text;
  • element_line() for grid and axis lines;
  • element_rect() for backgrounds and borders;
  • element_blank() to remove an element.

The setup chunk defines theme_guide(). Wrapping a project’s visual system in a function is more reliable than copying a long theme() call onto every plot.

ggplot(mpg, aes(displ, hwy, color = class)) +
  geom_point(alpha = 0.65, size = 2) +
  scale_color_brewer(palette = "Dark2") +
  labs(
    title = "Themes control layout; scales control mappings",
    subtitle = "The legend is below the plot and minor grid lines are removed",
    x = "Engine displacement (L)",
    y = "Highway fuel economy (mpg)",
    color = "Vehicle class"
  ) +
  guides(color = guide_legend(nrow = 2, byrow = TRUE,
                              override.aes = list(alpha = 1, size = 3))) +
  theme_guide(base_size = 13)
Displacement versus highway fuel economy scatterplot colored by vehicle class and styled with a reusable custom theme.

Figure 18: A reusable theme unifies type, grid, title, and legend placement.

When multiple layers map to the same variable and use the same scale name, ggplot2 usually merges their guides. Use guides() to adjust order, rows, or example marks, and theme(legend.position = "none") to remove an unnecessary legend.

9 Two-dimensional summaries: heatmaps

When both axes are categorical and a cell represents a count or value, geom_tile() is a useful choice.

heat <- as.data.frame(table(cyl = mpg$cyl, drv = mpg$drv))

ggplot(heat, aes(cyl, drv, fill = Freq)) +
  geom_tile(color = "white", linewidth = 0.8) +
  geom_text(aes(label = Freq), color = "white", fontface = "bold") +
  scale_fill_viridis_c(option = "B", begin = 0.2, end = 0.9) +
  labs(
    title = "geom_tile() maps two-dimensional combinations to color",
    x = "Number of cylinders",
    y = "Drive type",
    fill = "Vehicles"
  ) +
  coord_fixed()
Heatmap of vehicle counts for combinations of cylinder count and drive type, with values printed inside cells.

Figure 19: Heatmap of vehicle counts by cylinder count and drive type.

A continuous palette should have a perceptually consistent progression in lightness. Viridis palettes are generally friendly to common color-vision differences and grayscale printing. If labels disappear on light cells, choose text color conditionally from the fill value.

10 Using ggplot2 inside functions

Do not use the deprecated aes_string(). When a function receives column names as strings, use .data[[...]]:

make_scatter <- function(data, x, y, color = NULL) {
  if (is.null(color)) {
    mapping <- aes(x = .data[[x]], y = .data[[y]])
  } else {
    mapping <- aes(
      x = .data[[x]],
      y = .data[[y]],
      color = .data[[color]]
    )
  }

  ggplot(data, mapping) +
    geom_point(alpha = 0.65, size = 2) +
    labs(x = x, y = y, color = color) +
    theme_guide()
}

make_scatter(mpg, "displ", "hwy", "drv")

The function returns a ggplot object, so callers can keep adding layers:

make_scatter(mpg, "displ", "hwy", "drv") +
  geom_smooth(method = "lm", formula = y ~ x) +
  labs(title = "A plot returned by a function remains extensible")

11 Exporting high-quality figures

ggsave() saves the last displayed plot by default. Passing plot explicitly is safer in scripts and functions. Width, height, and dpi jointly determine raster image dimensions.

p <- ggplot(mpg, aes(displ, hwy, color = drv)) +
  geom_point(alpha = 0.7) +
  labs(x = "Engine displacement (L)", y = "Highway MPG") +
  theme_guide()

# PNG for documents or the web: 7 by 4.5 inches at 300 dpi
ggsave(
  filename = "mpg_scatter.png",
  plot = p,
  width = 7, height = 4.5, units = "in",
  dpi = 300, bg = "white"
)

# Vector PDF for papers and downstream layout
ggsave("mpg_scatter.pdf", p, width = 7, height = 4.5, units = "in")

Practical recommendations:

  • Web: PNG at 120–180 dpi is often enough; use 300 dpi for high-quality print;
  • Mostly lines and text: use vector PDF or SVG;
  • Many transparent points: PNG is often much smaller than a vector file;
  • Set width, height, units, and bg explicitly for consistent output;
  • Check type sizes at the final display dimensions, not only in a zoomed RStudio preview.

12 Common problems and a debugging checklist

Symptom Likely cause Fix
A strange legend contains a color name A constant was put inside aes() Put color = "red" outside aes()
Bars are gray or their interiors do not change color and fill were confused Bar and area interiors usually use fill
geom_bar() reports a y-related error Count bars were used for summarized values Use geom_col()
Separate lines are joined together Grouping is missing Map group, color, or linetype
A smooth or boxplot changes after zooming scale_* (limits=...) removed data Use coord_cartesian() for visual zooming
Points form a dark blob Overplotting Use smaller points, alpha, jitter, bins, or sampling
A histogram conclusion is unstable Default bins were accepted blindly Set and compare domain-relevant binwidth values
Category order is illogical Character/factor defaults are being used Set factor(levels=...) or use reorder()
Legends repeat or refuse to merge Layers use different variables or scale names Align mapping variables and labs(color/fill=...)
“Removed … rows” warning Missing values, out-of-range values, or scale limits Inspect is.na(), ranges, and limits

Do not add na.rm = TRUE merely to silence a warning. First determine why values are missing, whether deletion is appropriate, and whether deletion changes the conclusion.

13 Quick selection guide

Analytical goal First-choice layer Common additions
Relationship between two continuous variables geom_point() alpha, geom_smooth()
Continuous value over time geom_line() geom_point(), reference lines
Category frequencies geom_bar() position, horizontal y mapping
Pre-summarized category values geom_col() Ordering, error bars
One-variable distribution geom_histogram() / geom_density() Explicit binwidth / adjust
Comparing grouped distributions geom_boxplot() / geom_violin() Jittered points, sample sizes
Values for two categorical dimensions geom_tile() Continuous fill scale, labels
Repeated comparisons across groups facet_wrap() / facet_grid() Shared scales

14 A reliable plotting workflow

  1. Confirm the unit of analysis, missingness, variable types, and reasonable ranges.
  2. Draw raw data with the simplest geom; inspect outliers and overlap first.
  3. Add grouping, statistical summaries, or facets; avoid mapping too many aesthetics at once.
  4. Choose scales, category order, axis range, and color palette deliberately.
  5. Write a conclusion-oriented title, complete axis labels and units, a guide title, and data source.
  6. Check readability, color-vision accessibility, and grayscale output at the final size.
  7. Export with fixed dimensions and format using ggsave(), and retain the plotting code and session information.

15 Version and reproducibility information

This document uses modern ggplot2 syntax: linewidth for line width, after_stat() for computed variables, and .data[[...]] in programmatic mappings.

cat("R:", R.version.string, "\n")
#> R: R version 4.6.1 (2026-06-24)
cat("ggplot2:", as.character(packageVersion("ggplot2")), "\n")
#> ggplot2: 4.0.3
cat("knitr:", as.character(packageVersion("knitr")), "\n")
#> knitr: 1.51
cat("rmarkdown:", as.character(packageVersion("rmarkdown")), "\n")
#> rmarkdown: 2.31

You have now covered ggplot2’s most common grammar, geoms, statistical transformations, positions, facets, scales, coordinates, themes, annotations, programming interface, and export workflow. In real projects, make the data meaning and comparison task clear before adding visual decoration.