Week 3
Describing and visualizing Data
- Instructors
- Maclean Gaulin
Descriptive Statistics Overview
- Central tendency – where is the data located
- Mean (μ), median, mode
- Dispersion / variability – how spread out is the data
- Variance, standard deviation (σ), inter-quartile range, coefficient of variation (σ/μ), mean absolute deviation
- Shape – what are the “tails” of your data?
- Skewness / kurtosis, percentiles, min/max
- Frequency – categorical data
Key concept
descriptive statistics summarize the location, spread, and shape of data
Where: Central tendency
Key concept
central tendency metrics locate the 'center' of a distribution
How Spread Out: variability
Key concept
dispersion measures how spread out data points are around the center
Shape: Skewness
Key concept
skewness indicates asymmetry in the distribution tails
Tails: Kurtosis
Key concept
kurtosis measures the thickness of the distribution tails (outlier frequency)
Tails: Kurtosis
Key concept
heavy-tailed distributions contain more frequent extreme events than normal distributions
Shape – Kurtosis
Central tendency
- Mean (μ)
- Most common measure, basis of most statistics
- Weighs observations differently, sensitive to outliers
- Median
- The “typical” value, with half above & half below
- Immune to outliers
- Mode
- The “peak” of the distribution, most common value
- Can have multiple modes, but may be less informative
Key concept
median is robust to outliers, while the mean is highly sensitive
Dispersion / variability
- Standard deviation (σ)
- Most common dispersion measure, basis of most statistics
- Strongly affected by outliers (because math), normality
- Inter-quartile range
- Robust to outliers, good for skewed data, interpretable
- Coefficient of variation
- Dispersion as % of mean, comparable across units
- Mean absolute deviation (or w/ medians)
- Like σ, but less sensitive to outliers
Key concept
standard deviation is sensitive to outliers; IQR is robust
Shape
- Skewness – asymmetric “tails”
- Positive / right skewed: mean > median, more high values
- Negative / left skewed: mean < median, more low values
- Some models assume symmetrical data (skewness = 0)
- Kurtosis – heavy or light “tails”
- High heavy tails, low light tails
- Percentiles – bins of data
- Further resolution of full distribution, less of a summary
Key concept
skewed data violates standard normality assumptions in many statistical models
Frequency Information
- Count of values
- Most common value
- For ordinal data, 5-number summary
- Min, 25%, 50%, 75%, Max
| count | mean | std | min | 25% | 50% | 75% | max | |
|---|---|---|---|---|---|---|---|---|
| x | 142 | 54.3 | 16.8 | 22.3 | 44.1 | 53.3 | 64.7 | 98.2 |
| y | 142 | 47.8 | 26.9 | 2.9 | 25.3 | 46.0 | 68.5 | 99.5 |
Key concept
categorical analysis focuses on frequency, mode, and ordinal percentile distributions
What doN’T descriptives convey?
| Anscombe's quartet | ||||||||
|---|---|---|---|---|---|---|---|---|
| Dataset I | Dataset II | Dataset III | Dataset IV | |||||
| x | y | x | y | x | y | x | y | |
| 10.0 | 8.04 | 10.0 | 9.14 | 10.0 | 7.46 | 8.0 | 6.58 | |
| 8.0 | 6.95 | 8.0 | 8.14 | 8.0 | 6.77 | 8.0 | 5.76 | |
| 13.0 | 7.58 | 13.0 | 8.74 | 13.0 | 12.74 | 8.0 | 7.71 | |
| 9.0 | 8.81 | 9.0 | 8.77 | 9.0 | 7.11 | 8.0 | 8.84 | |
| 11.0 | 8.33 | 11.0 | 9.26 | 11.0 | 7.81 | 8.0 | 8.47 | |
| 14.0 | 9.96 | 14.0 | 8.1 | 14.0 | 8.84 | 8.0 | 7.04 | |
| 6.0 | 7.24 | 6.0 | 6.13 | 6.0 | 6.08 | 8.0 | 5.25 | |
| 4.0 | 4.26 | 4.0 | 3.1 | 4.0 | 5.39 | 19.0 | 12.5 | |
| 12.0 | 10.84 | 12.0 | 9.13 | 12.0 | 8.15 | 8.0 | 5.56 | |
| 7.0 | 4.82 | 7.0 | 7.26 | 7.0 | 6.42 | 8.0 | 7.91 | |
| 5.0 | 5.68 | 5.0 | 4.74 | 5.0 | 5.73 | 8.0 | 6.89 | |
| Mean | 9.0 | 7.5 | 9.0 | 7.5 | 9.0 | 7.5 | 9.0 | 7.5 |
| Variance | 10.0 | 3.75 | 10.0 | 3.75 | 10.0 | 3.75 | 10.0 | 3.75 |
| Correlation | 0.816 | 0.816 | 0.816 | 0.816 |
Key concept
identical summary statistics can hide vastly different data shapes
What doN’T descriptives convey?
| Anscombe's quartet | ||||||||
|---|---|---|---|---|---|---|---|---|
| Dataset I | Dataset II | Dataset III | Dataset IV | |||||
| x | y | x | y | x | y | x | y | |
| 10.0 | 8.04 | 10.0 | 9.14 | 10.0 | 7.46 | 8.0 | 6.58 | |
| 8.0 | 6.95 | 8.0 | 8.14 | 8.0 | 6.77 | 8.0 | 5.76 | |
| 13.0 | 7.58 | 13.0 | 8.74 | 13.0 | 12.74 | 8.0 | 7.71 | |
| 9.0 | 8.81 | 9.0 | 8.77 | 9.0 | 7.11 | 8.0 | 8.84 | |
| 11.0 | 8.33 | 11.0 | 9.26 | 11.0 | 7.81 | 8.0 | 8.47 | |
| 14.0 | 9.96 | 14.0 | 8.1 | 14.0 | 8.84 | 8.0 | 7.04 | |
| 6.0 | 7.24 | 6.0 | 6.13 | 6.0 | 6.08 | 8.0 | 5.25 | |
| 4.0 | 4.26 | 4.0 | 3.1 | 4.0 | 5.39 | 19.0 | 12.5 | |
| 12.0 | 10.84 | 12.0 | 9.13 | 12.0 | 8.15 | 8.0 | 5.56 | |
| 7.0 | 4.82 | 7.0 | 7.26 | 7.0 | 6.42 | 8.0 | 7.91 | |
| 5.0 | 5.68 | 5.0 | 4.74 | 5.0 | 5.73 | 8.0 | 6.89 | |
| Mean | 9.0 | 7.5 | 9.0 | 7.5 | 9.0 | 7.5 | 9.0 | 7.5 |
| Variance | 10.0 | 3.75 | 10.0 | 3.75 | 10.0 | 3.75 | 10.0 | 3.75 |
| Correlation | 0.816 | 0.816 | 0.816 | 0.816 |
Key concept
visualization reveals hidden structure the summary statistics erased
Why Visualize?
| Dataset | |
|---|---|
| x | y |
| 55.4 | 97.2 |
| 51.5 | 96.0 |
| 46.2 | 94.5 |
| 42.8 | 91.4 |
| 40.8 | 88.3 |
| 38.7 | 84.9 |
| 35.6 | 79.9 |
| 33.1 | 77.6 |
| 29.0 | 74.5 |
| 26.2 | 71.4 |
| 55.4 | 97.2 |
| … | … |
Key concept
identical statistics can mask different patterns
Why Visualize
- Human capacity for processing information is limited
- Human brain adapted to visual pattern recognition
- Improved understanding of complex data
- Identification of patterns, trends, and outliers
- Effective communication of analytical findings to various audiences
- Inform decisions – insight, reasoning, understanding
Key concept
human brains are optimized for visual pattern recognition
Leverage Strengths, Mitigate Weaknesses
- Leveraging strengths of human cognition
- Visual bandwidth1: About 20 Mbps (1Gbps total)
- Pre-attentive processing2: Instinctive processing
- Graphical perception3: Process some designs better
- Mitigate constraints of human cognition
- Avoid overloading limited “working memory”4
- Leverage external recognition
Key concept
visuals shift the burden from working memory to perception
Visualization as Communication
Key concept
visual communication must be tailored to the message and audience
Communicate by Storytelling
- A story is a set of data, observations, or events that are presented in a way or order intended to evoke a reaction or specific conclusion
- Can be as simple as well designed visualization
- Can be fully narrative, without tables or figures
Key concept
a story connects data to an intended conclusion
Data Example
Key concept
raw data without context fails to drive action
Storytelling Example
- As the global population grows, much of that growth is happening in cities. In 1950, 70% of the global population lived in rural areas. In 2050 that is expected to decline to 30%
Key concept
narratives help audiences remember and internalize facts
Visual Narrative
- Global populations are flocking to cities
- In 2050, that is expected to fall to 30%
Key concept
visual storytelling combines narrative structure with data visualization
How to tell stories
- Start with main idea
- What conclusion do you want your audience to draw?
- Consider audience and format, customize to match
- Find a hook – something to grab attention and anchor memory
- Storyboard – lay out the structure of your story
- Slides, post-it notes, pen and paper. Flexibility is key.
Key concept
begin with the main takeaway and storyboard the path to get there
Lab 3 - Visualizations
- Types of graphs
- Learning through visualization
Key concept
start with the question, not the chart type