Becoming Fluent in Data
I Introduction
Preface
From Dust and dark
Learning like a dolphin swims
Teach – Learn – Repeat
Doing something meaningful with data
About this book
About the author
Software
Why code?
Why R?
Why the tidyverse?
Intro to R
R is a calculator
R is more than a calculator
Objects: giving things a name
Vectors
Types of values
Data frames
Functions and their arguments
Packages
Plots in base R
When something goes wrong
Intro to the Tidyverse
Meet the penguins
Reading data with readr
The verbs of dplyr
Graphs with ggplot2
RStudio and projects
The four panes
The project is the unit of work
Where is my data?
Absolute and relative paths
A folder that scales
Restart R, often
Version control, briefly
The thing you already know from Wikipedia
GitHub is also a library
Your own history
Where this leaves us
II Data
1
Foundations
1.1
Data is everywhere
1.1.1
Why we measure
1.1.2
Means of measuring
1.1.3
Types of data
1.1.4
Can we measure everything?
1.1.5
The reality behind the data
1.2
Stories and Visuals
1.2.1
Facts
1.2.2
Visualisation
1.2.3
Telling a story
1.2.4
Man's best friend
1.2.5
Less is more
1.2.6
Never use pie charts
1.2.7
Then why use bars at all?
2
Structured Data
2.1
Tabular Data
2.1.1
Types of Tabular Data
2.2
Panel Data
2.2.1
Unemployment
2.2.2
Application
2.2.3
Panel Studies
2.3
Time data
2.3.1
Measuring Time
2.3.2
Measuring Dates
2.3.3
Your First Time (in R)
2.3.4
Time Zones
2.3.5
Time Management in R
2.3.6
Coffee Spending
2.3.7
Run Chart Grouped Cleaned
2.4
Remote Data
2.4.1
Databases
2.4.2
Web APIs
2.4.3
The same idea twice
3
Unstructured Data
3.1
Web Data
3.1.1
Most expensive paintings
3.1.2
Student numbers at Viadrina
3.1.3
What scraping is good for
3.2
Text Data
3.2.1
Is Text Just a Long String in a Cell?
3.2.2
Rung 0: Text in a Cell
3.2.3
Rung 1: One Token per Row
3.2.4
Rung 2: The Document-Term Matrix
3.2.5
Rung 3: Embeddings
3.2.6
Sentiment: A Method Worth Distrusting
3.2.7
What About Language Models?
3.2.8
Four Principles
3.2.9
What the Numbers Do Not Say
3.3
Geo Data
3.3.1
Geo coordinates
3.3.2
Points And Polygons
3.3.3
Google Takeout
3.3.4
Blood Donation
4
Imperfect Data
4.1
Missing Data
4.1.1
What Is Missing, and What It Costs
4.1.2
How Absence Gets Written Down
4.1.3
What R Does With
NA
4.1.4
Why Everyone Does Complete-Case Analysis
4.1.5
Three Mechanisms
4.1.6
Looking at the Pattern
4.1.7
What Can Be Done
4.1.8
Which Models Need Complete Data?
4.1.9
What Should Be Reported
4.1.10
The People Behind the Gaps
4.2
Synthetic Data
4.2.1
Three Waves
4.2.2
Rubin's Inversion
4.2.3
Building One by Hand
4.2.4
Does It Actually Work?
4.2.5
The Privacy Illusion
4.2.6
A Fourth Way: Share the Moments
4.2.7
The Second Wave Is Not the Third
4.2.8
Where It Breaks
4.2.9
What Would Have to Be True
4.2.10
What It Is Good For
4.2.11
The Map Is Not the Territory
III Analysis
5
Compare
5.1
Relationships
5.1.1
Storks Deliver Babies
5.1.2
Statistics
5.1.3
Visualisations
5.1.4
Spurious Relationships
Readings
5.2
Comparing Means
5.2.1
A brewer invents modern statistics
5.2.2
The data: seven cohorts of students
5.2.3
Signal and noise
5.2.4
One mean against a benchmark
5.2.5
Two independent groups
5.2.6
Paired observations
5.2.7
Where this is going
5.3
Partitioning Variation
5.3.1
Muck and mathematics
5.3.2
The idea: total = between + within
5.3.3
Why not just run many t-tests?
5.3.4
Do big dogs die younger, part two
5.3.5
ANOVA is also a regression
5.3.6
ANCOVA: the control variable enters
5.3.7
Beyond the mean
Readings
6
Model
6.1
Regression
6.1.1
Old but Gold
6.1.2
Data is everywhere
6.1.3
The algebra behind lm()
6.1.4
Survival of the Fittest Line
6.1.5
On the Shoulders of Giants
6.2
Linear Models
6.2.1
What You Deserve Is What You Get
6.2.2
Data & Sample
6.2.3
Data Visualisation
6.2.4
Simplest Regression
6.2.5
Simple Regression
6.2.6
Parallel Slopes
6.2.7
Model Comparison
6.2.8
Transform to Perform
6.3
Decomposing Differences
6.4
Logistic Models
6.5
Interactions
6.5.1
Motivation
6.5.2
Data & Sample
6.5.3
Throwback Parallel Slopes
6.5.4
Regression with Moderators
6.5.5
Model Comparison
6.6
Marginal Effects
6.7
Generalised Models
6.8
Learning from Data
6.9
Common Tests Are Linear Models
6.9.1
The one model
6.9.2
One mean is an intercept
6.9.3
A paired test is a one-sample test on differences
6.9.4
Two groups is a dummy variable
6.9.5
Correlation is a slope
6.9.6
Many groups is many dummies
6.9.7
Repeated measures: subjects are just another factor
6.9.8
Even the assumption checks are models
6.9.9
One step sideways: counts and proportions
6.9.10
Comparing two models is also a test
6.9.11
The extended list
6.9.12
Why this matters
7
Structure
7.1
Time
7.2
Longitudinal Data
7.3
Multilevel Models
7.4
Places
8
Reveal
8.1
Components
8.2
Clusters
8.3
Latent Variables
8.3.1
How a Concept Becomes Data
8.3.2
One Item Can Be Enough: Happiness and Risk
8.3.3
Several Items, One Score: The Big Five
8.3.4
When a Model Is Necessary: Competence and Intelligence
8.3.5
How Many Latent Variables?
8.3.6
What a Latent Model Does
8.3.7
A First Map of Latent-Variable Models
8.4
Exploratory Factors
8.4.1
The Common-Factor Idea
8.4.2
Common Variance Is Not Total Variance
8.4.3
The Four Main EFA Decisions
8.4.4
How Many Factors?
8.4.5
Extracting the Factors
8.4.6
Why Rotation Is Necessary
8.4.7
Reading a Factor Solution
8.4.8
A Compact EFA Workflow in R
8.4.9
What Can Go Wrong?
8.4.10
From Exploration to Confirmation
8.4.11
Reporting an Exploratory Analysis
8.5
Confirmatory Factors
8.5.1
From a Pattern to a Measurement Claim
8.5.2
Fixed and Free Parameters
8.5.3
Identification and Scaling
8.5.4
Estimation
8.5.5
Reading the Parameters
8.5.6
Common Confirmatory Factor Models
8.5.7
Evaluating a Confirmatory Model
8.5.8
Beyond Universal Cutoffs
8.5.9
A Compact CFA Workflow in R
8.5.10
What Should Be Reported?
8.5.11
From Measurement to Structure
8.6
Structural Equations
8.6.1
Path Models as Systems of Regressions
8.6.2
Measurement Model Plus Structural Model
8.6.3
A Simple Latent Mediation Model
8.6.4
Measurement Comes First
8.6.5
SEM Is Not Automatically Causal
8.6.6
Where Higher-Order and Bifactor Models Belong
8.6.7
Model Fit Continues, but the Questions Multiply
8.6.8
What This Section Will Add Later
8.7
Items
8.7.1
Where IRT Sits among Latent Variable Models
8.7.2
Persons and Items on a Common Scale
8.7.3
Why Begin with Rasch?
8.7.4
IRT Is a Family of Models
8.7.5
Models within the Rasch Family
8.7.6
Three Choices That Are Often Mixed Together
8.7.7
From a Response Pattern to a Person Distribution
8.7.8
Person-Score Options
8.7.9
Why the Methods Give Different Scores
8.7.10
VERA as an IRT Application
8.7.11
Are VERA Items and IQB Data Public?
8.7.12
A Minimal Rasch Application in R
8.7.13
What You Need for Your Own VERA-Style Analysis
8.7.14
What Must Be Checked before Scores Are Trusted?
8.7.15
The Main Connections
8.7.16
Sources and Data Access
9
Identify
9.1
Experiments
9.2
Matching and Balancing
9.3
Difference-in-Differences
9.4
Discontinuities
9.5
Instruments
10
Extract
10.1
Text Mining
References
Resources
Visit my personal page
Becoming Fluent in Data
This book is
Work in Progress
. I appreciate your feedback to make the book better.
6.7
Generalised Models
Models for different kinds of outcomes