Becoming Fluent in Data
I Introduction
Preface
Learning like a dolphin swims
Teach – Learn – Repeat
Doing something meaningful with data
About this book
Data literacy and data fluency
Who this book is for
How to read this book
What to expect
Open, modular, and evolving
Structure and visual language
About the author
Software
Why code?
Why R?
Why the Tidyverse?
Intro to R
R is a calculator
R is more than a calculator
Define objects
Plots
Intro to Tidyverse
Data with readr
Verbs of dplyr
Graphs with ggplot2
II Data
1
Data is everywhere
1.1
Why we measure
1.1.1
Women are having far fewer children.
1.1.2
Global surface temperature is rising.
1.2
Means of measuring
1.3
Types of data
1.3.1
Origin of data
1.3.2
Analysis of data
1.3.3
Structure of data
1.3.4
The level of access
1.4
Can we measure everything?
1.5
The reality behind the data
2
Stories and Visuals
2.1
Facts
2.2
Visualization
2.3
Telling a story
2.4
Man's best friend
2.5
Less is more
2.6
Grammar of Graphics
3
Tabular Data
3.1
Types of Tabular Data
3.1.1
Cross-section
3.1.2
Repeated cross-section
3.1.3
Time series
3.1.4
Panel data
4
Panel Data
4.1
Unemployment
4.1.1
On decline in Germany
4.1.2
Measurement
4.2
Application
4.2.1
Data Inspection
4.2.2
Data Preparation
4.2.3
Data Visualization
4.3
Panel Studies
5
Time data
5.1
Measuring Time
5.2
Measuring Dates
5.3
Your First Time (in R)
5.4
Time Zones
5.5
Time Management in R
5.5.1
Decimal Time
5.5.2
Time Formats
5.6
Coffee Spending
5.6.1
Spending Time of Day
5.6.2
Run Chart
5.6.3
Run Chart Grouped
5.7
Run Chart Grouped Cleaned
6
Web Data
6.1
Most expensive paintings
6.2
Student numbers at Viadrina
6.2.1
PDF scraping
6.2.2
Share of female students
6.2.3
Share of foreign students
6.2.4
Web scraping
6.2.5
Most recent student numbers
6.2.6
The long run trend
7
Geo Data
7.1
Geo coordinates
7.1.1
Where Are You?
7.1.2
Latitude and longitude
7.1.3
Angles and Degrees
7.1.4
Coordinate Reference System
7.1.5
Distance measurement
7.1.6
Points And Polygons
7.1.7
Shapefiles
7.2
Google Takeout
7.3
Blood Donation
8
Missing Data
8.1
Types of Missing Data
8.1.1
Missing Completely at Random (MCAR)
8.1.2
Missing at Random (MAR)
8.1.3
Missing Not at Random (MNAR)
8.2
Causes of Missing Data
8.3
Causes of Missing Data
8.4
Causes of Missing Data
8.5
Causes of Missing Data
8.6
Missing Data in R
III Analysis
9
Compare
9.1
Relationships
9.1.1
Storks Deliver Babies
9.1.2
Statistics
9.1.3
Visualizations
9.1.4
Spurious Relationships
Readings
9.2
Comparing Means
9.3
Partitioning Variation
10
Model
10.1
Regression
10.1.1
Old but Gold
10.1.2
Data is everywhere
10.1.3
For the truly dedicated
10.1.4
Survival of the Fittest Line
10.1.5
On the Shoulders of Giants
10.2
Linear Models
10.2.1
What You Deserve Is What You Get
10.2.2
Data & Sample
10.2.3
Data Visualization
10.2.4
Simplest Regression
10.2.5
Simple Regression
10.2.6
Parallel Slopes
10.2.7
Model Comparison
10.2.8
Transform to Perform
10.3
Decomposing Differences
10.4
Logistic Models
10.5
Interactions
10.5.1
Motivation
10.5.2
Data & Sample
10.5.3
Throwback Parallel Slopes
10.5.4
Regression with Moderators
10.5.5
Model Comparison
10.6
Marginal Effects
10.7
Generalized Models
10.8
Learning from Data
11
Structure
11.1
Time
11.2
Longitudinal Data
11.3
Multilevel Models
11.4
Places
12
Reveal
12.1
Components
12.2
Clusters
12.3
Latent Variables
12.3.1
How a Concept Becomes Data
12.3.2
One Item Can Be Enough: Happiness and Risk
12.3.3
Several Items, One Score: The Big Five
12.3.4
When a Model Is Necessary: Competence and Intelligence
12.3.5
How Many Latent Variables?
12.3.6
What a Latent Model Does
12.3.7
A First Map of Latent-Variable Models
12.4
Exploratory Factors
12.4.1
The Common-Factor Idea
12.4.2
Common Variance Is Not Total Variance
12.4.3
The Four Main EFA Decisions
12.4.4
How Many Factors?
12.4.5
Extracting the Factors
12.4.6
Why Rotation Is Necessary
12.4.7
Reading a Factor Solution
12.4.8
A Compact EFA Workflow in R
12.4.9
What Can Go Wrong?
12.4.10
From Exploration to Confirmation
12.4.11
Reporting an Exploratory Analysis
12.5
Confirmatory Factors
12.5.1
From a Pattern to a Measurement Claim
12.5.2
Fixed and Free Parameters
12.5.3
Identification and Scaling
12.5.4
Estimation
12.5.5
Reading the Parameters
12.5.6
Common Confirmatory Factor Models
12.5.7
Evaluating a Confirmatory Model
12.5.8
Beyond Universal Cutoffs
12.5.9
A Compact CFA Workflow in R
12.5.10
What Should Be Reported?
12.5.11
From Measurement to Structure
12.6
Structural Equations
12.6.1
Path Models as Systems of Regressions
12.6.2
Measurement Model Plus Structural Model
12.6.3
A Simple Latent Mediation Model
12.6.4
Measurement Comes First
12.6.5
SEM Is Not Automatically Causal
12.6.6
Where Higher-Order and Bifactor Models Belong
12.6.7
Model Fit Continues, but the Questions Multiply
12.6.8
What This Section Will Add Later
12.7
Items
12.7.1
Where IRT Sits among Latent Variable Models
12.7.2
Persons and Items on a Common Scale
12.7.3
Why Begin with Rasch?
12.7.4
IRT Is a Family of Models
12.7.5
Models within the Rasch Family
12.7.6
Three Choices That Are Often Mixed Together
12.7.7
From a Response Pattern to a Person Distribution
12.7.8
Person-Score Options
12.7.9
Why the Methods Give Different Scores
12.7.10
VERA as an IRT Application
12.7.11
Are VERA Items and IQB Data Public?
12.7.12
A Minimal Rasch Application in R
12.7.13
What You Need for Your Own VERA-Style Analysis
12.7.14
What Must Be Checked before Scores Are Trusted?
12.7.15
The Main Connections
12.7.16
Sources and Data Access
13
Identify
13.1
Experiments
13.2
Matching and Balancing
13.3
Difference-in-Differences
13.4
Discontinuities
13.5
Instruments
14
Extract
14.1
Text Mining
References
14.2
Resources
Visit my personal page
Becoming Fluent in Data
This book is
Work in Progress
. I appreciate your feedback to make the book better.
Chapter 14
Extract
Information from language
14.1
Text Mining
Finding patterns and information in text