This book is Work in Progress. I appreciate your feedback to make the book better.

Chapter 4 Imperfect Data

What is absent, and what is invented

Real datasets have holes. People skip questions, sensors fail, records are lost — and the pattern of what is missing is rarely accidental, which is precisely why deleting the incomplete rows and carrying on is the most common and most costly shortcut in empirical work.

This chapter takes absence seriously in both directions. First: what to do when values are missing, and how to tell the harmless kind from the kind that will bias every result. Then the mirror image — deliberately manufacturing data, whether to protect the people behind it, to test a method against a truth you control, or to check whether a finding could have arisen by chance alone.