Skip to content

References and Further Reading

The Ideas & Science pages connect Spillway's design to work in statistics, social science, human-computer interaction, decision theory, and machine learning.

This page separates two different uses of sources:

  • Direct support — a source substantiates a factual or scholarly claim made in the documentation.
  • Intellectual background — a source helps frame a concept or design question without implying that Spillway implements the source's exact model or method.

Individual essays should link here when they are pointing readers toward a field rather than making one source carry the whole argument.

Measurement, constructs, and uncertainty

  • Lee J. Cronbach and Paul E. Meehl, “Construct Validity in Psychological Tests,” Psychological Bulletin 52(4), 1955. A foundational account of construct validity and the difference between an observed measure and the latent concept it is intended to represent.
  • Josh Pasek, Gaurav Sood, and Jon A. Krosnick, “Misinformed About the Affordable Care Act? Leveraging Certainty to Assess the Prevalence of Misperceptions”, Journal of Communication 65(4), 2015. Directly relevant to the distinction between belief content and certainty.
  • Gabriel Miao Li, Josh Pasek, and Jon A. Krosnick, “A Certainty-Weighted, Belief-Based Model of Political Attitudes”, Political Psychology, 2025. Background for treating belief content, certainty, and influence as related but distinct.
  • Glenn W. Brier, “Verification of Forecasts Expressed in Terms of Probability,” Monthly Weather Review 78(1), 1950. Classic background for probabilistic calibration.

Evidence, dependence, and updating

  • Judea Pearl, Probabilistic Reasoning in Intelligent Systems, 1988. Background on probabilistic dependence and inference; Spillway does not attempt to implement one giant Bayesian network.
  • Thomas Bayes, “An Essay towards Solving a Problem in the Doctrine of Chances,” 1763, and the much larger body of Bayesian statistical work that followed. Spillway's “quasi-Bayesian” language refers to the discipline of updating beliefs with new evidence rather than to a literal single Bayesian model.
  • David J. Spiegelhalter, “Probabilistic Prediction in Patient Management and Clinical Trials,” Statistics in Medicine 5, 1986, among broader work on calibration and uncertainty. Useful background for separating probability quality from raw classification accuracy.

Decision theory and value of information

  • Howard Raiffa and Robert Schlaifer, Applied Statistical Decision Theory, 1961. Background for linking uncertainty to decisions rather than treating prediction as the endpoint.
  • Howard Raiffa, Decision Analysis, 1968. Classic introduction to decision-making under uncertainty.
  • Ron Howard, “Information Value Theory,” IEEE Transactions on Systems Science and Cybernetics 2(1), 1966. Background for asking whether learning more is worth the cost before acting.

Human behavior, affordances, and opportunity

  • James J. Gibson, The Ecological Approach to Visual Perception, 1979. Origin of the modern affordance concept: possibilities for action depend on the relation between an actor and an environment.
  • Donald A. Norman, The Design of Everyday Things, revised edition 2013. Human-computer-interaction background on perceived affordances, constraints, and design.
  • Herbert A. Simon, “A Behavioral Model of Rational Choice,” Quarterly Journal of Economics 69(1), 1955. Background for bounded rationality and decision-making under limited information and capacity.
  • Sendhil Mullainathan and Eldar Shafir, Scarcity, 2013. Broader background on how constrained attention and bandwidth change behavior; not a direct software prescription.

Digital traces are context-dependent measurements

  • Josh Pasek, Colleen A. McClain, Frank Newport, and Stephanie Marken, “Who’s Tweeting About the President? What Big Survey Data Can Tell Us About Digital Traces”, Social Science Computer Review 38(5), 2020. Directly relevant to the claim that platform behavior and survey responses are not interchangeable measures of the same construct.
  • danah boyd and Kate Crawford, “Critical Questions for Big Data,” Information, Communication & Society 15(5), 2012. Background on selection, context, and the limits of treating large behavioral traces as self-interpreting evidence.

Recommender systems, feedback loops, and proxy optimization

  • Recommender-systems research on implicit feedback treats clicks, views, purchases, and other behavior as noisy signals rather than direct declarations of preference. See Yifan Hu, Yehuda Koren, and Chris Volinsky, “Collaborative Filtering for Implicit Feedback Datasets,” ICDM, 2008, for one influential formulation.
  • Eli Pariser, The Filter Bubble, 2011, is a popular rather than technical reference for feedback between personalization and the information a person is subsequently exposed to.
  • Goodhart's law is often summarized as the problem that a measure can cease to be useful once it becomes a target. Charles Goodhart's original 1975 formulation concerned monetary policy; the broader principle is useful background for Spillway's concern about optimizing measurable proxies instead of underlying value.

Explainability and interpretable systems

  • Finale Doshi-Velez and Been Kim, “Towards A Rigorous Science of Interpretable Machine Learning,” 2017. Background for treating interpretability as something to evaluate rather than as a vague property.
  • Zachary C. Lipton, “The Mythos of Model Interpretability,” Queue 16(3), 2018. Useful background on the many different things people mean by “interpretable.”
  • Cynthia Rudin, “Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead,” Nature Machine Intelligence 1, 2019. A strong argument about high-stakes use; Spillway's system-level approach is related but not identical.

AI models and local-model comparisons

Model rankings change much faster than the conceptual sources above. Spillway therefore links to live resources rather than freezing a leaderboard into the documentation:

These resources measure different things. None is a Spillway-specific benchmark.

How to read these references

Spillway borrows questions and disciplines from these literatures more often than exact formulas.

For example, the project can take construct validity seriously without implementing a psychometric latent-variable model for every category. It can use value-of-information reasoning without solving a formal decision tree for every inference. It can preserve uncertainty without claiming full Bayesian posterior correctness.

That distinction is intentional. The goal is to prevent conceptual shortcuts from disappearing merely because the implementation uses pragmatic software techniques.

Social & Statistical Science · Computational Fallacies · Explainable AI · From Data to Knowledge


Documentation provenance: Iterative human–AI construction. See Documentation Provenance.