“X is associated with Y” and “changing X would change Y” are different claims. The first describes a pattern in observed data. The second asks about an intervention: what would happen to an outcome if a particular condition were changed while the relevant comparison remained meaningful.

That difference sounds elementary, yet causal language slips into ordinary summaries constantly. A dashboard rises after a policy launch. People who use a tool finish faster. A group with one habit reports better health. The sequence or association may be real. It does not, by itself, tell us what would have happened to the same population under another course.

Conceptual branching paths of wooden tiles with one brass lever changing a single route
A causal question changes one defined path and asks what follows. This original AI-assisted conceptual photograph is not a scientific experiment or causal diagram.
The intervention questionBefore accepting “X caused Y,” ask what precise change to X is imagined, compared with what alternative, for whom, over what time and through which evidence.

Association is evidence of a relationship, not its direction

An association says values co-vary under the way they were measured. It may be useful for prediction even when the causal story is unresolved. Smoke can predict fire without causing it. A symptom can predict disease without producing the disease. A user behavior can predict cancellation because it occurs after dissatisfaction has already begun.

Several structures can create the same surface pattern. X may affect Y. Y may affect X. A third factor may affect both. Selection into the dataset may depend on both variables. Measurement may change with context. The association may be ordinary sampling variation or a result selected from many analyses. Naming these possibilities does not prove the relationship is noncausal. It marks the work the causal claim still owes.

Prediction and intervention also optimize different questions. A model can accurately identify people at high risk without identifying an action that lowers their risk. Conversely, an intervention can have a modest average effect even when individual outcomes remain difficult to predict. The benchmark audit makes a similar distinction between test performance and operating consequence.

A causal effect compares potential outcomes

Modern causal inference often defines an effect by comparing potential outcomes under different interventions. For one person or unit, only one of those outcomes is normally observed at a given time. If a person takes an action, we do not simultaneously observe the same person, at the same moment and under identical conditions, not taking it.

This is the fundamental comparison problem, not a license for imagination. A study design tries to build a credible substitute: randomized groups, carefully adjusted observational comparisons, natural experiments, interrupted time series or other designs whose assumptions match the question. The unobserved outcome remains modeled. The discipline lies in making the comparison and assumptions explicit.

The site's essay A Counterfactual Is Not an Alternate Reality develops that point. Counterfactual language is useful because it forces the claim to say what would differ. It does not provide access to a hidden timeline.

Specify the causal question before selecting the method

“Does screen time cause harm?” is too broad to estimate. Which screen activity? Replaced with what? For which people? Which outcome? Over what interval? A defined causal question might compare disabling nonessential notifications with leaving current settings unchanged, among adults who receive more than a stated number per day, over four weeks, using a preselected measure of interruptions and a safety check for missed priority messages.

The intervention must be possible to describe. “Set stress to zero” is not one intervention; exercise, money, workload, medication and social support can change a stress measure through different mechanisms and side effects. Two actions that produce the same measured exposure are not automatically causally equivalent.

FieldQuestionCommon failure
PopulationFor whom or what does the claim apply?Generalizing beyond the people or systems studied
InterventionWhat exact change is made?Treating a vague state as one manipulable cause
ComparatorCompared with which realistic alternative?Using “nothing” when usual practice already changes outcomes
OutcomeWhat result is measured, and how?Substituting a convenient proxy for the consequence
TimeWhen does follow-up begin and end?Mixing immediate response with durable effect
EstimandAverage effect, subgroup effect, risk difference or something else?Letting one number stand for several questions
DecisionWhat action could the result change?Estimating an effect with no operational use

Design earns causal leverage; labels do not

Random assignment can make treatment groups comparable on average before an intervention, reducing many forms of confounding. It does not automatically repair poor adherence, missing outcomes, unblinded measurement, selective reporting, tiny samples or a population unlike the one where the decision will be made.

Observational evidence is not causally empty. Strong designs can exploit timing, thresholds, policy variation, repeated measures or rich subject-matter knowledge. But adjustment is not a ritual that converts a spreadsheet into a trial. A variable should be controlled because the causal structure justifies it. Controlling a consequence of the intervention can remove part of the effect; controlling a collider can create an association that was not there.

Causal diagrams help state those structural beliefs. They do not verify them. An arrow means the analyst is asserting a possible direct causal relation for the question at hand. The diagram makes assumptions discussable and can show which variables would block backdoor paths, but the truth still depends on knowledge, measurement and design.

Mechanism strengthens a claim without replacing identification

A plausible mechanism explains how an intervention could produce an outcome. It can guide measurement, identify delays, predict side effects and reveal why an effect might differ across contexts. It is not sufficient evidence that the proposed effect occurred. Many mechanisms are possible; several can operate at once; a mechanism demonstrated in a laboratory may be too weak in ordinary conditions.

The reverse mistake also happens: demanding a complete mechanism before accepting a well-identified effect. Reliable interventions sometimes precede full explanation. Evidence can establish that a change alters outcomes while leaving parts of the pathway unresolved. Keep effect identification, mechanism and transport to a new context as related but separate claims.

Average effects do not promise individual outcomes

A study may estimate an average causal effect across a population. That is not the effect for every member. Some people may benefit more, less or not at all; harms may be concentrated; subgroup estimates may be too imprecise to resolve the difference. The distribution audit asks which spread, tail and subgroup matter to the decision.

Do not infer an individual's alternate outcome from a population average with certainty. Nor should one vivid individual outcome erase a credible population effect. The levels answer different questions. Personal decisions may combine average evidence, known modifiers, values, cost, reversibility and qualified professional guidance where stakes are high.

The eight-step causal-claim audit

  1. Rewrite the claim as a change: “If we did X rather than C…”
  2. Name the population, outcome and follow-up period.
  3. Draw a simple timeline so causes do not occur after their alleged effects.
  4. List common causes of exposure and outcome before looking at adjustment results.
  5. Identify selection, missingness and measurement processes that could create the pattern.
  6. Ask which design feature—not which adjective—supports the comparison.
  7. Separate estimated effect, proposed mechanism and transfer to the present context.
  8. State what evidence or result would weaken the conclusion and what decision the estimate changes.

Use the claim–evidence–inference worksheet to record the bridge. For a scientific paper, the paper-reading protocol adds sampling, bias, uncertainty and applicability checks.

What the evidence can and cannot carry

Sourced fact: causal effects require a defined contrast between potential outcomes, and randomized or observational designs identify that contrast only under assumptions. Inference: public explanations become more honest when they specify intervention, comparator, outcome and time instead of using “cause” as a synonym for importance. Judgment: how much causal certainty is enough depends on the decision's stakes, reversibility and alternatives. Not claimed: this audit can turn weak data into a reliable causal estimate.

Causation is not a prestige label added after analysis. It is a structure: a change, a comparison, a pathway, a population, a time horizon and assumptions that someone else can inspect.

Primary and authoritative sources


END OF TRANSMISSION 050

Keep the question. Test the model.

Choose the narrowest claim the evidence can carry, then leave room for revision.