You change one part of your routine and feel better. Was the change responsible? Perhaps. Perhaps the week was quieter, expectations did the work, measurement changed behavior, or the result was ordinary variation. A one-person experiment cannot eliminate every alternative. It can make the question narrower and the answer less vulnerable to memory.
This guide adapts ideas from N-of-1 trial design to ordinary, low-risk choices such as notification settings, meeting placement, reading time, workspace arrangement or the timing of a familiar routine. It is not a clinical-trial protocol and not permission to start, stop, delay or compare medicines, supplements, treatment, sleep deprivation, restrictive diets or risky physical practices. Health decisions belong with a qualified clinician.

Step 1: ask a question one person can answer
Good personal questions are narrow, observable and reversible. “Does placing my phone outside the room during a 60-minute writing block change completed draft words and perceived concentration?” is testable. “Does technology destroy my creativity?” is not.
Write the population honestly: you, in these weeks, under these conditions. The result may help you choose a routine. It does not establish what everyone should do. If the question concerns other people, workplace policy or a household member, their outcomes and consent cannot be absorbed into your experiment.
Step 2: draw the safety boundary first
Use this protocol only when both conditions are ordinary and acceptable, stopping is easy, and a disappointing result will not cause meaningful harm. Do not self-experiment with prescription or nonprescription drugs, supplements, allergens, substance use, dangerous temperatures, fasting, deliberate sleep loss, pain, breathing restriction or intense exertion. Do not postpone care to finish a data series.
If symptoms, treatment or safety are involved, speak with a clinician. Formal N-of-1 trials can be valuable in medicine, but they use professional oversight, defined outcomes, ethical safeguards and design choices that a casual tracker does not reproduce.
Step 3: write the prediction and decision before collecting data
Precommitment prevents the target from moving after the result. Record four lines:
- Prediction: which condition do you expect to help, by roughly how much?
- Primary outcome: the one measure that decides the comparison.
- Guardrail: one cost that must not worsen.
- Decision rule: what result would make you keep, reject or retest the change?
Choose one primary outcome. You can record supporting observations, but twenty measures make it easy to select the flattering one afterward. The measure should be close to the question: completed practice problems for study structure, minutes to fall into the planned task for a workspace change, or remembered claims from a reading session for a reading method.
Step 4: measure a baseline without improving it
Observe the current routine for several comparable sessions before intervening. A baseline shows ordinary variation and whether the outcome can be recorded consistently. If the measurement changes the behavior, that is part of the system: continue until the novelty settles or admit that measurement and intervention cannot be separated.
Record context that plausibly matters—day, duration, location, unusual interruption and sleep quality—without turning the exercise into surveillance. Context helps explain a strange session; it should not become a license to delete every result you dislike.
Step 5: compare one change with a credible control
Keep duration, task type and time window as similar as practical. Change one main condition. If condition A is “phone on desk,” condition B might be “phone in another room.” Do not simultaneously add a new playlist, drink, desk and reward, because the bundle cannot tell you which part mattered.
A no-change control is often enough. A sham or placebo condition is rarely practical for everyday routines, and pretending to be blinded when you know the condition adds theater rather than rigor. Name the expectation effect instead. Sometimes the belief that a ritual should work is part of its useful effect; the experiment simply cannot separate it from mechanism.
Step 6: alternate or randomize the order
Improvement over time can make the later condition look better. Weekdays can differ from weekends. Randomizing or counterbalancing order reduces these patterns. For six comparable sessions, shuffle three A cards and three B cards, reveal one just before each session, and follow the assigned condition.
If the effect lingers, rapid alternation is inappropriate. Formal N-of-1 research treats carryover and washout as important design problems. For ordinary routines, either allow enough return-to-baseline time or choose a question without meaningful carryover. If you cannot tell how long the effect persists, the study may not support a clean comparison.
Step 7: repeat enough to see variability
One good day is an anecdote. Repeated A–B periods show whether the difference survives different mornings and ordinary disruption. There is no universal magic number for a casual personal experiment. Choose the number of sessions in advance, make it long enough to include variation, and short enough that the surrounding life is reasonably comparable.
Stop early for a guardrail failure, not because the favorite condition is losing. If work, relationships, mood or safety deteriorate, restore the prior routine and record why. The reversibility test helps decide whether a change is suitable for experimentation at all.
Step 8: record immediately and minimally
Use one row per session: date, assigned condition, primary outcome, guardrail, and one context note. Record before checking earlier results. A simple paper sheet is enough. More tracking does not produce more truth if the data become inconsistent or the burden changes the routine being tested.
For subjective outcomes, define the scale in advance. “Concentration: 1 means repeatedly unable to resume; 3 means several recoverable distractions; 5 means sustained work with no memorable diversion.” Anchors do not make the rating objective, but they make your use of it less elastic.
Step 9: compare distributions, not only averages
List the results for A and B. Compare the median, the spread and the worst session. Ask whether one condition produces a small average improvement but occasional unacceptable failures. For a short, informal experiment, a graph or complex significance test can imply more precision than the design earned.
Look for overlap. If most A and B sessions are similar, the practical conclusion may be that the choice does not matter enough to maintain. If the result is large but rests on one unusual day, repeat. If the effect reverses by context, the useful rule may be conditional rather than universal.
Step 10: write a narrow conclusion
Use this template: “For me, during these sessions, condition B was associated with [difference] in [primary outcome], while [guardrail] did/did not worsen. Because [main limitation], I will [keep/reject/retest] it until [review date].”
Do not upgrade association to mechanism. Do not generalize to people you did not study. Do not call a null result proof that the change never works. The experiment is evidence for one decision, not a new identity.
Worked example: an interruption boundary
Question: during comparable 60-minute editing blocks, does keeping the phone in another room reduce the time needed to complete one planned section? Baseline: four sessions with the existing setup. Conditions: A, phone face down on desk; B, phone in another room with one emergency contact allowed through a separate channel. Order: three shuffled A and three shuffled B sessions.
Primary outcome: whether the planned section reaches the defined edit state within the hour. Guardrail: no missed urgent family message. Stop rule: restore the phone if the emergency path fails. Decision: keep B only if it improves completion across more than one session without failing the guardrail.
This design still has weaknesses: the person cannot be blinded, tasks differ, and anticipation may alter behavior. Writing those limits is part of the result.
One-page protocol
- Write one narrow, low-risk question.
- Name the excluded safety domains.
- Choose one primary outcome and one guardrail.
- Record the predicted direction and decision rule.
- Measure several baseline sessions.
- Define A and B with one main difference.
- Randomize or counterbalance the order when practical.
- Preselect the number of sessions and stop rule.
- Record each session immediately.
- Compare typical result, variability and worst case.
- Write one narrow conclusion and review date.
Method notes and further reading
- Lillie and colleagues, The n-of-1 clinical trial: the ultimate strategy for individualizing medicine?, on randomization, carryover, washout and blinding.
- Vohra and colleagues, CONSORT extension for reporting N-of-1 trials.
- Hawksworth and colleagues, A methodological review of randomized N-of-1 trials.
- For interpretation discipline, read How to Read a Scientific Paper Without Borrowing Its Confidence, Keep a Decision Journal and Uncertainty Is Information, Not Failure.
END OF FIELD GUIDE 038
Keep the question. Test the model.
Choose the narrowest claim the evidence can carry, then leave room for revision.