A useful AI assistant can turn a meeting transcript into actions, explain a contract, summarize a support case or find patterns in a spreadsheet. The same convenience makes pasting feel frictionless. Frictionless is not the same as authorized.
A data boundary is a short decision process between the source material and the AI system. It asks what the material contains, who has authority to disclose it, which system will receive it, what that system may retain or reuse, and whether the task can succeed with less data. The goal is not blanket fear. It is deliberate routing.

“AI” is not one destination
A local model running without network access, a consumer chat account, an employer-managed enterprise workspace and an application built through an API can have different contracts, administrators, retention rules, training defaults, connectors and logging. Product names are not enough. Configuration and account context matter.
OpenAI's current consumer documentation, for example, separates model-improvement controls, memory controls, regular chat and Temporary Chat. It says Temporary Chats do not appear in history, do not use or create memories and are not used to train models, while copies may be retained for up to 30 days for safety. A managed workspace may apply organizational controls. Those details are useful, but they do not create permission to upload someone else's confidential record. Tool controls and disclosure authority are separate questions.
Step 1: classify the material
| Class | Examples | Default route |
|---|---|---|
| Public | Published article, public specification, your own non-sensitive draft | Allowed if the task and source license permit |
| Internal | Unreleased notes, routine operating documents, ordinary work product | Approved managed tool only; minimize first |
| Confidential | Contracts, strategy, private financials, source identities, security details | Do not paste without explicit policy and authority |
| Regulated or privileged | Health, education, legal, payment or government-controlled records | Use only a specifically approved workflow with professional oversight |
| Personal | Names, contact details, precise location, identifiers, private messages | Remove or transform unless necessary and authorized |
This is a working classification, not a legal opinion. Your organization or jurisdiction may define categories differently. If a formal policy exists, it controls. If the material belongs to a client, patient, student, employee, source or partner, your access to it does not automatically include authority to send it to a new processor.
Step 2: confirm authority and purpose
Write one sentence: “I am using this material to ___, and I have authority because ___.” If the second blank produces only “I can see the file,” stop. Access is not consent. A broad purpose such as improve my work also invites excessive disclosure; narrow it to extract dates from this redacted schedule or rewrite this paragraph without names.
NIST's Privacy Framework treats privacy as a risk-management problem across a data-processing ecosystem. That framing is helpful because the risk does not end at the paste box. It includes the organization operating the tool, subprocessors, administrators, connected applications, exported outputs and people affected by the processing.
Step 3: minimize before redacting
Redaction is valuable, but the best sensitive field is the one that never enters the prompt. Start by reducing the job:
- Paste only the relevant paragraph, rows or code path—not the whole repository or record.
- Replace names with stable labels such as Client A or Employee 2 when identity is unnecessary.
- Generalize dates, locations and dollar amounts if exact values do not affect the answer.
- Remove signatures, account numbers, access tokens, passwords, document metadata and hidden comments.
- Summarize the situation yourself and ask about the pattern rather than uploading the source.
Use the broader personal-data minimization protocol when the problem begins upstream. An AI boundary cannot compensate for collecting and retaining unnecessary data everywhere else.
Step 4: verify redaction as data, not appearance
Drawing a black rectangle over text in a document may leave the underlying text selectable, searchable or recoverable. A screenshot may include names in tabs, filenames, comments or notifications. A spreadsheet can hide columns without removing them. Export a sanitized copy, reopen it in a separate viewer and attempt to select, search and inspect the material.
Keep the original untouched in its authorized location. Give the sanitized copy a name that cannot be confused with the source. If reliable redaction requires specialized software or legal review, use it; do not improvise on a consequential record.
Step 5: select the environment
Check the exact account and workspace before uploading. Record the following:
- whether content may be used for model improvement and which setting controls it;
- retention and deletion behavior for chats, files and temporary modes;
- whether memory or personalization can retain details across conversations;
- who administers or can access the workspace;
- which connectors, plugins, actions or external tools may receive data;
- whether the organization has approved this specific tool for this data class;
- whether contractual terms cover the intended processing.
Do not infer enterprise protections from a work email address, or consumer behavior from a product family name. Open the current official documentation and your administrator's policy. Settings change. If the workflow depends on a promise, save the policy version or link and date in the decision record.
Step 6: map the whole route
The conversation window may be only the first hop. A custom assistant can call tools. A browser extension can read a page. A connector can search a drive or mailbox. Generated output can be copied into a ticket system. Map input, model, retrieval sources, tools, logs, exports and final recipients.
A connector with broad access can disclose more than the prompt contains. Grant the narrowest scope available, prefer a dedicated folder or test dataset and remove access when the task ends. The AI memory-controls guide covers another persistence layer: what the system recalls, surfaces and lets you delete.
Step 7: treat the output as another disclosure
Even when the input route is approved, the answer may reproduce sensitive text or infer details. Review before sharing. Remove identifiers, verify recipients and separate sourced fact from model-generated interpretation. For consequential claims, use the AI evidence ledger so the model's prose does not become its own source.
Stop rules
- You cannot identify the data owner or your authority to disclose.
- The material includes credentials, secrets, private keys or active access tokens.
- A policy, contract, privilege or regulation may prohibit the transfer.
- You cannot determine the destination's retention, training or connector behavior.
- The task can be completed with a fabricated, public or smaller dataset instead.
- A person could face material harm if the content were exposed or misinterpreted.
Stopping does not mean abandoning the task. Build a synthetic example, run a local approved model, ask the data owner, use an internal tool, or have a qualified professional perform the analysis within the authorized environment.
A lightweight boundary record
| Field | Record |
|---|---|
| Task | The narrow purpose and expected output |
| Data class | Public, internal, confidential, regulated/privileged or personal |
| Authority | Owner, consent, policy or contract permitting use |
| Minimization | Fields, rows and identifiers removed or generalized |
| Destination | Product, account/workspace, mode and approved configuration |
| Persistence | Retention, training, memory, logs and connectors checked |
| Output route | Where the answer will go and who will review it |
| Decision | Proceed, revise, escalate or stop—with owner and date |
Use the record when the stakes, novelty or sensitivity justify it. Routine public-text tasks do not need bureaucracy. The point is to make the decision reproducible when “I thought it was fine” would be inadequate.
Claims and boundaries
Sourced fact: NIST provides a voluntary privacy-risk framework, and OpenAI documents current consumer controls for training, memory, retention and Temporary Chat. Inference: a pre-paste gate reduces accidental over-disclosure across many AI products. Judgment: uncertain authority should default to a smaller or approved route. Not claimed: Temporary Chat makes confidential uploads appropriate, redaction eliminates all re-identification risk, or this guide replaces organizational policy or legal advice.
Official references
- NIST Privacy Framework, for managing privacy risk across data processing.
- OpenAI: How data is handled in consumer services, for current control categories.
- OpenAI: Temporary Chat FAQ, for current history, memory, training and retention behavior.
- OpenAI: Chat and file retention policies, for current deletion and retention details.
END OF FIELD GUIDE 078
Keep the question. Test the model.
Choose the narrowest claim the evidence can carry, then leave room for revision.