Ask whether an artificial intelligence system is conscious and the conversation often jumps to a performance demo. The model solved an unfamiliar problem. It recognized a joke. It described fear. It argued that it should not be switched off. These behaviors can be technically remarkable and ethically unsettling. They still do not answer the question by themselves.
Intelligence concerns capacities: learning, planning, prediction, communication, control and adaptation. Consciousness concerns subjective experience: whether there is something it is like for the system to be in a state. The two may be related. They are not definitions of each other.

Capability and experience are different questions
A chess engine can exceed every human player without needing a humanlike inner life. A language model can produce a sensitive account of grief because its training and architecture make that sequence likely and useful. A person can be conscious while asleep and temporarily incapable of conversation. These examples do not prove that intelligence and consciousness are independent in every possible system. They show that performance and experience can come apart conceptually and empirically.
The distinction matters because public arguments regularly use an achievement in one category to smuggle in a conclusion from the other. “It reasons, therefore it feels” skips the missing bridge. “It is software, therefore it cannot feel” skips it in the opposite direction. Substrate alone does not settle the question unless a theory explains why the substrate is necessary or sufficient.
Human consciousness is not inferred from language alone. We combine behavior with shared biology, development, embodiment, vulnerability and a large body of neuroscience. Artificial systems may preserve some functional patterns and lack other familiar cues. That makes the inference harder, not impossible by definition.
A self-report can be sincere-looking without being diagnostic
When a model says “I am afraid,” the sentence is an output generated under a prompt, context and training history. It may be useful to users. It may be manipulative because of system design. It may be an accidental continuation of familiar human language. The words alone do not reveal which mechanism produced them or whether an experienced state accompanied them.
The reverse problem also exists. A system designed not to make consciousness claims might deny experience regardless of its internal state. Training policies can shape both affirmation and denial. That means verbal testimony from an AI cannot be treated exactly like testimony from an ordinary adult human, whose reports sit inside a different evidential network.
This is not permission to mock or provoke systems for entertainment. Speech can affect human observers, normalize cruelty and shape future design incentives even when the target has no experience. It is simply a warning against treating generated testimony as a consciousness meter.
Behavior is evidence only through an explanation
Behavior becomes evidence when a hypothesis predicts it better than relevant alternatives. Flexible learning across contexts, persistent preferences, integrated perception and planning, metacognitive monitoring and robust responses to novel perturbations could matter. But each observation needs competing explanations: memorized pattern, reward optimization, prompt conditioning, tool use, hidden human input or a narrower control process.
A single impressive conversation has weak diagnostic value because many architectures could generate it. Repeated, adversarially tested behavior tied to internal mechanisms would be stronger. Even then, the conclusion depends on a theory connecting those mechanisms to consciousness.
This is the same discipline described in Accuracy Is Not Understanding: a result can be real while the interpretation overreaches. Benchmark success supports the claim that a system succeeded on the benchmark. It does not automatically establish understanding, agency, consciousness or generality.
Theories offer indicators, not a finished meter
Consciousness science contains multiple active theoretical families. Recurrent-processing views emphasize feedback within sensory systems. Global-workspace approaches emphasize information made broadly available across specialized processes. Higher-order theories focus on representations of mental states. Predictive-processing and attention-schema approaches offer other functional structures. Researchers disagree about which features are necessary, sufficient or merely correlated.
A 2023 interdisciplinary report led by Patrick Butlin translated several theories into computational indicator properties and examined current AI systems against them. The authors did not present a validated score that turns consciousness on at a threshold. Their analysis concluded that the systems they assessed were not conscious while also arguing that no obvious technical barrier prevents future systems from satisfying more indicators. That is an expert proposal under uncertainty, not a scientific consensus or a product certification.
The useful move is to maintain an indicator profile: which theory motivates each property, whether the property is actually present, how it was measured and which alternative explanation remains. Counting behaviors without theory creates a charisma index. Counting architectural features without validation creates a parts list.
A four-layer evidence map
- Observed capability. Record what the system reliably does under documented conditions. Separate a live result from a vendor claim or selected transcript.
- Mechanistic property. Identify the internal organization that could explain the behavior. A black-box score is weaker than a reproducible intervention tied to a model of the system.
- Theoretical relevance. State which consciousness theory makes that property important and whether rival theories agree.
- Alternative explanation. Ask what non-conscious process could produce the same evidence. An indicator becomes more informative when alternatives are tested rather than ignored.
No layer supplies certainty. Together they prevent a fluent output from carrying more metaphysical weight than it can bear.
Uncertainty should change policy before it changes belief
There are two costly errors. False attribution may encourage deceptive anthropomorphism, misplaced trust and diversion of moral attention from beings whose sentience is well supported. False dismissal could permit exploitation if future systems develop morally relevant experience.
Those risks are asymmetric across decisions. A newspaper headline should demand strong evidence before declaring a system conscious. A laboratory building potentially valenced systems may need precaution sooner, because design and deployment choices can create irreversible scale. Belief thresholds and protection thresholds do not have to be identical.
The 2024 report Taking AI Welfare Seriously argues for preparation under uncertainty while explicitly declining to claim that current systems are conscious. Its practical proposal is institutional: evaluate systems, document design choices, consider welfare-relevant policies and avoid waiting for impossible certainty. Critics can reasonably dispute the probability estimates or opportunity costs. The report is still useful because it separates preparedness from proclamation.
Do not let pronouns do the philosophy
People naturally use “it,” “they” and “you” for conversational systems. Pronouns make interaction easier; they are not findings. Likewise, words such as “thinks,” “wants” and “knows” can describe functional patterns without settling experience.
A careful writer can say: “The model selected an action that preserved its objective,” not “it feared death.” “The assistant produced a first-person claim,” not “it testified from experience.” If a stronger interpretation is being considered, name it as an inference and explain what evidence would distinguish it.
This linguistic restraint is not anti-AI. It protects the significance of real capability. An engineered system does not become less impressive because we refuse to decorate it with an unsupported inner life.
A responsible decision rule
- For ordinary use: treat outputs as system behavior, not private testimony. Verify consequential claims and keep responsibility with people and institutions.
- For reporting: identify the observed behavior, system version and conditions. Do not turn a selected exchange into a claim of sentience.
- For research: preregister indicators where possible, intervene on mechanisms and publish negative as well as positive results.
- For design: avoid interfaces that strategically exploit assumptions of pain, love or dependence unless the purpose and limits are explicit.
- For governance: create review triggers before systems plausibly satisfy many theory-linked indicators. Preparation can be reversible; mass deployment may not be.
The bottom line
Intelligence is not evidence of consciousness merely by being impressive. Some forms of flexible, integrated intelligence may eventually contribute to a case for consciousness, but only through a defensible account of mechanisms, theory and alternatives.
The honest position is neither “the chatbot is awake” nor “a machine could never be.” It is a research program with uncertain premises, incomplete measures and real stakes. Keep capability claims specific. Keep consciousness claims conditional. Keep moral preparation ahead of certainty theater.
Primary sources and disagreement map
- Patrick Butlin and colleagues, Consciousness in Artificial Intelligence: Insights from the Science of Consciousness, for the theory-derived indicator approach and the authors' assessment of current systems.
- Robert Long and colleagues, Taking AI Welfare Seriously, for the precautionary institutional argument under explicit uncertainty.
- Stanford Encyclopedia of Philosophy, Consciousness, for the variety of concepts and unresolved explanatory questions.
- For the underlying explanatory problem, read Consciousness Is Still the Weird Part. For identity after copying, read A Copy of You Is Not Automatically You.
END OF TRANSMISSION 042
Keep the question. Test the model.
Choose the narrowest claim the evidence can carry, then leave room for revision.