Set the evidence boundary first
This guide covers interface behavior: how an AI chat product represents conversation state, constrains actions, preserves information, asks for approval, exposes uncertainty, and recovers from disruption. It also explains how to verify that a named implementation follows those rules.
The boundary matters. An interface can accurately display a completed state while the underlying answer is wrong. It can meet its component contract while violating a safety policy. It may look accessible in one fixture but fail with a keyboard, assistive technology, zoom, or real content. Each claim needs its own evidence.
| Claim | Relevant evidence | What it does not prove | |
|---|---|---|---|
| Interface state is represented correctly | The named state appears as declared | State contract plus an observation under a recorded condition | Model accuracy or task success |
| Visual system is applied | The named consumer follows the visual system | Versioned tokens, mappings, component use, and rendered inspection | Correct chat behavior or release readiness |
| Interaction is accessible | The tested interaction meets its recorded expectation | Accessibility tests in named components and representative consumers | Universal conformance outside the tested scope |
| Action is safe to execute | The named action satisfies its safety contract | Product safety policy, authorization rules, and execution controls | That a polished confirmation dialog is sufficient |
| Product is ready to release | The named release meets all applicable gates | Combined product, safety, accessibility, operational, and business evidence | Any isolated screenshot, story, or token-file check |
Never infer completion from prose
A model can produce a confident sentence before a tool finishes, after it fails, or while approval remains unresolved. Treat generated content and execution state as separate evidence channels.
Map the conversation lifecycle
Start with transitions, not a component inventory. A message component is easy to draw. The harder question is what must change as a request moves from local composition to submission, streaming, tool execution, approval, and recovery.
- 1
Idle
No draft or operation is active. The composer is available, and any conversation-level status is current.
- 2
Composing
The user has an unsent draft. Preserve it across harmless navigation and recoverable connection changes according to the product's stated persistence policy.
- 3
Submitting
The request has left the local composer, but the system hasn't established a response stream. Prevent accidental duplicate submission or make the deduplication rule explicit.
- 4
Streaming
Assistant content is arriving. Show that the answer is provisional, expose any supported interruption action, and preserve the user request that caused the stream.
- 5
Waiting for a tool
The assistant can't continue until an external operation returns. Distinguish queued, running, delayed, and failed activity when those differences affect user action.
- 6
Waiting for approval
A proposed action requires a person to approve, reject, or revise it. Identify the exact action and scope. Continued prose must not imply that approval was granted.
- 7
Completed
The in-scope response or operation reached its declared terminal condition. System evidence must support completion; tone is not evidence.
- 8
Interrupted
Generation or execution stopped before its intended terminal condition. Preserve useful partial output, label its status, and offer only recovery actions the system can honor.
- 9
Cancelled
The user or system intentionally ended the operation. Record whether cancellation was requested, acknowledged, and effective, especially when an external tool may already have acted.
- 10
Partially completed
Some declared work succeeded and some didn't. Name the completed and unresolved parts instead of collapsing the result into success or failure.
- 11
Failed
The operation can't continue without a retry, changed input, restored dependency, or human intervention. Preserve the request and enough diagnostic context for the supported recovery path.
Treat these states as a starting contract, not a universal finite-state machine. A research assistant, coding agent, support bot, and transaction agent need different substates. Add a state when it changes visible evidence, permitted actions, persistence, ownership, or recovery. If it changes none of those, it's probably an implementation detail.
Copy the state-contract matrix
Use one record for each state or high-risk transition. A compact matrix helps the team spot omissions, while the full record keeps every decision inspectable.
STATE CONTRACT
State: <stable name>
Entry condition: <event and prior state>
Exit conditions: <allowed next states>
Visible evidence: <status, labels, timestamps, progress, partial output>
Permitted user actions: <send, stop, edit, approve, reject, retry, resume>
Blocked actions: <actions hidden or disabled, with reason>
Persistence: <draft, messages, tool results, approvals, partial output>
Provenance or uncertainty: <source, freshness, confidence limits, unresolved status>
Recovery: <automatic and user-triggered routes>
Protected surfaces: <content or controls that must not change>
Owner: <role that owns this contract>
Required observations: <conditions and expected results>
Open questions: <unresolved decisions with owners>
Evidence disposition: required | conditional | unresolved | not_applicableThe four evidence dispositions stop an empty field from passing as a decision. Required means acceptance depends on it. Conditional means it becomes required under a named condition. Unresolved marks a missing decision or observation. Not applicable needs a reason because it's a scoped conclusion, not a shortcut.
| Test content | Inspection criteria and failure condition | |
|---|---|---|
| Idle or composing | Empty composer, long draft, and an attachment if supported | Inspect readiness, draft retention, and available actions. Fail if stale status remains or a promised draft disappears without disclosure. |
| Submitting | One request plus a repeated send attempt | Inspect pending evidence and duplicate handling. Fail if both requests can create the same unintended effect or the user cannot tell whether submission was accepted. |
| Streaming | Long prose, code, citations, and an active Stop action | Inspect provisional status, partial-output retention, and interruption controls. Fail if prose appears terminal while execution remains active. |
| Tool running | A delayed operation with a stable invocation identity | Inspect operation name, scope, status, and safe actions. Fail if the UI hides whether the tool is queued, active, delayed, or failed when that difference affects recovery. |
| Approval pending | A concrete action, resolved target, consequence, and stale-approval condition | Inspect scope, available decisions, persistence, and revalidation. Fail if approval can apply to an ambiguous or changed target. |
| Interrupted or cancelled | Received partial output and a stop or cancel request | Inspect request, acknowledgement, retained content, and any work that may continue. Fail if a requested stop is displayed as effective before acknowledgement. |
| Partially completed | Several sub-operations with at least one success and one failure | Inspect successful and unresolved scopes separately. Fail if retrying can duplicate successful work without warning or control. |
| Failed | A recoverable dependency failure and a non-recoverable case | Inspect usable explanation, preserved input, and valid next actions. Fail if the interface offers a retry that cannot work under the recorded condition. |
| Completed | A response whose terminal condition is independently available | Inspect terminal system evidence and retained provenance. Fail if completion is inferred only from generated wording. |
Review transitions as well as states
The most dangerous ambiguity often sits between rows: stop requested versus stop acknowledged, approval shown versus approval recorded, or reconnect started versus the operation actually resumed.
Separate content state from system state
The transcript holds the content. The execution lifecycle holds the system state. They often move together, but they aren't the same thing. Keep both available to the interface so generated language cannot overwrite operational truth.
- Content state describes the answer: draft text, citations, code, attachments, partial sections, and revisions.
- System state describes the work: request accepted, stream active, tool queued, approval pending, cancellation requested, or operation complete.
- Provenance state describes where a claim or result came from and whether its source is available, current, or unresolved.
- Action state describes what the user may safely do now, including whether retry could duplicate an external effect.
Suppose an assistant writes, "The invoice has been created," then opens an approval request for the tool that would create it. The content claims completion while the system is still approval-pending. The interface should follow the execution record: label the action as proposed, keep approval controls attached to the exact scope, and prevent the prose from becoming the status source.
The same rule applies to partial success. If three files were updated and a fourth failed, preserve the successful results, identify the failed scope, and allow a targeted retry only when the operation supports it. A generic error banner discards the information needed for safe recovery.
Define component contracts without freezing one visual treatment
Once the lifecycle is explicit, components can represent it. Define each component by its responsibility, inputs, outputs, states, and protected behavior. A particular spinner, card shape, or animation shouldn't become the behavioral contract.
- Message: authorship, content status, timestamps where relevant, edits, attachments, and per-message actions.
- Provenance: source identity, availability, freshness or retrieval condition, and the relationship between a source and a claim.
- Tool activity: operation identity, scope, status, meaningful progress evidence, result boundary, and failure detail.
- Approval: proposed action, affected target, consequence, expiry or revalidation rule, and approve, reject, or revise outcomes.
- Destructive confirmation: the exact irreversible effect, resolved target, safeguards, and evidence that authorization was recorded.
- Interruption: stop request, acknowledgement, retained partial output, operations that may continue, and supported next actions.
- Retry: retry scope, idempotency assumptions, duplicate-effect risk, attempt status, and what changes before another attempt.
- Reconnection: last known server state, reconciliation method, stale-state treatment, and whether the user must decide how to continue.
- Recovery: preserved inputs and outputs, the valid restart point, responsible owner, and the condition that closes the incident.
The same contract may appear as inline status, a timeline, a compact activity row, or a detailed panel. Desktop, mobile, embedded assistants, and agent workspaces can use different treatments while preserving the same behavioral rules.
Give the visual layer a verified upstream source
After writing the chat-state contract, browse published kits for semantic colors, typography, spacing, modes, and implementation guidance. Keep approval, persistence, interruption, and recovery behavior in project-owned contracts.
Keep visual inputs upstream of chat behavior
A visual design system should own reusable appearance decisions such as semantic light and dark color roles, typography, spacing, layout guidance, motifs, and focus treatment. The AI-chat project should own lifecycle rules, component behavior, authorization, persistence, provenance, and recovery.
Identity Forge can supply the upstream visual layer through complete kits, DESIGN.md guidance, and implementation exports. It doesn't provide an AI-chat component library, decide what an approval authorizes, validate model output, or approve a product release. That boundary prevents a visual artifact from being accepted as behavioral proof.
Token specimen · real values
Ambient Sage
Live renderAmbient Sage's actual tokens — the same values its exports use.
The published Ambient Sage kit is a concrete visual input rather than a hypothetical theme. Its public page documents a warm-sage palette, Plus Jakarta Sans for general typography, and JetBrains Mono for technical strings. Those facts can guide the visual implementation. They don't determine what streaming means, when a retry is safe, or whether an approval remains valid.
| Upstream visual system | Project-owned chat system | |
|---|---|---|
| Color and modes | Semantic light and dark roles | Meaning of pending, approved, interrupted, and failed in product context |
| Typography | Families, roles, weights, scale, and usage guidance | Treatment of streamed content, code, provenance, and operational status |
| Spacing and layout | Reusable rhythm and layout guidance | Composer behavior, activity placement, mobile adaptation, and transcript structure |
| Components | Visual foundations and shared primitive treatment | Lifecycle, actions, persistence, authorization, and recovery logic |
| Acceptance | Artifact and visual-consumer evidence | Runtime observations across named chat states and transitions |
Trace a hypothetical interrupted stream
This example is intentionally hypothetical. It shows how to structure the record. It is not a tested Identity Forge implementation, usability result, browser observation, or product benchmark.
NAMED CLAIM
The chat preserves a useful partial response when a user stops streaming.
DECLARED
Source version: chat-contract v0.4
Rule: During streaming, Stop requests interruption. Received text remains visible
and is labelled "Stopped" after server acknowledgement.
Permitted next actions: continue in a new turn; regenerate from the original request.
Protected surfaces: original user message, received text, cited sources.
DELIVERED
Artifact: state-contract.json, hypothetical build candidate 184
State IDs: streaming, interruption_requested, interrupted
Required event: stream.stop_acknowledged
MAPPED
Consumer: assistant response component
Status region: response footer
Stop control: composer action slot
Persistence: transcript store is expected to retain received chunks and terminal state
COMPONENT CONTRACT
On stop request: disable repeated Stop, expose the pending request, await acknowledgement.
On acknowledgement: mark response interrupted, retain received content, expose valid recovery.
On timeout: show unresolved interruption; do not claim the operation stopped.
EXPECTED OBSERVATION
Given an active stream with received text, when Stop is requested and acknowledged,
the text remains, status becomes Stopped, focus follows the project's interaction
specification, and no later chunks are appended.
OBSERVED
Unresolved. No runtime observation is supplied in this hypothetical example.
CORRECTION OWNER
Assign only after locating the first divergent layer.
RETEST TRIGGER
Any change to stream events, transcript persistence, response status, Stop behavior,
reconnection logic, or the affected visual-system roles.This record supports only a narrow claim. A Stop button doesn't prove interruption works, and a client request doesn't prove the server stopped. The record also doesn't establish that the flow is accessible, safe, or ready to release. Those claims remain open until the required observations exist.
Use a complete acceptance record
Accept one named claim at a time. Bind the record to a source version, delivered artifact, consumer, state or transition, conditions, and observations. This stops evidence from one fixture or build from quietly becoming a claim about the whole product.
AI-CHAT ACCEPTANCE RECORD
Claim ID:
Claim:
Scope:
Out of scope:
Declared evidence
- Source and version:
- State or transition:
- Entry and exit conditions:
- Visible evidence:
- Permitted and blocked actions:
- Persistence rule:
- Provenance or uncertainty rule:
- Recovery rule:
- Protected surfaces:
- Owner:
- Disposition: required | conditional | unresolved | not_applicable
Delivered evidence
- Artifact and version:
- Relevant identifiers or events:
- Generated or distributed output inspected:
- Disposition:
Mapped evidence
- Project mapping:
- Component contract:
- Approved exceptions:
- Named consumers:
- Disposition:
Observed evidence
- Build or release identifier:
- Consumer:
- Mode and viewport:
- Input and content condition:
- Network, latency, and tool condition:
- Keyboard, motion, and assistive-technology condition:
- Expected result:
- Actual result:
- Evidence location:
- Disposition:
Open items
- Item:
- Owner:
- Required before:
- Retest trigger:
Decision: accept | revise | block
Decision rationale:
Decision owner:
Decision date:Choose accept, revise, or block
- Accept when every required field for the named scope is supported, conditional evidence has been resolved for the tested condition, exceptions are approved, and observations match expectations.
- Revise when the contract is sound but a bounded artifact, mapping, component, or consumer mismatch has a clear correction owner and doesn't invalidate the test itself.
- Block when required evidence is missing, authorization or destructive-action boundaries are ambiguous, the observed state contradicts the declared contract, or recovery could duplicate or conceal an external effect.
Not applicable is still a claim
Record why a field doesn't apply. Provenance may be outside the scope of a purely generative brainstorming response, for example, but it may become required when the same component presents retrieved facts.
Verify representative consumers and conditions
Choose conditions likely to expose a false assumption. The verification set doesn't need to test every transcript, but it must name the consumers and conditions that support the acceptance claim.
- Light and dark modes, including status roles that must remain distinguishable without relying on color alone
- Short and long responses, code blocks, tables, citations, attachments, and content that wraps unpredictably
- Desktop and constrained mobile layouts, including the composer, approval controls, and tool activity
- Keyboard focus order, focus visibility, status announcements, and supported assistive-technology paths
- Reduced-motion conditions for streaming indicators, progress, and transitions
- Normal, delayed, and disconnected streaming, including reconnection with stale local state
- Tool success, tool failure, delayed results, and results that arrive after interruption
- Approval accepted, rejected, revised, expired, and disconnected before the decision is recorded
- Cancellation requested before and after an external effect may have begun
- Partial completion where retrying the entire operation could duplicate successful work
For each observation, record the version, consumer, mode, viewport, content, network or tool condition, expected result, actual result, and evidence location. A passing desktop light-mode stream doesn't transfer automatically to mobile dark mode, approval pending, or reconnection.
Accessibility evidence follows the same boundary. A component fixture can show that one control has a visible focus state under one setup, but it cannot establish whole-product accessibility conformance. Test the shared component, then retest representative product consumers with their real composition and content.
Route mismatches to the first divergent layer
When the interface disagrees with the contract, inspect the chain in ownership order. Correct the first layer that differs from the accepted decision. Patching the final screen may hide the symptom while other consumers remain wrong.
- 1
Check the upstream decision
Is the intended state, action, persistence rule, or visual role explicit, current, approved, and owned? If it's missing or contradictory, resolve it before changing the downstream implementation.
- 2
Check the emitted artifact
Does the reviewed source version produce the expected identifiers, events, tokens, or guidance? If the source is correct, a stale or incomplete artifact may be the first divergence.
- 3
Check the project mapping
Does the application map the artifact to the intended runtime event, state store, semantic role, and consumer? Look for stale aliases and locally duplicated values.
- 4
Check the component contract
Does the component implement the permitted actions, status evidence, persistence boundary, and recovery behavior? A visually correct component can still violate the lifecycle.
- 5
Check approved exceptions
Is the difference deliberate, scoped, current, and approved? An undocumented exception is a mismatch, not a design decision.
- 6
Check the rendered consumer
Does composition, content, routing, viewport, mode, latency, or a local override change the result? Record the precise condition before assigning ownership.
Prompt changes belong in this diagnosis only when the prompt is the first layer that owns the divergent decision. Repeatedly telling a coding agent to make a state clearer can't replace a defined state, the correct artifact, and a mapping into the component contract.
Write the highest-risk transition before another screen
Choose the transition with the greatest consequence if the interface lies: approval to execution, streaming to interruption, tool failure to retry, or partial completion to recovery. Fill in one state contract and one acceptance record for a named consumer. Don't accept it until you have observed the visible status, safe actions, persistence, and recovery path under the recorded condition.
Sources
- Designing AI chat interfaces: Anatomy, patterns, pitfalls: The captured page treats AI chat as a distinct interface category with its own anatomy, controls, states, and uncertainty concerns.
- Agentic Design System - From Chatbot to Orchestration: The captured article argues that agent-facing components need machine-readable intent, constraints, validation, and human oversight rather than visual definitions alone.
- AI Design Systems: A Practical Guide for Designers (2026): The captured guide describes unpredictable output, drift, explicit constraints, and human review as central concerns in AI-assisted design-system work.
- What is chatbot design?: IBM distinguishes chatbot UI from UX and recommends planning conversation paths that include confusion, detours, dead ends, and human escalation.
- Creating a chatbot UI Design System.: The captured case study organizes reusable chatbot patterns around recurring flows, errors, timeouts, dismissal, help, and human handoff.
- Design Systems with AI: The captured practitioner article presents design systems as guardrails for generated interfaces and emphasizes testing before treating generated work as ready.
- Identity Forge: Identity Forge publicly describes a pipeline that supplies visual design-system decisions and agent-ready implementation artifacts.
- Identity Forge kit gallery: The public gallery presents visual kits with fonts, colors, tokens, and component rules for use by coding agents.
- Ambient Sage Design Kit: The public Ambient Sage page documents a real kit with a warm-sage palette, Plus Jakarta Sans typography, JetBrains Mono for technical strings, and implementation-oriented visual guidance.
- AI UI review checklist: The Identity Forge guide separates task completeness, design-system conformance, resilience, and accessibility evidence rather than treating visual polish as release proof.
- Design system accessibility checklist: test the system, then retest the product: The Identity Forge checklist explains that documented intent and isolated component tests do not establish accessibility conformance in representative product consumers.