Get started

Design system for AI chat apps: define states, actions, and recovery

An AI chat design system needs more than message bubbles, a composer, and visual tokens. It needs a contract that identifies the active state, the evidence the interface must show, what the user may safely do, what persists, and how the conversation recovers. Without that contract, fluent text can hide a pending tool, a failed action, or a request that still needs approval.

Updated October 5, 2026

Set the evidence boundary first

This guide covers interface behavior: how an AI chat product represents conversation state, constrains actions, preserves information, asks for approval, exposes uncertainty, and recovers from disruption. It also explains how to verify that a named implementation follows those rules.

The boundary matters. An interface can accurately display a completed state while the underlying answer is wrong. It can meet its component contract while violating a safety policy. It may look accessible in one fixture but fail with a keyboard, assistive technology, zoom, or real content. Each claim needs its own evidence.

ClaimRelevant evidenceWhat it does not prove
Interface state is represented correctlyThe named state appears as declaredState contract plus an observation under a recorded conditionModel accuracy or task success
Visual system is appliedThe named consumer follows the visual systemVersioned tokens, mappings, component use, and rendered inspectionCorrect chat behavior or release readiness
Interaction is accessibleThe tested interaction meets its recorded expectationAccessibility tests in named components and representative consumersUniversal conformance outside the tested scope
Action is safe to executeThe named action satisfies its safety contractProduct safety policy, authorization rules, and execution controlsThat a polished confirmation dialog is sufficient
Product is ready to releaseThe named release meets all applicable gatesCombined product, safety, accessibility, operational, and business evidenceAny isolated screenshot, story, or token-file check
Keep each acceptance claim within the evidence that can support it.

Never infer completion from prose

A model can produce a confident sentence before a tool finishes, after it fails, or while approval remains unresolved. Treat generated content and execution state as separate evidence channels.

Map the conversation lifecycle

Start with transitions, not a component inventory. A message component is easy to draw. The harder question is what must change as a request moves from local composition to submission, streaming, tool execution, approval, and recovery.

  1. 1

    Idle

    No draft or operation is active. The composer is available, and any conversation-level status is current.

  2. 2

    Composing

    The user has an unsent draft. Preserve it across harmless navigation and recoverable connection changes according to the product's stated persistence policy.

  3. 3

    Submitting

    The request has left the local composer, but the system hasn't established a response stream. Prevent accidental duplicate submission or make the deduplication rule explicit.

  4. 4

    Streaming

    Assistant content is arriving. Show that the answer is provisional, expose any supported interruption action, and preserve the user request that caused the stream.

  5. 5

    Waiting for a tool

    The assistant can't continue until an external operation returns. Distinguish queued, running, delayed, and failed activity when those differences affect user action.

  6. 6

    Waiting for approval

    A proposed action requires a person to approve, reject, or revise it. Identify the exact action and scope. Continued prose must not imply that approval was granted.

  7. 7

    Completed

    The in-scope response or operation reached its declared terminal condition. System evidence must support completion; tone is not evidence.

  8. 8

    Interrupted

    Generation or execution stopped before its intended terminal condition. Preserve useful partial output, label its status, and offer only recovery actions the system can honor.

  9. 9

    Cancelled

    The user or system intentionally ended the operation. Record whether cancellation was requested, acknowledged, and effective, especially when an external tool may already have acted.

  10. 10

    Partially completed

    Some declared work succeeded and some didn't. Name the completed and unresolved parts instead of collapsing the result into success or failure.

  11. 11

    Failed

    The operation can't continue without a retry, changed input, restored dependency, or human intervention. Preserve the request and enough diagnostic context for the supported recovery path.

Treat these states as a starting contract, not a universal finite-state machine. A research assistant, coding agent, support bot, and transaction agent need different substates. Add a state when it changes visible evidence, permitted actions, persistence, ownership, or recovery. If it changes none of those, it's probably an implementation detail.

Copy the state-contract matrix

Use one record for each state or high-risk transition. A compact matrix helps the team spot omissions, while the full record keeps every decision inspectable.

STATE CONTRACT

State: <stable name>
Entry condition: <event and prior state>
Exit conditions: <allowed next states>
Visible evidence: <status, labels, timestamps, progress, partial output>
Permitted user actions: <send, stop, edit, approve, reject, retry, resume>
Blocked actions: <actions hidden or disabled, with reason>
Persistence: <draft, messages, tool results, approvals, partial output>
Provenance or uncertainty: <source, freshness, confidence limits, unresolved status>
Recovery: <automatic and user-triggered routes>
Protected surfaces: <content or controls that must not change>
Owner: <role that owns this contract>
Required observations: <conditions and expected results>
Open questions: <unresolved decisions with owners>
Evidence disposition: required | conditional | unresolved | not_applicable
Copyable record for one AI-chat state

The four evidence dispositions stop an empty field from passing as a decision. Required means acceptance depends on it. Conditional means it becomes required under a named condition. Unresolved marks a missing decision or observation. Not applicable needs a reason because it's a scoped conclusion, not a shortcut.

Test contentInspection criteria and failure condition
Idle or composingEmpty composer, long draft, and an attachment if supportedInspect readiness, draft retention, and available actions. Fail if stale status remains or a promised draft disappears without disclosure.
SubmittingOne request plus a repeated send attemptInspect pending evidence and duplicate handling. Fail if both requests can create the same unintended effect or the user cannot tell whether submission was accepted.
StreamingLong prose, code, citations, and an active Stop actionInspect provisional status, partial-output retention, and interruption controls. Fail if prose appears terminal while execution remains active.
Tool runningA delayed operation with a stable invocation identityInspect operation name, scope, status, and safe actions. Fail if the UI hides whether the tool is queued, active, delayed, or failed when that difference affects recovery.
Approval pendingA concrete action, resolved target, consequence, and stale-approval conditionInspect scope, available decisions, persistence, and revalidation. Fail if approval can apply to an ambiguous or changed target.
Interrupted or cancelledReceived partial output and a stop or cancel requestInspect request, acknowledgement, retained content, and any work that may continue. Fail if a requested stop is displayed as effective before acknowledgement.
Partially completedSeveral sub-operations with at least one success and one failureInspect successful and unresolved scopes separately. Fail if retrying can duplicate successful work without warning or control.
FailedA recoverable dependency failure and a non-recoverable caseInspect usable explanation, preserved input, and valid next actions. Fail if the interface offers a retry that cannot work under the recorded condition.
CompletedA response whose terminal condition is independently availableInspect terminal system evidence and retained provenance. Fail if completion is inferred only from generated wording.
A practical review matrix. These rows are requirements to inspect, not claims that any particular product has passed.

Review transitions as well as states

The most dangerous ambiguity often sits between rows: stop requested versus stop acknowledged, approval shown versus approval recorded, or reconnect started versus the operation actually resumed.

Separate content state from system state

The transcript holds the content. The execution lifecycle holds the system state. They often move together, but they aren't the same thing. Keep both available to the interface so generated language cannot overwrite operational truth.

  • Content state describes the answer: draft text, citations, code, attachments, partial sections, and revisions.
  • System state describes the work: request accepted, stream active, tool queued, approval pending, cancellation requested, or operation complete.
  • Provenance state describes where a claim or result came from and whether its source is available, current, or unresolved.
  • Action state describes what the user may safely do now, including whether retry could duplicate an external effect.

Suppose an assistant writes, "The invoice has been created," then opens an approval request for the tool that would create it. The content claims completion while the system is still approval-pending. The interface should follow the execution record: label the action as proposed, keep approval controls attached to the exact scope, and prevent the prose from becoming the status source.

The same rule applies to partial success. If three files were updated and a fourth failed, preserve the successful results, identify the failed scope, and allow a targeted retry only when the operation supports it. A generic error banner discards the information needed for safe recovery.

Define component contracts without freezing one visual treatment

Once the lifecycle is explicit, components can represent it. Define each component by its responsibility, inputs, outputs, states, and protected behavior. A particular spinner, card shape, or animation shouldn't become the behavioral contract.

  • Message: authorship, content status, timestamps where relevant, edits, attachments, and per-message actions.
  • Provenance: source identity, availability, freshness or retrieval condition, and the relationship between a source and a claim.
  • Tool activity: operation identity, scope, status, meaningful progress evidence, result boundary, and failure detail.
  • Approval: proposed action, affected target, consequence, expiry or revalidation rule, and approve, reject, or revise outcomes.
  • Destructive confirmation: the exact irreversible effect, resolved target, safeguards, and evidence that authorization was recorded.
  • Interruption: stop request, acknowledgement, retained partial output, operations that may continue, and supported next actions.
  • Retry: retry scope, idempotency assumptions, duplicate-effect risk, attempt status, and what changes before another attempt.
  • Reconnection: last known server state, reconciliation method, stale-state treatment, and whether the user must decide how to continue.
  • Recovery: preserved inputs and outputs, the valid restart point, responsible owner, and the condition that closes the incident.

The same contract may appear as inline status, a timeline, a compact activity row, or a detailed panel. Desktop, mobile, embedded assistants, and agent workspaces can use different treatments while preserving the same behavioral rules.

Give the visual layer a verified upstream source

After writing the chat-state contract, browse published kits for semantic colors, typography, spacing, modes, and implementation guidance. Keep approval, persistence, interruption, and recovery behavior in project-owned contracts.

Keep visual inputs upstream of chat behavior

A visual design system should own reusable appearance decisions such as semantic light and dark color roles, typography, spacing, layout guidance, motifs, and focus treatment. The AI-chat project should own lifecycle rules, component behavior, authorization, persistence, provenance, and recovery.

Identity Forge can supply the upstream visual layer through complete kits, DESIGN.md guidance, and implementation exports. It doesn't provide an AI-chat component library, decide what an approval authorizes, validate model output, or approve a product release. That boundary prevents a visual artifact from being accepted as behavioral proof.

Token specimen · real values

Ambient Sage

Live render

Ambient Sage's actual tokens — the same values its exports use.

Color tokensSemantic roles with HEX / HSL / CMYK

Color tokens

Ambient Sage
light · HEX · HSL · CMYK

Core

#F3F4EF

background

H 72 · C0, 0, 2, 4

#1A1C17

foreground

H 84 · C7, 0, 18, 89

#E5E6E0

card

H 70 · C0, 0, 3, 10

#ECEEE8

muted

H 80 · C1, 0, 3, 7

#D8D9D2

border

H 68.57 · C0, 0, 3, 15

Brand

#FEE951

primary

H 52.72 · C0, 8, 68, 0

#1A1C17

primary-fg

H 84 · C7, 0, 18, 89

#E5E6E0

secondary

H 70 · C0, 0, 3, 10

#F7E464

accent

H 52.24 · C0, 8, 60, 3

#FEE951

ring

H 52.72 · C0, 8, 68, 0

Semantic

#C0392B

destructive

H 5.64 · C0, 70, 78, 25

#FFFFFF

destructive-fg

H 0 · C0, 0, 0, 0

#2D7238

success

H 129.57 · C61, 0, 51, 55

#C97D12

warning

H 35.08 · C0, 38, 91, 21

#545651

muted-fg

H 84 · C2, 0, 6, 66

Charts

#FEE951

chart-1

H 52.72 · C0, 8, 68, 0

#4A8FD4

chart-2

H 210 · C65, 33, 0, 17

#6BBF8A

chart-3

H 142.14 · C44, 0, 28, 25

#E07498

chart-4

H 340 · C0, 48, 32, 12

#E8A24B

chart-5

H 33.25 · C0, 30, 68, 9

Type scaleHeading, body, and mono in the kit's fonts

Typography

Ambient Sage

Scale: compact-product

Density: balanced

Heading · Plus Jakarta Sans · 1.875rem

Ship beautiful product faster

Subheading · Plus Jakarta Sans · 1.375rem

A warm-sage neutral-surface mobile kit with a single vivid yellow accent, flat tonal cards, and oversized display numerals.

Body · Plus Jakarta Sans · 1rem

Ambient Sage uses a near-white warm-sage canvas (#f3f4ef) with card panels distinguished only by a tonal shift to #e5e6e0, never by shadows or borders. A single vivid yellow (#fee951) is the only saturated color and appears sparingly at component scale as orbs, button fills, and focus rings. Primary data values render as oversized bold hero numerals with a small superscript unit. Typography is a friendly rounded geometric (Plus Jakarta Sans) with no uppercase and no tight tracking, while JetBrains Mono is reserved for hex codes and technical strings. Generous rounding and luminance-only contrast give the whole system a calm, minimal feel.

Mono · JetBrains Mono · 0.8125rem

npx shadcn add ambientsage.json

Aa

Plus Jakarta Sans · Heading

400500600700

Aa

Plus Jakarta Sans · Body

400500600700

ABCDEFGHIJKLM NOPQRSTUVWXYZ

abcdefghijklmnopqrstuvwxyz

0123456789 & @ # % →

Radius & spacingCorner radius, elevation, and spacing steps

Tokens

Ambient Sage primitives
density: balanced

Radius scale

sm · 0.375rem
md · 0.75rem
lg · 1.25rem
xl · 1.75rem

Component radius

button
card
input

Elevation

level 1
level 2
level 3
level 4

Spacing · base 4px

1x
2x
3x
4x
6x
8x
Ambient Sage is a verified public kit used here only to show upstream visual inputs. This specimen is not evidence that the kit was tested with an AI-chat runtime, component contract, interruption flow, or accessibility setup.

The published Ambient Sage kit is a concrete visual input rather than a hypothetical theme. Its public page documents a warm-sage palette, Plus Jakarta Sans for general typography, and JetBrains Mono for technical strings. Those facts can guide the visual implementation. They don't determine what streaming means, when a retry is safe, or whether an approval remains valid.

Upstream visual systemProject-owned chat system
Color and modesSemantic light and dark rolesMeaning of pending, approved, interrupted, and failed in product context
TypographyFamilies, roles, weights, scale, and usage guidanceTreatment of streamed content, code, provenance, and operational status
Spacing and layoutReusable rhythm and layout guidanceComposer behavior, activity placement, mobile adaptation, and transcript structure
ComponentsVisual foundations and shared primitive treatmentLifecycle, actions, persistence, authorization, and recovery logic
AcceptanceArtifact and visual-consumer evidenceRuntime observations across named chat states and transitions
Place decisions in the layer that can enforce and verify them.

Trace a hypothetical interrupted stream

This example is intentionally hypothetical. It shows how to structure the record. It is not a tested Identity Forge implementation, usability result, browser observation, or product benchmark.

NAMED CLAIM
The chat preserves a useful partial response when a user stops streaming.

DECLARED
Source version: chat-contract v0.4
Rule: During streaming, Stop requests interruption. Received text remains visible
and is labelled "Stopped" after server acknowledgement.
Permitted next actions: continue in a new turn; regenerate from the original request.
Protected surfaces: original user message, received text, cited sources.

DELIVERED
Artifact: state-contract.json, hypothetical build candidate 184
State IDs: streaming, interruption_requested, interrupted
Required event: stream.stop_acknowledged

MAPPED
Consumer: assistant response component
Status region: response footer
Stop control: composer action slot
Persistence: transcript store is expected to retain received chunks and terminal state

COMPONENT CONTRACT
On stop request: disable repeated Stop, expose the pending request, await acknowledgement.
On acknowledgement: mark response interrupted, retain received content, expose valid recovery.
On timeout: show unresolved interruption; do not claim the operation stopped.

EXPECTED OBSERVATION
Given an active stream with received text, when Stop is requested and acknowledged,
the text remains, status becomes Stopped, focus follows the project's interaction
specification, and no later chunks are appended.

OBSERVED
Unresolved. No runtime observation is supplied in this hypothetical example.

CORRECTION OWNER
Assign only after locating the first divergent layer.

RETEST TRIGGER
Any change to stream events, transcript persistence, response status, Stop behavior,
reconnection logic, or the affected visual-system roles.
Hypothetical source-to-consumer record

This record supports only a narrow claim. A Stop button doesn't prove interruption works, and a client request doesn't prove the server stopped. The record also doesn't establish that the flow is accessible, safe, or ready to release. Those claims remain open until the required observations exist.

Use a complete acceptance record

Accept one named claim at a time. Bind the record to a source version, delivered artifact, consumer, state or transition, conditions, and observations. This stops evidence from one fixture or build from quietly becoming a claim about the whole product.

AI-CHAT ACCEPTANCE RECORD

Claim ID:
Claim:
Scope:
Out of scope:

Declared evidence
- Source and version:
- State or transition:
- Entry and exit conditions:
- Visible evidence:
- Permitted and blocked actions:
- Persistence rule:
- Provenance or uncertainty rule:
- Recovery rule:
- Protected surfaces:
- Owner:
- Disposition: required | conditional | unresolved | not_applicable

Delivered evidence
- Artifact and version:
- Relevant identifiers or events:
- Generated or distributed output inspected:
- Disposition:

Mapped evidence
- Project mapping:
- Component contract:
- Approved exceptions:
- Named consumers:
- Disposition:

Observed evidence
- Build or release identifier:
- Consumer:
- Mode and viewport:
- Input and content condition:
- Network, latency, and tool condition:
- Keyboard, motion, and assistive-technology condition:
- Expected result:
- Actual result:
- Evidence location:
- Disposition:

Open items
- Item:
- Owner:
- Required before:
- Retest trigger:

Decision: accept | revise | block
Decision rationale:
Decision owner:
Decision date:
Copyable acceptance record for one bounded implementation claim

Choose accept, revise, or block

  • Accept when every required field for the named scope is supported, conditional evidence has been resolved for the tested condition, exceptions are approved, and observations match expectations.
  • Revise when the contract is sound but a bounded artifact, mapping, component, or consumer mismatch has a clear correction owner and doesn't invalidate the test itself.
  • Block when required evidence is missing, authorization or destructive-action boundaries are ambiguous, the observed state contradicts the declared contract, or recovery could duplicate or conceal an external effect.

Not applicable is still a claim

Record why a field doesn't apply. Provenance may be outside the scope of a purely generative brainstorming response, for example, but it may become required when the same component presents retrieved facts.

Verify representative consumers and conditions

Choose conditions likely to expose a false assumption. The verification set doesn't need to test every transcript, but it must name the consumers and conditions that support the acceptance claim.

  • Light and dark modes, including status roles that must remain distinguishable without relying on color alone
  • Short and long responses, code blocks, tables, citations, attachments, and content that wraps unpredictably
  • Desktop and constrained mobile layouts, including the composer, approval controls, and tool activity
  • Keyboard focus order, focus visibility, status announcements, and supported assistive-technology paths
  • Reduced-motion conditions for streaming indicators, progress, and transitions
  • Normal, delayed, and disconnected streaming, including reconnection with stale local state
  • Tool success, tool failure, delayed results, and results that arrive after interruption
  • Approval accepted, rejected, revised, expired, and disconnected before the decision is recorded
  • Cancellation requested before and after an external effect may have begun
  • Partial completion where retrying the entire operation could duplicate successful work

For each observation, record the version, consumer, mode, viewport, content, network or tool condition, expected result, actual result, and evidence location. A passing desktop light-mode stream doesn't transfer automatically to mobile dark mode, approval pending, or reconnection.

Accessibility evidence follows the same boundary. A component fixture can show that one control has a visible focus state under one setup, but it cannot establish whole-product accessibility conformance. Test the shared component, then retest representative product consumers with their real composition and content.

Route mismatches to the first divergent layer

When the interface disagrees with the contract, inspect the chain in ownership order. Correct the first layer that differs from the accepted decision. Patching the final screen may hide the symptom while other consumers remain wrong.

  1. 1

    Check the upstream decision

    Is the intended state, action, persistence rule, or visual role explicit, current, approved, and owned? If it's missing or contradictory, resolve it before changing the downstream implementation.

  2. 2

    Check the emitted artifact

    Does the reviewed source version produce the expected identifiers, events, tokens, or guidance? If the source is correct, a stale or incomplete artifact may be the first divergence.

  3. 3

    Check the project mapping

    Does the application map the artifact to the intended runtime event, state store, semantic role, and consumer? Look for stale aliases and locally duplicated values.

  4. 4

    Check the component contract

    Does the component implement the permitted actions, status evidence, persistence boundary, and recovery behavior? A visually correct component can still violate the lifecycle.

  5. 5

    Check approved exceptions

    Is the difference deliberate, scoped, current, and approved? An undocumented exception is a mismatch, not a design decision.

  6. 6

    Check the rendered consumer

    Does composition, content, routing, viewport, mode, latency, or a local override change the result? Record the precise condition before assigning ownership.

Prompt changes belong in this diagnosis only when the prompt is the first layer that owns the divergent decision. Repeatedly telling a coding agent to make a state clearer can't replace a defined state, the correct artifact, and a mapping into the component contract.

Write the highest-risk transition before another screen

Choose the transition with the greatest consequence if the interface lies: approval to execution, streaming to interruption, tool failure to retry, or partial completion to recovery. Fill in one state contract and one acceptance record for a named consumer. Don't accept it until you have observed the visible status, safe actions, persistence, and recovery path under the recorded condition.

Sources

  • Designing AI chat interfaces: Anatomy, patterns, pitfalls: The captured page treats AI chat as a distinct interface category with its own anatomy, controls, states, and uncertainty concerns.
  • Agentic Design System - From Chatbot to Orchestration: The captured article argues that agent-facing components need machine-readable intent, constraints, validation, and human oversight rather than visual definitions alone.
  • AI Design Systems: A Practical Guide for Designers (2026): The captured guide describes unpredictable output, drift, explicit constraints, and human review as central concerns in AI-assisted design-system work.
  • What is chatbot design?: IBM distinguishes chatbot UI from UX and recommends planning conversation paths that include confusion, detours, dead ends, and human escalation.
  • Creating a chatbot UI Design System.: The captured case study organizes reusable chatbot patterns around recurring flows, errors, timeouts, dismissal, help, and human handoff.
  • Design Systems with AI: The captured practitioner article presents design systems as guardrails for generated interfaces and emphasizes testing before treating generated work as ready.
  • Identity Forge: Identity Forge publicly describes a pipeline that supplies visual design-system decisions and agent-ready implementation artifacts.
  • Identity Forge kit gallery: The public gallery presents visual kits with fonts, colors, tokens, and component rules for use by coding agents.
  • Ambient Sage Design Kit: The public Ambient Sage page documents a real kit with a warm-sage palette, Plus Jakarta Sans typography, JetBrains Mono for technical strings, and implementation-oriented visual guidance.
  • AI UI review checklist: The Identity Forge guide separates task completeness, design-system conformance, resilience, and accessibility evidence rather than treating visual polish as release proof.
  • Design system accessibility checklist: test the system, then retest the product: The Identity Forge checklist explains that documented intent and isolated component tests do not establish accessibility conformance in representative product consumers.