---
title: Streaming & defer
description: Answer immediately with the cheap parts of a page and stream the expensive parts into the same response with Defer.
order: 8
---

# Streaming & defer

Use `<Defer>` when one slow region would hold back an otherwise cheap page. The shell answers with
ordinary JSX fallback content, then Kovo streams the real region later in the same response and
morphs it into place.

Kovo has three different streaming surfaces:

- `respond.stream()` and raw `endpoint()` responses are app-owned protocols for downloads, webhooks,
  or custom integrations. CSV/TSV/spreadsheet exports stay in that app-owned bucket; Kovo does not
  present them as a safe-by-default framework lane. They do not run the enhanced mutation apply
  path.
- `<Defer>` is first-render streaming. It replaces a fallback inside the document response.
- Streaming mutations are post-submit streams. They keep the normal mutation lifecycle and stream
  Kovo wire chunks back through the enhanced form response.

## The app shape

Author the boundary with the public JSX primitive from `@kovojs/server`:

```tsx
import { Defer } from '@kovojs/server';

<Defer
  target="product-grid"
  priority="after-paint"
  fallback={<section aria-busy="true">Loading products...</section>}
  render={() => <ProductGrid />}
/>;
```

In a page, that boundary sits beside the cheap content that can render immediately:

```tsx
import { component } from '@kovojs/core';
import { Defer, route } from '@kovojs/server';

export const ProductGrid = component({
  queries: { productGrid },
  render: ({ productGrid }: { productGrid: { items: unknown[] } }) => (
    <section>{productGrid.items.map(renderProduct)}</section>
  ),
});

export const productPage = route('/products', {
  page: () => (
    <main>
      <h1>Products</h1>
      <Defer
        target="product-grid"
        priority="after-paint"
        fallback={<section aria-busy="true">Loading products...</section>}
        render={() => <ProductGrid />}
      />
    </main>
  ),
  stylesheets: commerceStylesheets,
});
```

`fallback` and `render` are normal JSX, so text and attributes go through the same escaping rules as
the rest of the page. The response is one chunked HTML document. Internally, Kovo emits the owned
wire placeholder and the later fragment:

```html
<!doctype html>
<html>
  <body>
    <main class="product-page">
      <kovo-defer target="product-grid" state="pending"></kovo-defer>
      --kovo-boundary
      <kovo-query name="productGrid">{"items":[…],"nextCursor":"p2"}</kovo-query>
      <kovo-fragment target="product-grid"
        ><link rel="stylesheet" href="/assets/site.css" />
        <section kovo-c="product-grid" kovo-deps="product">…</section>
      </kovo-fragment>
      --kovo-boundary--
    </main>
  </body>
</html>
```

The vocabulary is the mutation response's — `<kovo-query>` then `<kovo-fragment>` — arriving during first
render instead of after a POST. It reads top to bottom in view-source, like everything else on the
wire.

## Run it

Use `curl -N` against a route with `priority="after-paint"`:

```sh
curl -N http://localhost:3000/products
```

You should see the shell and `<kovo-defer ... state="pending">` frame first, then the later
`<kovo-query>` and `<kovo-fragment target="product-grid">` chunk. Change the boundary to
`priority="critical"` and that split disappears because the region renders inline with the shell.

## Pick a priority

`priority` controls whether a region renders in the shell or streams later:

- `critical` is the default. It renders inline with the shell, so it does **not** defer.
- `after-paint` streams after the shell has painted.
- `visible` waits until the region scrolls into view, then streams through the shared
  IntersectionObserver-gated path.

If the smallest example is meant to defer, say `priority="after-paint"` explicitly.

## Replace the fallback

The framework-emitted `<kovo-defer>` element carries the fallback you authored on `<Defer>`. A
skeleton, spinner, or static summary paints with the shell at first byte. When the matching
`<kovo-fragment target="…">` chunk arrives, the morph layer patches it in. Because it morphs rather
than replaces, the swap preserves focus, scroll position, selection, CSS transitions, and the state
of any islands nested in the fallback. Patched-in islands are inert-until-touched like everything
else, and `on:visible` observers attach to them normally.

`mode="append"` is available on deferred fragments as the explicit append vocabulary, the same as
mutation fragments — useful for streaming list pages.

## Keep deferred queries ahead of their consumers

The guarantee to rely on: deferred query JSON arrives before or with its consumers. That's why
Kovo emits a fragment's query values in the same chunk as the fragment whose `data-bind` attributes
read them, so a deferred island never renders against missing data.

Priority is declared on the route/component surface that owns the late region. Finer priority
semantics and query-JSON placement under HTTP/1.1 fallbacks are still open design areas. The
before-or-with guarantee is the contract you can depend on.

## Stylesheets for late fragments

A deferred fragment may use StyleX atoms or document CSS the shell never referenced. Fragment chunks
declare their stylesheets, and the links ride inside the framework-emitted fragment — present before
the content paints, deduped by `href` within the response:

```html
<kovo-fragment target="product-grid">
  <link rel="stylesheet" href="/assets/site.css" />
  <section kovo-c="product-grid">...</section>
</kovo-fragment>
```

The same Kovo stylesheet contract applies as everywhere: StyleX rules are extracted from source at
build time, document CSS is shipped as a declared asset, and the fragment lists the stylesheet it
needs. See [styling with StyleX](/guides/styling/).

## The client side

On a server-rendered stream the inline loader handles this. A custom shell installs the same
document-scoped runtime with one experimental API:

```ts
import { installKovoClient } from '@kovojs/browser/client';

const client = installKovoClient({
  importModule: (url) => import(url),
  root: document,
});
await client.ready;
```

Each applied chunk behaves exactly like a mutation response landing: query values update their
bindings, fragments morph into their targets. Kovo owns the store, morph adapter, request settings,
and module allowlist. When the shell goes away, call `await client.dispose()`.

## When to reach for it

Projected children all ship in the initial HTML — every tab panel, dialog body, accordion content.
There's no client-side lazy mount. So the question `<Defer>` answers is about server render cost at
first paint, not payload size.

**Use it when** a subtree is expensive to produce and the rest of the page isn't: recommendations
behind a slow model, an analytics panel aggregating wide tables, third-party-data sections with
unpredictable latency. The cheap 95% of the page paints at first byte, and the slow section streams in
seconds later with no client round-trip.

**Don't use it for:**

- _Big-but-cheap subtrees._ Streaming reorders HTML; it doesn't shrink it. A long static page is fine
  as a long static page.
- _Below-the-fold JS deferral._ That's `on:visible`, which defers executing JavaScript, not HTML.
- _Data that updates after load._ That's a query with refetch or a mutation response. Defer is a
  first-render mechanism only.
- _Chat token rendering._ That's a streaming mutation response: one enhanced form POST whose chunks
  append fragments, append escaped text, and then reconcile to server truth.
- _Navigation._ Pages are complete documents; defer streams within one response, it doesn't splice
  between pages.

A reasonable default: render everything inline until a route's server time is dominated by one
identifiable subtree, then defer exactly that subtree. The wire stays readable either way, and
`kovo explain page` keeps listing the route's queries — deferred or not — as one surface.

## Streaming mutation responses

Use a streaming mutation when the user has submitted a real form and the response should render
progressively, such as a chat assistant answer. It is not an SSE subscription and it is not
`<Defer>`; it is one enhanced mutation POST response. The server still runs CSRF, input schema
validation, guards, replay/idempotency, and the mutation transaction before user-visible assistant
chunks are emitted.

Keep one `Kovo-Idem` token on one delivery path. Kovo binds the request's normalized, negotiated path
to buffered or stream vocabulary; typed failures still use a buffered body. A retry that switches
paths returns a generic 422 conflict without rerunning the mutation or exposing the earlier response
body. A same-path successful streaming retry replays the complete settled response, including
`<kovo-done>`. `Kovo-Stream` on responses is framework-owned — application result values and response
hooks cannot add, remove, or override it.

The wire remains the mutation vocabulary plus one narrow text-source primitive:

```html
<kovo-fragment target="messages" mode="append">
  <article class="message user">How do I ship this?</article>
</kovo-fragment>

<kovo-fragment target="messages" mode="append">
  <article class="message assistant">
    <div data-stream-text="assistant:a1" aria-live="polite" aria-atomic="true"></div>
  </article>
</kovo-fragment>

<kovo-text target="assistant:a1">Start with the typed mutation path.</kovo-text>
<kovo-text target="assistant:a1" mode="checkpoint">Start with the typed mutation path.</kovo-text>

<kovo-fragment target="messages">
  <section kovo-fragment-target="messages">...canonical server-rendered messages...</section>
</kovo-fragment>
```

`<kovo-text>` appends escaped text to a declared `data-stream-text` source. HTML-looking model output
stays text. If the UI wants Markdown, citations, tables, or code highlighting while the answer is
arriving, declare a renderer for the source and let that app-owned renderer transform the accumulated
text into presentation. Kovo still owns the source buffer and the final server-rendered fragment or
query reconciliation.

Use `mode="checkpoint"` on long streams when the server wants to replace the accumulated source text
with canonical text so far. Use a final `<kovo-fragment>` or `<kovo-query>` update to reconcile the
message or message list with server truth. A partial text stream is never the final authority.

For accessible chat shells, put live-region semantics on the assistant message container or source
element, not on each token. Prefer `aria-live="polite"` and a stable status/message element so screen
readers receive coalesced updates instead of one announcement per token. Keep the submitted user
message and assistant shell as ordinary fragments, so no-JS users still get the normal mutation
fallback and the final page remains meaningful.

If the stream aborts, validation fails, a guard/session check fails, a stream target is missing, or a
deploy build token is stale, the runtime must recover to server truth: mark the submitted UI failed,
refetch the affected target, or navigate through the normal form path. Do not leave a partial answer
presented as confirmed.

## Degradation

The stream is one HTML response, so the no-JS story degrades the way the rest of the framework does:
the document still arrives and completes, and the fallback content is what a non-JS visitor keeps for
deferred regions. Keep fallbacks honest — a meaningful placeholder or summary, not an empty box — for
the same reason the no-JS form path stays a real form.

## Handle failure

Two failure cases should be obvious to the reader:

- If the deferred work runs too long, put a real timeout on the boundary so the fallback can switch
  to an error state instead of hanging forever.
- If a late render fails, the placeholder should re-emit as `state="error"` with copy that tells the
  user what failed and what they can retry.

For example:

```tsx
import { Defer } from '@kovojs/server';

<Defer
  target="product-grid"
  timeoutMs={30_000}
  fallback={<section aria-busy="true">Loading products...</section>}
  render={() => <ProductGrid />}
/>;
```

The user-facing rule is simple: a partial stream may stay partial briefly, but it must resolve to
server truth or to a visible failure state.

## Next

- [Styling with StyleX](/guides/styling/) — the stylesheet contract these chunks use.
- [Mutations](/guides/mutations/) — the typed form lifecycle streaming mutations preserve.
- [Queries & invalidation](/guides/queries/) — the query values deferred chunks deliver.
- [Stability & Versioning](/getting-started/stability/#import-boundaries) — when a custom shell should import
  `@kovojs/browser/client`.

<details>
<summary>Spec & diagnostics</summary>

Defer as a first-render reuse of the fragment protocol, morph survival, the before-or-with ordering
guarantee, and no-JS degradation: SPEC §8. The shared `<kovo-query>`/`<kovo-fragment>` wire vocabulary,
`mode="append"`, streaming mutation responses, `<kovo-text>`, checkpoints, interruption behavior, and
server-truth reconciliation: SPEC §9.1. `respond.stream()` as an app-owned escape hatch: SPEC §6.4.
`on:visible` and inert-until-touched islands: SPEC §4.7. Projected children shipping in initial HTML:
SPEC §4.5. Priority hinting and open stream-ordering areas: SPEC §13.3. Stylesheets for late
fragments: SPEC §13.1. Defer vs. post-load data updates: SPEC §9.3. `kovo explain page` as one
surface: SPEC §5.3. App-authored `defer(...)` as a JSX child is **KV244**; author `<Defer>` and let
Kovo emit `<kovo-defer>`.

API reference: [@kovojs/browser](/api/browser/), [@kovojs/core](/api/core/), [@kovojs/server](/api/server/).

</details>
