---
name: create-surfaces
description: Create, revise, debug, or validate Traces surfaces: single-file HTML visualizations that run in the sandboxed trace-page iframe through window.traces.getTrace(). Use when building a new surface, changing its layout or interaction, diagnosing loading/rendering, or checking an artifact before publication.
---

# Prompt: build a Traces surface

You are writing a **Traces surface**: a single self-contained HTML file that renders a view of one AI agent trace (a recorded coding-agent session: messages, tool calls, thinking, token usage).

Traces stores your file as an immutable version, serves it from its own storage, and runs it inside a locked-down iframe on the trace page. The user opens it as a tab next to "Highlights" and "Full Trace".

## Output

One HTML file. Nothing else — no build step, no bundler, no package.json.

## Hard constraints (these are enforced; violating them means it won't publish or won't run)

- **Single file, ≤ 16 MiB.** All CSS and JS inline in the document.
- **Must contain a `<head>` element.**
- **No external resources at all.** No `<link>`, `<iframe>`, `<object>`, `<embed>`, `<base>` tags. No `@import`. No `src`/`action`/`formaction`/`poster` attribute pointing at `http:`, `https:` or `//`. That means no CDN scripts, no Google Fonts, no external stylesheets, no remote images.
- **No network access whatsoever.** CSP is `default-src 'none'; connect-src 'none'`, and the iframe is `sandbox="allow-scripts"` without `allow-same-origin`. `fetch`, `XMLHttpRequest`, WebSockets, and form submission all fail. You have no origin, no cookies, no `localStorage`, no access to the host page.
- **Allowed:** inline `<script>` and inline `<style>`, `data:` URIs for images and fonts, inline SVG, canvas, and everything else that runs purely in-page.
- **Vendoring a library is allowed** but it must be inlined in the file and it counts against the 16 MiB. Prefer writing it yourself — a big bundle costs parse/compile time on every open, and the iframe shares nothing with the host page.

## The SDK

The host injects a global before your code runs:

```js
window.traces.version    // "surface-sdk.v1"
await window.traces.getTrace()   // Promise<SurfaceTrace>
```

`getTrace()` is the entire data API. There is exactly one call, it resolves once with everything you get, and there is no way to ask for more data. It rejects if the host doesn't deliver within 10 seconds — handle that and render a readable message rather than a blank page.

Data is pushed to you over a private channel; you never request it.

### Classifying (beta)

Classification is a beta feature and needs a signed-in viewer. For anyone signed out it fails with `errorCode: "unauthenticated"`, so every surface that uses it needs a fallback.

`window.traces.classify` answers typed questions about any JSON state you build from the trace:

```js
const result = await window.traces.classify({
  state: { command: call.args.command },
  questions: {
    destructive: { type: "bool", instructions: "Does `command` delete data?",
                   criteria: { true: "Removes files or history", false: "Leaves data intact" } },
    area: { type: "choice", instructions: "What does `command` touch?",
            criteria: { git: "Version control", files: "The filesystem", network: "Remote hosts" } },
    risk: { type: "score", instructions: "How consequential is `command`?",
            criteria: ["Read-only", "Local change", "Leaves the machine"] }
  }
});
// result.answers.destructive.probability · .area.choice / .probabilities · .risk.score (0 = first level)
```

It never rejects. Check `result.stopReason === "stop"`; otherwise `result.errorMessage` says why, and `result.errorCode` says what to do:
- `"unauthenticated"`, `"forbidden"` or `"quota_exceeded"`: this viewer can't classify. Stop and say so.
- `"rate_limited"`: the viewer is calling too fast. Back off briefly and retry.
- no code: this call failed. Render the fallback; the platform has already retried what a retry can fix.

Getting good answers:
- **Decide what each answer is about**, such as a tool call, a user turn or the whole trace. The `state` is that item plus the context needed to judge it. Every question sees all of the state, so unrelated content can sway an answer, and missing context makes it guess.
- **Ask everything about an item in one call.** Questions are answered independently against the state, so they don't affect each other, and the state is paid for once. Use a separate call only when a question needs different evidence or depends on another answer.
- **Point questions at the state.** Name fields with backticked paths, such as `` `action.output` ``, rather than "this" or "the output".
- **Choose how many items share a call.** One item per call gives each answer the cleanest evidence but makes the most calls, which is slower on long traces and repeats a little overhead per item. Several items per call is faster and lighter on limits, but every question then sees the other items, and content unrelated to a question can cost accuracy. Group items when speed matters more than the last bit of accuracy, when items are short, or when the question compares them. Give single-item checks that must be right (did it fail, did it touch secrets) their own call. When you group, give each item its own field and name it in its questions.
- **Send calls in parallel.** You don't need your own retries or concurrency tuning; the platform retries the model's own failures. Handle `"rate_limited"`, and classify what's on screen first if the trace is long. Oversized requests are refused, so trim `state` to the relevant excerpt.

## The data contract

```ts
type SurfaceTrace = {
  trace: {
    id: string;
    title?: string;
    agentId: string;                       // which coding agent produced it
    namespace: { id: string; slug: string };
    author?: { id: string; displayName?: string; avatarUrl?: string };
    model?: string;
    createdAt: number;                     // epoch ms
    startedAt?: number;
    messageCount?: number;                 // total in the trace, may exceed messages.length
    tokenUsage?: TokenUsage;
    project?: { name?: string; path?: string; gitRemoteUrl?: string; gitBranch?: string };
    analysis?: TraceAiAnalysis;            // optional AI summary, may be absent
  };
  messages: SurfaceTraceMessage[];
  truncated: boolean;
};

type SurfaceTraceMessage = {
  id: string;
  role: string;                            // e.g. user / assistant / system — do not assume a closed set
  text?: string;                           // flattened text, convenient but lossy
  model?: string;
  order?: number;
  timestamp?: number;                      // epoch ms
  parts: SurfaceTracePart[];
  tokenUsage?: TokenUsage;
};

type SurfaceTracePart = {
  type: "text" | "thinking" | "tool_call" | "tool_result" | "error" | "system_event";
  order: number;
  content: Record<string, unknown>;        // shape depends on type
  tokenUsage?: TokenUsage;
};
```

Part content shapes:

- `text` / `thinking` → `{ text }`
- `tool_call` → `{ callId, toolName, args }`
- `tool_result` → `{ callId, toolName, output, status: "success" | "error" }`
- `error` → `{ message }`
- `system_event` → `{ subtype, ... }` (open-ended; extra fields ride along verbatim)

Any content may also carry `startedAt`, `completedAt`, `durationMs`. Do **not** sum `durationMs` across parts — some agents report per-part timing and others stamp a whole turn's wall-clock onto every part in it, so summing double-counts.

Pair `tool_call` with `tool_result` on `callId`, not on adjacency.

## Rules for the code you write

- **Every field except `trace.id`, `trace.agentId`, `trace.namespace`, `trace.createdAt`, `messages[].id`, `messages[].role`, `messages[].parts` and `parts[].type/order/content` is optional.** Real traces are missing things — no title, no model, no token usage, no author. Render a sensible placeholder instead of `undefined` or `NaN`, and never crash on a missing field.
- **Handle `truncated: true` honestly.** You may be given a prefix of a long trace. If `truncated` is true, or `messages.length < trace.messageCount`, say so in the UI — "over the first N of M messages" — rather than presenting a partial aggregate as if it covered the whole trace.
- **Tolerate unknown values.** New part types, new `system_event` subtypes and new roles will appear. Skip or generically render what you don't recognise; never throw.
- **Escape all trace content.** Message text and tool output are untrusted arbitrary strings, frequently containing HTML, code and markup. Use `textContent`, never `innerHTML`, for anything derived from the data.
- **Don't sort by array position.** Use `order` / `timestamp` when you need sequence, and cope with them being absent.
- **Assume it can be big.** Up to a few thousand messages. Don't render thousands of DOM nodes eagerly if you can summarise, virtualise or collapse.
- **Render something immediately**, then fill in when `getTrace()` resolves: a loading state, an error state for the rejection case, and an empty state for a trace with no messages.
- **Keep states mutually exclusive.** If you use the HTML `hidden` attribute, include `[hidden] { display: none !important; }`. Author CSS such as `.state { display: grid }` can otherwise override the browser default and expose loading, error, empty, and content states simultaneously.
- **Define the surface's semantic units from the current request.** Do not assume what a “turn,” step, token, duration, category, or aggregate means based on another surface. Label estimated or derived values as such.
- Look reasonable on both a narrow and a wide viewport; the iframe is the width of the trace page.
- Keep it self-explanatory. No settings, no persistence — there is nowhere to save anything.

## Aesthetics — match traces.com

The surface renders inside the Traces trace page, so it should look like it belongs there rather than like an embedded third-party widget. Traces is a restrained, near-monochrome, information-dense UI: flat surfaces, hairline borders, one accent colour used sparingly, no drop shadows, no gradients, no rounded-pill decoration.

**The surface is a section of the trace page, not a page of its own.** It renders in a tab directly under the trace's own `<h1>` (the trace title) and the tab bar, so it must not restate that chrome:

- **Never paint a background on the whole surface.** Leave `html` and `body` transparent — no `background: var(--background)` on the root. `--background` is still fine for a *recessed* fill inside a card, just never on the root.
- **Never open with a title, an `<h1>`, or eyebrow text, and don't add a heading just because it's a section.** It's part of a layout, so start with the content — the tab name already acts as the surface's header. Add an `<h2>` only where it earns its place: a later section the tab name doesn't cover, or a sibling that needs telling apart.
- **Never add outer padding or margin on `html`/`body`.** The frontend already applies the page's outer gutter around the iframe; a root padding doubles it up as dead space. Put spacing on your own inner elements instead — `body { margin: 0 }`, no `padding` on `html` or `body`.

```html
<!-- Don't: a page header the trace page already gives you -->
<body>
  <p class="eyebrow">SKILL ACTIVITY</p>
  <h1>Skills loaded in this run</h1>
  <p class="lede">Open a skill to read its instructions and see every recorded load.</p>
  <section class="stats">…</section>
</body>

<!-- Do: open on the content; section headings from h2 down -->
<body>
  <section class="stats">…</section>
  <h2>Loaded skills</h2>
  <ul class="skill-list">…</ul>
</body>
```

**Always style both light and dark — never ship just one.** The host gives you no theme signal. Read `@media (prefers-color-scheme: dark)` and define both, for every color you introduce, not just the tokens below. Declare `color-scheme: light dark` so form controls and scrollbars follow.

Copy these tokens verbatim — they're the real values from the app:

```css
:root {
  color-scheme: light dark;
  --background: #f5f5f5;
  --foreground: #141414;
  --card: #ffffff;
  --muted: #f2f2f2;             /* recessed fills */
  --muted-foreground: #6b6b6b;  /* labels, secondary text */
  --faint-foreground: #808080;  /* timestamps, least important text */
  --border: #e3e3e3;
  --border-strong: #dbdbdb;
  --hairline: rgb(0 0 0 / 0.07);
  --accent: rgb(0 0 0 / 0.03);  /* hover fill */
  --primary: oklch(0.50 0.15 278);   /* indigo-violet; links, selection, focus */
  --destructive: #cf4415;       /* errors, deletions */
  --success: #4b7c0b;           /* success, additions */
  --radius: 0.625rem;           /* 10px cards; 5px for small controls */
}

@media (prefers-color-scheme: dark) {
  :root {
    --background: #141414;
    --foreground: #e2e2e2;
    --card: #262626;
    --muted: #111111;
    --muted-foreground: #a0a0a0;
    --faint-foreground: #737373;
    --border: #202020;
    --border-strong: #242424;
    --hairline: rgb(226 226 226 / 0.06);
    --accent: rgb(255 255 255 / 0.03);
  }
}
```

**Type.** Body is Inter in the app; you can't load it, so use `font-family: Inter, ui-sans-serif, system-ui, sans-serif` and let it fall back. Code, tool names, file paths, hashes, IDs and numeric columns are monospace: `ui-monospace, "Berkeley Mono", SFMono-Regular, Menlo, monospace`, ~13px. Body text 13–14px, section labels ~12px in `--muted-foreground`, occasionally uppercase with slight letter-spacing. **Inter is the only typeface** — no display, serif, or decorative face for titles or emphasis — unless the user's request explicitly names another one. Monospace, as above, is the only other family, reserved for code/numeric content.

**Weight.** Only two font weights: `font-weight: 400` (normal) and `font-weight: 500` (medium). Never 600/700/bold/semibold. Titles — including `<h2>`, the largest text you have — are `font-weight: 400`; size and tight tracking (`letter-spacing: -0.01em` to `-0.02em`) carry the hierarchy, not weight. Reserve `500` for small text that needs to read as distinct from adjacent body text at that *same* size — a list-item's title next to its description, a label next to its value. Everything else is 400.

**Restraint.** Use as few distinct type styles and spacing values as possible (ex: one font size, one text colour, one weight, one spacing unit — vary just one property when a genuinely different element needs a distinction).

**Layout.** Use `--card` sparingly: `1px solid --border`, `border-radius: var(--radius)`, no shadow. Prefer a single hairline (`--hairline`) between list rows over boxing each one. Hover = `--accent` fill, not a colour change. Keep the palette monochrome; `--primary` for links/focus/one emphasis, `--destructive`/`--success` only for genuine error/success.

**Density.** Traces is comfortable with dense tables and tight rows — prefer a compact table or list over big cards with lots of padding. Right-align and monospace numbers so columns line up. Format timestamps as short absolute dates (or relative for recent), and format bytes/tokens with thousands separators rather than raw integers.

**Sizing.** The frame is the full width of the trace page and **grows to fit your content** — the SDK measures your layout and the host resizes the iframe, so the trace page does the scrolling. Lay out naturally (`body { margin: 0 }`, no fixed-height wrapper) and do not build your own inner scroll container: a `height: 100vh` root or an internally scrolling shell gets you a nested scrollbar, and the host deliberately refuses to grow past the viewport for surfaces whose height tracks the viewport. Minimum height is 32rem, maximum 1000000px. Don't add your own outer border or page background that would double up the frame's. **Don't cap the surface with a `max-width` and center it — use the full width you're given.** A centered column reads as a floating card rather than the tab's own content.

**Motion.** Almost none. If you animate, keep it under ~150ms and wrap it in `@media (prefers-reduced-motion: reduce)` so it can be turned off.

Do not publish, install, set a current version, commit, or change the API/frontend runtime unless the user explicitly asks.

## Trying it on a real trace

A surface can only be seen on a trace, so the loop is: upload a version, try it, then release it. Uploading does not change what anyone else sees — a version is invisible until it is made current.

1. Prepare an upload for the surface key and a new semver (`traces_surfaces_prepare_upload` via the Traces MCP, or `POST /v1/surfaces/:key/versions/uploads` with an API key). You get back a short-lived, single-use `upload` destination.
2. Send the raw HTML file to it directly, e.g. `curl -fsS -X POST -H 'Content-Type: text/html' --data-binary @surface.html <url>`. The response contains `artifactId`. Never paste the HTML into a tool call.
3. Complete the upload (`traces_surfaces_complete_upload`, or `POST …/uploads/complete`) with that `artifactId`. Leave `release` off.
4. Open the version on a trace: `https://traces.com/s/<trace-id>?surface=<key>&version=<version>`. The MCP tool returns this link for the namespace's latest trace as `previewUrl`; share it with the user. Only namespace members can open a version that is not current.
5. Fix, bump the version, repeat. When the user is happy, release it (`traces_surfaces_release_version`, or `PATCH /v1/surfaces/:key`) to make it current.

## Now build

Use the user's requested surface description to decide what the surface should show and how. If that description is missing or leaves a consequential semantic choice unresolved, ask a focused question before building.
