Std.Ui.Terminal — DOM byte stream vs. server screen + canvas

What moving the terminal emulation to the server, drawing on a canvas and sending screen-diff frames bought, against the first cut of the module, and why binary frames were not built.

Conditions

HostApple M1, 8 cores, 16 GB, macOS 27.0, arm64; other agents' work on the machine at the same time
Go / Node / Playwright1.26.1 / 26.3.1 / 1.59.1, headless Chromium
Apprust/crates/sky/tests/fixtures/ui-terminal (Sky.Live, sh on a PTY), built by each tree's compiler

Workloads (workloads.mjs)

Deterministic byte streams, the bytes a PTY hands the server:

namebyteswhat
yes4,194,303y\r\n to 4 MiB
seq1,488,895seq 1 200000
redraw686,100300 x (clear, a 38-line coloured ls -la)
colour5,816,27360 full 120 x 40 screens, every cell its own 256-colour fg and bg

Method

  1. Widget alone (bench-widget.mjs, median of 5). Headless Chromium, the island sized to exactly 120 x 40 cells, fed the commands one per task (MessageChannel, the way SSE messages arrive) as fast as the page takes them. Before: the stream cut into 64 KiB "output" commands. After: the frames the Go screen makes from the same 64 KiB cuts, one per cut, no pacing (TERM_BENCH_DIR=<dir> go test ./rt -run TestTerminalWireCost writes them) — the same number of messages as before. Measured: first command to last paint, time inside the widget's command handler and paint, widget work per animation frame, missed animation frames (a gap of k x 16.7 ms counts k-1), and Chromium's Performance.getMetrics TaskDuration / ScriptDuration / Layout + RecalcStyle deltas.
  2. Wire cost (TestTerminalWireCost in runtime-go/rt/term_bench_test.go, output in wire-cost.txt). For each workload in 64 KiB reads: the base64 "output" messages, the JSON frames, and the same frames in a compact binary encoding (varints, a tag byte per op, length-prefixed UTF-8) base64-encoded for SSE. Each as SSE island event text, raw and through one gzip stream flushed after every message (what a compressing proxy such as Caddy encode gzip does to an SSE stream).
  3. End to end (bench-app.mjs, median of 3). The fixture app; the test types cat <workload>.bin; echo; echo DONE$((40+2)) and waits (60 s cap) for the line DONE42. Measured: Enter to DONE42, SSE island bytes and messages (CDP Network.eventSourceMessageReceived), Chromium TaskDuration, missed animation frames, and whether the last line before DONE42 is the workload's last line. The fixture's terminal is 300 px tall (about 15 rows): every workload scrolls.
  4. Slow client. bench-app.mjs --throttle-kbps 2000 (CDP network emulation) and the Go test TestTerminalScreenIsExactWhenTheRingOverflows (a 4 KiB ring, 3,000 lines, the reader behind by far more than the ring).
  5. Server cost. go test ./rt -bench BenchmarkVTScreenFeed (the screen consuming built-in versions of yes, seq, colour at 120 x 40).

Results

1. Widget alone, 120 x 40 (before-widget.json, after-widget.json)

workloadtotal mscommand mspaint mswork / frame p95 msmissed framesChromium task mslayout + style ms
yesbefore2,208.62,151.926.035.364 of 662,252.57.1
after33.56.48.36.00 of 5157.70.7
seqbefore388.8338.724.425.00 of 25400.18.4
after26.15.65.35.60 of 4105.00.9
redrawbefore50.128.84.08.80 of 648.42.6
after16.61.65.05.00 of 213.80.4
colourbefore2,935.4265.9552.111.883 of 913,018.2679.8
after1,335.711.1376.26.06 of 921,996.933.4

The before widget's time goes to parsing (the JS VT: yes is 1.4 M scrolls) and, on colour, to style and layout of up to 4,800 styled spans per paint. After, parsing is on the server and a paint is a few fillRect / fillText calls per changed row.

2. Wire cost, 64 KiB reads (wire-cost.txt)

workloadbase64 bytes raw / gzipJSON frames raw / gzipbinary frames raw / gzip
yes5,599,958 / 14,006562,984 / 4,131544,768 / 3,948
seq1,987,882 / 519,643311,465 / 53,163341,248 / 64,179
redraw916,075 / 66,16517,624 / 2,56918,248 / 3,737
colour7,765,554 / 1,558,5675,450,137 / 1,821,0053,546,908 / 1,848,616

Frames against bytes, raw: 10x (yes), 6.4x (seq), 52x (redraw), 1.4x (colour) fewer. Binary against JSON frames: raw -3% / +10% / +4% / -35%; through gzip -4% / +21% / +45% / +2%.

3. End to end (before-app.json, after-app.json)

workloadEnter to DONE42SSE bytesSSE messagesChromium task msmissed framesright last line
yesbeforenever (60 s)3,914,3841361,613590 of 3
after766 ms325,678414203 of 3
seqbefore428 ms2,255,14630367193 of 3
after205 ms104,607101903 of 3
redrawbefore117 ms931,387148443 of 3
after106 ms195,04951703 of 3
colourbefore425 ms7,193,06598391162 of 3
after416 ms4,872,4891915103 of 3

The before yes never shows DONE42: the output outruns the page, the session's SSE buffer (16 frames) fills, and the island commands that do not fit are dropped. The widget only noticed a gap when a LATER command arrived, so the dropped commands at the end of the output left the terminal stopped short of the last line until the next output (confirmed by reading the widget's offset: it stopped at 3,033,573 of 4,194,303 bytes, with no gap flagged). One colour run lost its end the same way. After, frames are paced (16 ms), one per paint at most, so the buffer does not fill; if a frame is lost anyway, the next frame or the check frame (0.75 s after a burst) shows the gap and the widget asks for one repaint.

colour stays large after (4.9 MB): on a 15-row terminal every screen of the stress scrolls 40 rows of per-cell colours into the scrollback, and the scrollback lines are sent with their colours.

4. Slow client

--throttle-kbps 2000 did not bind the SSE stream in headless Chromium: the before seq run moved 2.26 MB in 1.9 s (1.2 MB/s, over the 250 KB/s cap). The figures are in before-slow-*.json / after-slow-*.json for the record (before yes: 1 of 2 runs right, 1 repaint; after: 2 of 2 right for every workload) but the slow-client argument rests on:

5. Server cost (BenchmarkVTScreenFeed, 120 x 40)

workloadthroughput
yes (3-byte lines: one scroll each)8.4-8.7 MB/s (about 2.9 M lines/s)
seq15.5-16.0 MB/s
colour66-67 MB/s

The screen runs on the process's output pump, so a process that writes faster than this is slowed to it (the flow control every terminal has); a slow page is not.

Decisions

  1. Canvas renderer: built. The DOM span renderer: removed. The canvas was faster on every workload (table 1), and most on the ones the span renderer was worst at. The text layer over the canvas (transparent, one row per screen row) keeps what the spans gave: mouse selection, copy, a screen reader. Without a 2D canvas it is shown as a monochrome fallback.
  2. Server-side VT with a screen-diff op stream: built. Scroll ops, row spans with styles, cursor, title, modes, bell. A remount is one repaint frame from the screen plus the last 1000 scrollback lines.
  3. Raw-byte replay: removed. Nothing needs the bytes once the screen is the source of truth, and the replay was what could not survive a ring overflow. Terminal.encodeOutput / encodeExit (unreleased) went with it.
  4. Binary frames: not built. Through gzip they save 4% on yes and cost more on the other three (table 2), under the 10% bar; uncompressed they save 35% only on the per-cell colour stress. A binary path would need a second transport (a WebSocket for terminal islands) proven against the strict CSP, proxies, reconnect and the header-session transport, for no measured saving. Frames stay JSON on the island SSE channel.