What a headless Sky.Http.Server + Std.Db read costs

Every capacity figure this project has ever measured is a Sky.Live SSE workload on examples/19-skyforum (see skylive-interaction-cost.md). The v1 audit flagged that the other shape a real trial will run — a stateless JSON API that reads Std.Db and returns JSON, no SSE, no sessions, no CSRF — had zero measured capacity. This document records the first measurement of that shape, the app that produced it, the harness, and — as with every perf doc here — what the measurement does not cover.

Rule for quoting anything here (inherited from the SSE doc): no number may be repeated without the conditions attached to it. The raw run and its env.txt are archived at runs/http-metadata-local-20260906/.

The app under test

examples/65-metadata-service — the shape of the internal "core-metadata" trial workload:

The example now SHIPS on SQLite; the PostgreSQL numbers below stand. examples/65-metadata-service/sky.toml ships [database] driver = "sqlite" (a single in-process file, no cluster/bundle/DSN) so the example runs anywhere — including CI's build-run gate, which starts the bare ./app binary with no sky run, no cluster, and no injected DSN (an embedded = true app exits on start there, because it cannot reach a PostgreSQL cluster). Because Std.Db is dialect-safe, this is a one-line config swap back to embedded PostgreSQL ([database] embedded = true) with no code change — that PostgreSQL configuration is the production target, and it is what every number in the "The measurement" section below was measured on. The PostgreSQL figures are real and were not re-run; they are not restated for SQLite.

A local SQLite data point for the shipped default (same host as the run below — Apple M1 Mac mini, Sky 48a6a4be, closed-loop load/loadgen.go, 5 s per level after a 1 s warm-up):

Endpointconcreq/sp50 msp99 mserr %
GET /metadata/:key (indexed single-row)85,695.71.382.310.00
GET /metadata/:key (indexed single-row)645,686.28.2945.720.00
GET /healthz (server ceiling)6438,613.11.236.750.00

On this host the SQLite single-row read path lands in the same ~5.7k req/s band as embedded PostgreSQL — the DB read, not HTTP/JSON, is the bound in both (the /healthz framework ceiling is ~7× higher). Treat this as an order-of-magnitude sanity point, not a substitute for the PostgreSQL sweep: SQLite is single-writer and single-file, so it does not carry the production tier's concurrency or multi-instance story.

EndpointReadVerified
GET /healthznone (server ceiling)200 {"status":"ok"}
GET /metadata/:keyone indexed row by PK (Db.findOneByField)200 JSON object / 404 on miss
GET /metadata?limit=Nfirst N rows ordered by key (Db.query)200 {"count":…,"items":[…]}

sky check is clean; all four responses (incl. the 404) were curl-verified before the load run.

How to reproduce

# 1. run the app (embedded PostgreSQL, binds :8137)
cd examples/65-metadata-service && sky run src/Main.sky

# 2. drive load from another shell (closed-loop Go harness, stdlib only)
./load/run-load.sh                          # default sweep, localhost:8137
LEVELS="512,1024,2048,4096" DUR=8s ./load/run-load.sh

The harness (load/loadgen.go) is closed-loop / constant-concurrency: at each level, N goroutines each loop "send → measure → repeat" for the level's duration — the model wrk/hey use. /metadata/:key requests a random svc-0001…svc-0500 key each time. No oha/hey/wrk/bombardier was present on the host, so a Go harness (the most accurate option available) was written; the client uses a 2048-conn pool and a 10 s per-request timeout.

The measurement

Host: Apple M1 Mac mini, 8 cores, 16 GiB, macOS 26.5.2, Go 1.26.1, Sky b83e9493. 8 s per level after a 1 s warm-up. The app, PostgreSQL, the load generator and go all shared the same 8 cores, and the host's 5-minute load average was elevated (~8–10) from concurrent builds. This is a co-located single-laptop baseline, not an isolated benchmark.

GET /metadata/:key — indexed single-row read (the hot path)

concreq/sp50 msp90 msp99 msmax mserr %
14,054.80.240.270.301.130.00
85,603.11.401.872.334.710.00
165,765.32.584.326.1511.060.00
325,671.55.049.8315.2229.690.00
645,664.510.2419.0429.1647.220.00
1285,645.119.4740.6369.80167.580.00
2565,677.434.8891.11169.95355.430.00
5125,639.466.60196.17379.03788.850.00
10245,658.5130.32401.79784.642,154.710.00
20485,657.9262.38795.741,552.624,545.860.00
40965,533.3531.521,614.973,087.376,288.090.00

GET /metadata?limit=50 — range read, 50 rows

concreq/sp50 msp90 msp99 msmax mserr %
12,139.90.460.490.582.050.00
85,097.31.522.062.7910.470.00
325,133.35.4910.9916.9730.560.00
1285,022.121.6346.0180.12217.090.00
2565,009.139.31103.57193.52442.380.00
5125,059.273.78217.03427.551,000.380.00

GET /healthz — no DB touch (the framework ceiling)

concreq/sp50 msp99 mserr %
114,566.00.070.100.00
1637,817.70.371.220.00
6439,054.31.226.610.00
25637,860.86.4915.330.00
51237,542.613.4824.220.00

Full sweeps: runs/http-metadata-local-20260906/sweep.txt and knee-search.txt.

What the numbers say

What this does NOT establish — and what remains

This is a local-laptop baseline. It is not the trial's capacity. It proves three things and only three: the shape works end-to-end (Sky.Http.Server + Std.Db + embedded PostgreSQL, real 200s with JSON bodies), it has a per-request cost floor (~0.24 ms p50 single-row at concurrency 1), and its local ceiling on an 8-core M1 sharing its cores with the DB and the load generator is ~5.6k req/s with zero errors to 4096 concurrent.

It does not establish the number a capacity claim needs, because:

  1. Wrong machine. The trial runs on a GCE instance (a specific e2/n2/c3 family + size), not an M1. Prior work in this repo has been wrong by 2.5–5× carrying a laptop/container number to real GCE hardware, and by roughly an order of magnitude sizing on the wrong resource — see the CPU-binds-before-memory and count-physical-cores-not-vCPUs sections of skylive-interaction-cost.md. No extrapolation to a GCE number is made here, deliberately.
  2. Co-located and contended. The app shared 8 cores with PostgreSQL, the load generator, and go, under an elevated load average. A real deployment separates the load driver (and often the database) onto other hosts; the ceiling would move.
  3. No burst/soak behaviour. The SSE doc found a rested burstable e2 overstates sustained capacity by ~2.7× on its first run; that decay is unmeasured for this shape.

What remains (blocked on instance access): the at-scale run of this exact app on the target GCE instance family — sky db provision --embed (or a Cloud SQL DSN) on the instance, the load driven from a separate bench host, sweeping concurrency until the real knee (SLA p99 breach or genuine errors) appears, with MemAvailable and CPU sampled alongside. That run — not this one — produces the capacity number the trial can be sized on. The app and harness in examples/65-metadata-service/ are written to run unchanged against a remote URL=http://<instance-ip>:8137.