RFC-003 — Recording skill orchestration (v0.1)
- Status: Draft
- Authors: @yiidtw
- Created: 2026-05-09
- Related: RFC-001 (function-based v0.1), RFC-002 (Clipper skills catalog), SPEC.md § “amem is its own best demo,” archived guide/self-recording.md (superseded by this RFC)
TL;DR
Replace the existing chrome.tabCapture self-recording skeleton with a
scripted, librarian-driven recording pipeline. A YAML script declares a
sequence of Chrome operations; the librarian drives Chrome through the
bridge (which executes them in amem Clipper) while simultaneously running
macOS window-level screencapture -v -l<windowID> against the Chrome
window. Output: an mp4 at ~/.amem/recordings/<uuid>.mp4 plus an
amem://recording/<uuid> URI. Optional handoff to video-use for
post-processing (transcribe, cut filler, captions).
The whole thing surfaces as one MCP tool — amem_record_demo — and one
skill card (cws-demo, RFC-002 §3c). v0.1 ships one built-in script
(the CWS demo). Two more (feature-update demo, tutorial template) follow
when the third script is requested.
Why librarian-side and not extension-side: chrome.tabCapture cannot
record the sidepanel UI itself (sidepanels are out-of-tab surfaces),
which is exactly the surface our CWS demo needs to show. Window-level
macOS capture fixes that and only that.
Motivation
amem inherits crossmem’s principle — “a product that is its own best
demo” (SPEC.md § README, also docs/src/guide/self-recording.md). We
record our own marketing video by driving our own extension. This is real
load-bearing infrastructure, not a marketing gimmick: every CWS
re-submission, every feature announcement, every tutorial benefits.
The Day 1 skeleton (docs/src/guide/self-recording.md) used
chrome.tabCapture from an offscreen document. Three problems made us
abandon it for v0.1:
chrome.tabCapturecannot capture sidepanels. It captures the tab’s rendered area; the sidepanel is outside that surface. The single most important thing our demo must show — the skills catalog sidepanel — is invisible to tabCapture. We could render the catalog inside a tab as a workaround, but then we are demoing a fake.- No driving model. The skeleton assumed “an orchestrator” sends
start_recordingand “drives the extension UI.” There is no orchestrator design — just a hand-wave. v0.1 must ship a real one. - MV3 service worker lifecycle is hostile to long recordings. Service workers can be evicted under memory pressure; offscreen documents help but add coordination complexity. Doing this in the librarian (a long-running Rust process) is simpler and more reliable.
Moving the recorder to the librarian means it can use macOS-native
screencapture (window-level, captures everything in the Chrome window
including all of Chrome’s chrome) and orchestrate the demo via the bridge.
Same component split as RFC-001: librarian = brain, clipper = sensor.
Proposal
1. YAML script schema
Scripts live in two places:
crates/amem-librarian/builtin_scripts/*.yaml— built-in templates (CWS demo, etc.), compiled into the binary~/.amem/recordings/scripts/*.yaml— user scripts (v0.1 supports these but does not have a UI for managing them; CLI only)
Schema:
# ~/.amem/recordings/scripts/cws-demo.yaml
name: "CWS demo (v0.1)"
duration_target: 45s # advisory; total run time including waits
window:
app: "Google Chrome"
match: "active" # or { title_regex: "..." } for multi-window
record_cursor: true # passed to screencapture
output:
format: mp4
resolution: source # or 1080p, 720p (downscale via ffmpeg)
filename: "cws-demo-{ts}.mp4"
steps:
- id: open-arxiv
type: navigate
url: "https://arxiv.org/abs/1706.03762"
wait_for: "h1.title"
- id: capture-arxiv
type: caption
text: "amem captures any arxiv paper you open"
duration: 3s
- id: open-skills
type: click_selector
selector: "[data-amem-tab='skills']"
wait_after: 500ms
- id: show-skills
type: caption
text: "Three skills shipped — and your agent can call any of them"
duration: 4s
- id: click-linkedin
type: click_selector
selector: "[data-amem-skill='linkedin-inbox'] button.run"
wait_for: "[data-amem-skill='linkedin-inbox'] .digest"
- id: settle
type: wait
duration: 2s
- id: terminal-finish
type: terminal
cmd: "amem recall --json 'attention is all you need'"
show: "stdout"
duration: 4s
2. Step types
Each step is one variant; steps execute sequentially (no v0.1 parallelism).
| Type | Purpose | Driver |
|---|---|---|
navigate | Load URL in active tab | chrome_navigate over bridge |
click_selector | Click an element | chrome_click |
wait | Pause for N ms / N s | librarian sleep |
wait_for (also a field on other steps) | Block until selector / URL appears | chrome_wait |
extract_dom | Read DOM, save to script vars (for later assertions/captions) | chrome_extract |
caption | Render an overlay caption for N seconds | librarian draws into a borderless overlay window |
terminal | Run a binary command, optionally render its stdout in an overlay | librarian process + overlay |
The terminal step is interesting: it lets a recording show CLI usage
alongside the browser. The librarian opens a small floating window (via
its own UI process, not the Chrome window) that renders the command and
its stdout in monospace; screencapture picks up the overlay because we
target the Chrome window but composite the overlay on top of it before
each frame. Implementation detail: macOS lets us position a borderless
NSWindow above the target Chrome window and screencapture -l<chromeWinId>
will include it (Quartz compositing, same as visible-tab capture).
If the overlay approach proves unreliable across macOS versions, fallback
v0.1: split the recording — capture Chrome window for browser steps,
capture full screen for terminal steps, stitch in post via video-use.
Decision deferred until we hit a Sonoma/Sequoia regression in testing.
caption steps render a similar borderless overlay at the bottom-center
of the Chrome window with a translucent black bar.
3. macOS implementation — window-level screencapture
Core command:
screencapture -v -l <chromeWindowId> -V <duration_seconds> output.mov
-v— start recording immediately (no UI)-l <windowId>— target a specific window. We get this fromCGWindowListCopyWindowInfo(.optionOnScreenOnly)filtered to bundle identifiercom.google.Chrome(or Brave / Arc / Edge variants — we hardcode the major Chromium IDs).-V <seconds>— fixed duration. v0.1 sets this toduration_target + 10s buffer; we kill the process early if all steps finish before the timer.- Output
.mov; converted to.mp4viaffmpeg -i in.mov -c:v libx264 -crf 20 out.mp4post-recording.
Why window-level not full-screen:
- Privacy: full-screen captures the user’s desktop background, dock, notifications, other apps. Window-level captures only the Chrome window’s pixel rect.
- Aesthetics: window-level avoids us cropping the recording in post.
- Privacy posture matches the rest of amem (“data stays local; we don’t capture what we don’t need”).
Permission flow: macOS requires Screen Recording permission. On
first invoke, the librarian prompts the user via TCC. If denied, we
return { code: "SCREEN_RECORDING_DENIED", how_to_fix } so the agent
can surface the system-settings link.
Linux / Windows: out of scope for v0.1. The recording skill is
gated to macOS in v0.1 (#[cfg(target_os = "macos")]); on other OSes
the skill renders as kind:"error" with “Recording requires macOS in
v0.1.” Linux pipeline (probably wf-recorder for Wayland +
scrot/ffmpeg-x11grab for X11) is a follow-up RFC.
4. video-use integration (optional post-processor)
The raw mp4 from §3 is usable as-is — but for marketing-grade output we
want transcription, filler-word cutting, captions, and consistent
encoding. video-use is the operator’s existing tool for that; integrating
is a one-liner:
#![allow(unused)]
fn main() {
// crates/amem-librarian/src/skills/record_demo.rs
async fn record_demo(script: &Script) -> Result<RecordingOutput> {
let raw_mov = run_screencapture(&script.window, script.duration_target).await?;
let raw_mp4 = ffmpeg_convert(&raw_mov).await?;
let final_mp4 = if script.post_process.unwrap_or(false) {
video_use::process(&raw_mp4, &script.post_options).await?
} else {
raw_mp4
};
Ok(RecordingOutput {
amem_uri: format!("amem://recording/{}", uuid),
mp4_path: final_mp4,
})
}
}
post_process: false is the v0.1 default (raw recording). Setting
post_process: true in the YAML opts in. video-use is treated as an
optional dependency: if it isn’t installed, the field is ignored with a
warning.
The script’s post_options mirror video-use’s CLI flags one-to-one (so
the integration stays a thin shim, not a redesign).
5. amem_record_demo MCP tool
Exposed indirectly through amem_invoke_skill("cws-demo") per RFC-001
§3d, but also as a direct tool for ad-hoc agent use:
amem_record_demo(
script_path?: string, // path to YAML
script_inline?: string, // YAML literal (mutually exclusive with script_path)
output_dir?: string, // defaults to ~/.amem/recordings/
)
-> {
amem_uri: "amem://recording/<uuid>",
mp4_path: "/Users/.../<uuid>.mp4",
duration_s: number,
steps_run: number,
}
Errors:
| Code | Meaning |
|---|---|
SCREEN_RECORDING_DENIED | macOS TCC denied screen recording |
BRIDGE_UNAVAILABLE | Cannot reach amem Clipper to drive Chrome |
WINDOW_NOT_FOUND | No Chrome window matched script.window |
STEP_FAILED | Some step failed; partial recording saved with partial: true |
RECORDING_IN_PROGRESS | Another recording is active; refuse to start a second |
UNSUPPORTED_OS | Not macOS in v0.1 |
The tool is blocking: it returns when the recording finishes (typically 30–90s for v0.1 scripts). MCP clients already render “long-running tool call” affordances.
6. Use cases
The same pipeline serves three concrete needs, each justifying ship in v0.1:
6a. CWS promo video (the SPEC’s “own best demo” principle)
The CWS listing needs a 30–60 second demo. We script it (§1), the librarian records it, video-use polishes it, we upload. When v0.1.x patches change the UI, we re-run the same script. The demo never goes stale.
6b. Feature update demos (post-v0.1)
When v0.2 ships custom-skill installation, we want a demo of that flow. Same pipeline, new YAML script. v0.2 itself is a script.
6c. Tutorials
docs.amem.sh user guide pages can embed inline mp4s recorded from
canonical YAML scripts in crates/amem-librarian/builtin_scripts/. When
the UI changes, regenerate.
7. Why the librarian records, not the clipper
The cleanest framing of this RFC’s central decision:
| Option | Records sidepanel UI? | Privacy? | Lifecycle? | Cross-OS path? |
|---|---|---|---|---|
A. chrome.tabCapture (extension) | ❌ No (sidepanel is out-of-tab) | ✅ tab-only | ⚠️ MV3 service worker eviction risk | ✅ identical everywhere |
B. getDisplayMedia (extension) | ✅ User picks the window | ✅ user opt-in per recording | ⚠️ same MV3 risks | ✅ uniform |
| C. macOS window screencapture (librarian) | ✅ entire Chrome window | ✅ window-only, no desktop leak | ✅ Rust process is long-lived | ❌ macOS-only v0.1 |
Option B is the second-best choice; we considered it seriously. We rejected it for v0.1 because:
getDisplayMediarequires a user picker dialog every recording — incompatible with a “click Run, get a recording” agent-driven flow.- The MV3 service worker would still need to coordinate with the librarian for the script driver, doubling the moving parts.
- We want the librarian to own the timeline anyway — it already owns the script, the bridge, the storage, the post-process step. Adding the recording itself keeps responsibility in one place.
Option C costs us OS portability in v0.1 — accepted tradeoff. Linux / Windows recording lands when there is a non-macOS user.
Privacy
- Window-level screencapture: only the Chrome window’s pixels are captured. Desktop, dock, menu bar, other apps, notifications — all outside the frame.
- The recorded mp4 lives in
~/.amem/recordings/<uuid>.mp4. Never uploaded by the librarian. The user canamem upload(future) or manually drag-drop to a CWS listing. - macOS Screen Recording permission is requested on first invoke; the user can revoke at any time in System Settings.
- Caption / terminal overlays render the script’s literal text; no enrichment, no LLM rewriting (v0.1).
video-use(when used) is a local binary; transcription runs locally via whisper. No cloud calls in the default config.- The script itself is recorded into the mp4 metadata’s
commentfield (so anyone with the mp4 can see what was scripted). The user can disable withmetadata: falsein the YAML.
Failure modes
| Mode | Cause | Mitigation |
|---|---|---|
| Screen Recording permission denied | First-run user hasn’t granted TCC | Tool returns SCREEN_RECORDING_DENIED with open System Settings → Privacy → Screen Recording instruction; sidepanel card shows the same |
| Chrome window not found | User has no Chrome window open, or app is Brave/Arc | YAML window.app allow-list extended to common Chromium variants; tool errors with WINDOW_NOT_FOUND listing detected windows |
| Step times out | Selector never appears | Step times out at wait_for budget (default 30s); recording stops; partial: true flag in result so user knows |
| Bridge disconnects mid-recording | Librarian-clipper WS drops | Recording continues (it’s screencapture, not bridge-driven pixels), but subsequent steps can’t drive Chrome; we abort gracefully and save what we have |
screencapture produces 0-byte file | Known macOS bug in some Sonoma builds when target window is fully occluded | Pre-check: bring Chrome to front + verify visibility before starting; document the Apple bug in docs/troubleshooting.md |
| Multiple recordings requested concurrently | Two agents both call amem_record_demo | Librarian-wide recording lock; second call gets RECORDING_IN_PROGRESS |
| Output mp4 huge | High-resolution display, long demo | Default resolution: source for v0.1; users can opt to 1080p / 720p in YAML; ffmpeg downscale handled in §4 |
| Terminal overlay flickers / drops frames | macOS compositor under load | Documented; user can switch to “split capture + post-stitch” mode |
| Wrong Chrome window picked (multi-window) | User has 3 Chrome windows; we pick wrong | YAML window.match accepts title_regex; default is “active window” which uses the focused one at recording start |
Concrete work
In rough order:
- (
amem-librarian) YAML script parser + schema validator — ~0.5d - (
amem-librarian) macOS window-id resolver via QuartzCGWindowListCopyWindowInfo(FFI throughcore-graphicscrate) — ~0.5d - (
amem-librarian)screencapturedriver: spawn, monitor, kill-early on completion — ~0.5d - (
amem-librarian) Step executor: bridge-drivennavigate/click_selector/wait_for/extract_dom(mostly reuses RFC-001 bridge wrappers) — ~1d - (
amem-librarian) Caption / terminal overlay window (NSWindow with borderless Cocoa view) — ~1.5d (riskiest step; if compositing misbehaves, fall back to post-stitch mode and trim to ~0.5d) - (
amem-librarian)ffmpegmp4 conversion + optionalvideo-usehandoff — ~0.5d - (
amem-librarian)amem_record_demoMCP tool surface + structured errors — ~0.5d - (
amem-librarian) Built-incws-demo.yamlscript + smoke test on the actual amem Clipper UI — ~1d (this is the dogfood — script will reveal UI bugs) - (
docs.amem.sh) Replaceguide/self-recording.mdcontent with pointers to this RFC + how-to for users — ~0.25d
Total: ~6.25d (5d if overlay step falls back to post-stitch).
Rejected alternatives
- Keep
chrome.tabCapturefrom offscreen documents. Cannot capture sidepanel; demo would have to fake the catalog UI. Disqualifying. - Use
getDisplayMediafrom the extension. User picker dialog every time; service worker lifecycle headaches; doubles the moving parts. See §7 Option B. - Run a generic OBS / ffmpeg pipe and let the user start/stop manually. Loses the “agent-driven scripted demo” capability that makes self- recording load-bearing. We’d be back to manual demo production.
- Embed a video-use clone in the librarian. video-use is its own product with its own scope; reimplementing transcription / cut-filler in amem is out of scope. Optional handoff is the right boundary.
- Ship Linux / Windows recording in v0.1. Tripled scope for an audience we don’t currently have. macOS-only is acceptable v0.1 posture.
- Record from a JS-only stack (puppeteer + ffmpeg-screen). Loses the user’s Chrome profile and login state (per RFC-001 §3a / browser automation rules). Non-starter.
- Skip captions / terminal overlays. Dropping captions makes the resulting mp4 unsuitable for CWS listings without manual editing, which defeats the whole “own best demo” pipeline. Worth the implementation cost.
Open questions
- Overlay window rendering technology. Native NSWindow + Cocoa view
(proposed) or a small SwiftUI helper app the librarian shells out to?
Soft preference: Cocoa from Rust via the
cocoa/objc2crates; SwiftUI helper if FFI gets miserable. - Caption style. Plain translucent black bar with white text v0.1; do we offer themes / fonts? Soft preference: defer to v0.2; one hardcoded style in v0.1.
- Whether to embed the YAML script in the mp4 metadata by default.
Useful for reproducibility, mildly leaky for privacy if the user
shares the mp4. Soft preference: opt-in (
metadata: truein YAML). - video-use handoff API stability. v0.1 calls it as a binary subprocess; if video-use grows a Rust crate API later, switch over. Not blocking.
- Recording lock granularity. Per-librarian (current proposal) or per-script-id? Soft preference: per-librarian (simpler; scripted recording is sequential by nature).
- Should the sidepanel show a live preview of the recording? Tempting but adds complexity (need to read the in-progress mp4 or echo screencapture frames). Defer to v0.2.
Roll-out
- Day 4 of v0.1 build: YAML parser + window resolver + screencapture driver land. Smoke test: record a 5-second blank capture.
- Day 5: step executor + bridge wiring. Smoke test: 10-second scripted recording opens arxiv, clicks something, finishes.
- Day 6: caption / terminal overlay. Risk day; fallback path ready.
- Day 7 (CWS submission day): full
cws-demo.yamlruns, video-use polishes, we have the listing video. Submit. - Post-v0.1: track real script usage. When the third user-script exists (per RFC-001 §6 rule-of-three trigger), re-evaluate whether YAML schema needs versioning, whether scripts need a UI, whether Linux support belongs in v0.2.
- Future: cloud-render mode (a remote machine runs the script and returns the mp4) for users without macOS. Out of scope here; tracked in a future RFC.