Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

RFC-003 — Recording skill orchestration (v0.1)

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-05-09
  • Related: RFC-001 (function-based v0.1), RFC-002 (Clipper skills catalog), SPEC.md § “amem is its own best demo,” archived guide/self-recording.md (superseded by this RFC)

TL;DR

Replace the existing chrome.tabCapture self-recording skeleton with a scripted, librarian-driven recording pipeline. A YAML script declares a sequence of Chrome operations; the librarian drives Chrome through the bridge (which executes them in amem Clipper) while simultaneously running macOS window-level screencapture -v -l<windowID> against the Chrome window. Output: an mp4 at ~/.amem/recordings/<uuid>.mp4 plus an amem://recording/<uuid> URI. Optional handoff to video-use for post-processing (transcribe, cut filler, captions).

The whole thing surfaces as one MCP tool — amem_record_demo — and one skill card (cws-demo, RFC-002 §3c). v0.1 ships one built-in script (the CWS demo). Two more (feature-update demo, tutorial template) follow when the third script is requested.

Why librarian-side and not extension-side: chrome.tabCapture cannot record the sidepanel UI itself (sidepanels are out-of-tab surfaces), which is exactly the surface our CWS demo needs to show. Window-level macOS capture fixes that and only that.

Motivation

amem inherits crossmem’s principle — “a product that is its own best demo” (SPEC.md § README, also docs/src/guide/self-recording.md). We record our own marketing video by driving our own extension. This is real load-bearing infrastructure, not a marketing gimmick: every CWS re-submission, every feature announcement, every tutorial benefits.

The Day 1 skeleton (docs/src/guide/self-recording.md) used chrome.tabCapture from an offscreen document. Three problems made us abandon it for v0.1:

  1. chrome.tabCapture cannot capture sidepanels. It captures the tab’s rendered area; the sidepanel is outside that surface. The single most important thing our demo must show — the skills catalog sidepanel — is invisible to tabCapture. We could render the catalog inside a tab as a workaround, but then we are demoing a fake.
  2. No driving model. The skeleton assumed “an orchestrator” sends start_recording and “drives the extension UI.” There is no orchestrator design — just a hand-wave. v0.1 must ship a real one.
  3. MV3 service worker lifecycle is hostile to long recordings. Service workers can be evicted under memory pressure; offscreen documents help but add coordination complexity. Doing this in the librarian (a long-running Rust process) is simpler and more reliable.

Moving the recorder to the librarian means it can use macOS-native screencapture (window-level, captures everything in the Chrome window including all of Chrome’s chrome) and orchestrate the demo via the bridge. Same component split as RFC-001: librarian = brain, clipper = sensor.

Proposal

1. YAML script schema

Scripts live in two places:

  • crates/amem-librarian/builtin_scripts/*.yaml — built-in templates (CWS demo, etc.), compiled into the binary
  • ~/.amem/recordings/scripts/*.yaml — user scripts (v0.1 supports these but does not have a UI for managing them; CLI only)

Schema:

# ~/.amem/recordings/scripts/cws-demo.yaml
name: "CWS demo (v0.1)"
duration_target: 45s        # advisory; total run time including waits
window:
  app: "Google Chrome"
  match: "active"           # or { title_regex: "..." } for multi-window
record_cursor: true         # passed to screencapture
output:
  format: mp4
  resolution: source        # or 1080p, 720p (downscale via ffmpeg)
  filename: "cws-demo-{ts}.mp4"

steps:
  - id: open-arxiv
    type: navigate
    url: "https://arxiv.org/abs/1706.03762"
    wait_for: "h1.title"

  - id: capture-arxiv
    type: caption
    text: "amem captures any arxiv paper you open"
    duration: 3s

  - id: open-skills
    type: click_selector
    selector: "[data-amem-tab='skills']"
    wait_after: 500ms

  - id: show-skills
    type: caption
    text: "Three skills shipped — and your agent can call any of them"
    duration: 4s

  - id: click-linkedin
    type: click_selector
    selector: "[data-amem-skill='linkedin-inbox'] button.run"
    wait_for: "[data-amem-skill='linkedin-inbox'] .digest"

  - id: settle
    type: wait
    duration: 2s

  - id: terminal-finish
    type: terminal
    cmd: "amem recall --json 'attention is all you need'"
    show: "stdout"
    duration: 4s

2. Step types

Each step is one variant; steps execute sequentially (no v0.1 parallelism).

TypePurposeDriver
navigateLoad URL in active tabchrome_navigate over bridge
click_selectorClick an elementchrome_click
waitPause for N ms / N slibrarian sleep
wait_for (also a field on other steps)Block until selector / URL appearschrome_wait
extract_domRead DOM, save to script vars (for later assertions/captions)chrome_extract
captionRender an overlay caption for N secondslibrarian draws into a borderless overlay window
terminalRun a binary command, optionally render its stdout in an overlaylibrarian process + overlay

The terminal step is interesting: it lets a recording show CLI usage alongside the browser. The librarian opens a small floating window (via its own UI process, not the Chrome window) that renders the command and its stdout in monospace; screencapture picks up the overlay because we target the Chrome window but composite the overlay on top of it before each frame. Implementation detail: macOS lets us position a borderless NSWindow above the target Chrome window and screencapture -l<chromeWinId> will include it (Quartz compositing, same as visible-tab capture).

If the overlay approach proves unreliable across macOS versions, fallback v0.1: split the recording — capture Chrome window for browser steps, capture full screen for terminal steps, stitch in post via video-use. Decision deferred until we hit a Sonoma/Sequoia regression in testing.

caption steps render a similar borderless overlay at the bottom-center of the Chrome window with a translucent black bar.

3. macOS implementation — window-level screencapture

Core command:

screencapture -v -l <chromeWindowId> -V <duration_seconds> output.mov
  • -v — start recording immediately (no UI)
  • -l <windowId> — target a specific window. We get this from CGWindowListCopyWindowInfo(.optionOnScreenOnly) filtered to bundle identifier com.google.Chrome (or Brave / Arc / Edge variants — we hardcode the major Chromium IDs).
  • -V <seconds> — fixed duration. v0.1 sets this to duration_target + 10s buffer; we kill the process early if all steps finish before the timer.
  • Output .mov; converted to .mp4 via ffmpeg -i in.mov -c:v libx264 -crf 20 out.mp4 post-recording.

Why window-level not full-screen:

  • Privacy: full-screen captures the user’s desktop background, dock, notifications, other apps. Window-level captures only the Chrome window’s pixel rect.
  • Aesthetics: window-level avoids us cropping the recording in post.
  • Privacy posture matches the rest of amem (“data stays local; we don’t capture what we don’t need”).

Permission flow: macOS requires Screen Recording permission. On first invoke, the librarian prompts the user via TCC. If denied, we return { code: "SCREEN_RECORDING_DENIED", how_to_fix } so the agent can surface the system-settings link.

Linux / Windows: out of scope for v0.1. The recording skill is gated to macOS in v0.1 (#[cfg(target_os = "macos")]); on other OSes the skill renders as kind:"error" with “Recording requires macOS in v0.1.” Linux pipeline (probably wf-recorder for Wayland + scrot/ffmpeg-x11grab for X11) is a follow-up RFC.

4. video-use integration (optional post-processor)

The raw mp4 from §3 is usable as-is — but for marketing-grade output we want transcription, filler-word cutting, captions, and consistent encoding. video-use is the operator’s existing tool for that; integrating is a one-liner:

#![allow(unused)]
fn main() {
// crates/amem-librarian/src/skills/record_demo.rs
async fn record_demo(script: &Script) -> Result<RecordingOutput> {
    let raw_mov = run_screencapture(&script.window, script.duration_target).await?;
    let raw_mp4 = ffmpeg_convert(&raw_mov).await?;

    let final_mp4 = if script.post_process.unwrap_or(false) {
        video_use::process(&raw_mp4, &script.post_options).await?
    } else {
        raw_mp4
    };

    Ok(RecordingOutput {
        amem_uri: format!("amem://recording/{}", uuid),
        mp4_path: final_mp4,
    })
}
}

post_process: false is the v0.1 default (raw recording). Setting post_process: true in the YAML opts in. video-use is treated as an optional dependency: if it isn’t installed, the field is ignored with a warning.

The script’s post_options mirror video-use’s CLI flags one-to-one (so the integration stays a thin shim, not a redesign).

5. amem_record_demo MCP tool

Exposed indirectly through amem_invoke_skill("cws-demo") per RFC-001 §3d, but also as a direct tool for ad-hoc agent use:

amem_record_demo(
  script_path?:   string,        // path to YAML
  script_inline?: string,        // YAML literal (mutually exclusive with script_path)
  output_dir?:    string,        // defaults to ~/.amem/recordings/
)
  -> {
    amem_uri:    "amem://recording/<uuid>",
    mp4_path:    "/Users/.../<uuid>.mp4",
    duration_s:  number,
    steps_run:   number,
  }

Errors:

CodeMeaning
SCREEN_RECORDING_DENIEDmacOS TCC denied screen recording
BRIDGE_UNAVAILABLECannot reach amem Clipper to drive Chrome
WINDOW_NOT_FOUNDNo Chrome window matched script.window
STEP_FAILEDSome step failed; partial recording saved with partial: true
RECORDING_IN_PROGRESSAnother recording is active; refuse to start a second
UNSUPPORTED_OSNot macOS in v0.1

The tool is blocking: it returns when the recording finishes (typically 30–90s for v0.1 scripts). MCP clients already render “long-running tool call” affordances.

6. Use cases

The same pipeline serves three concrete needs, each justifying ship in v0.1:

6a. CWS promo video (the SPEC’s “own best demo” principle)

The CWS listing needs a 30–60 second demo. We script it (§1), the librarian records it, video-use polishes it, we upload. When v0.1.x patches change the UI, we re-run the same script. The demo never goes stale.

6b. Feature update demos (post-v0.1)

When v0.2 ships custom-skill installation, we want a demo of that flow. Same pipeline, new YAML script. v0.2 itself is a script.

6c. Tutorials

docs.amem.sh user guide pages can embed inline mp4s recorded from canonical YAML scripts in crates/amem-librarian/builtin_scripts/. When the UI changes, regenerate.

7. Why the librarian records, not the clipper

The cleanest framing of this RFC’s central decision:

OptionRecords sidepanel UI?Privacy?Lifecycle?Cross-OS path?
A. chrome.tabCapture (extension)❌ No (sidepanel is out-of-tab)✅ tab-only⚠️ MV3 service worker eviction risk✅ identical everywhere
B. getDisplayMedia (extension)✅ User picks the window✅ user opt-in per recording⚠️ same MV3 risks✅ uniform
C. macOS window screencapture (librarian)✅ entire Chrome window✅ window-only, no desktop leak✅ Rust process is long-lived❌ macOS-only v0.1

Option B is the second-best choice; we considered it seriously. We rejected it for v0.1 because:

  1. getDisplayMedia requires a user picker dialog every recording — incompatible with a “click Run, get a recording” agent-driven flow.
  2. The MV3 service worker would still need to coordinate with the librarian for the script driver, doubling the moving parts.
  3. We want the librarian to own the timeline anyway — it already owns the script, the bridge, the storage, the post-process step. Adding the recording itself keeps responsibility in one place.

Option C costs us OS portability in v0.1 — accepted tradeoff. Linux / Windows recording lands when there is a non-macOS user.

Privacy

  • Window-level screencapture: only the Chrome window’s pixels are captured. Desktop, dock, menu bar, other apps, notifications — all outside the frame.
  • The recorded mp4 lives in ~/.amem/recordings/<uuid>.mp4. Never uploaded by the librarian. The user can amem upload (future) or manually drag-drop to a CWS listing.
  • macOS Screen Recording permission is requested on first invoke; the user can revoke at any time in System Settings.
  • Caption / terminal overlays render the script’s literal text; no enrichment, no LLM rewriting (v0.1).
  • video-use (when used) is a local binary; transcription runs locally via whisper. No cloud calls in the default config.
  • The script itself is recorded into the mp4 metadata’s comment field (so anyone with the mp4 can see what was scripted). The user can disable with metadata: false in the YAML.

Failure modes

ModeCauseMitigation
Screen Recording permission deniedFirst-run user hasn’t granted TCCTool returns SCREEN_RECORDING_DENIED with open System Settings → Privacy → Screen Recording instruction; sidepanel card shows the same
Chrome window not foundUser has no Chrome window open, or app is Brave/ArcYAML window.app allow-list extended to common Chromium variants; tool errors with WINDOW_NOT_FOUND listing detected windows
Step times outSelector never appearsStep times out at wait_for budget (default 30s); recording stops; partial: true flag in result so user knows
Bridge disconnects mid-recordingLibrarian-clipper WS dropsRecording continues (it’s screencapture, not bridge-driven pixels), but subsequent steps can’t drive Chrome; we abort gracefully and save what we have
screencapture produces 0-byte fileKnown macOS bug in some Sonoma builds when target window is fully occludedPre-check: bring Chrome to front + verify visibility before starting; document the Apple bug in docs/troubleshooting.md
Multiple recordings requested concurrentlyTwo agents both call amem_record_demoLibrarian-wide recording lock; second call gets RECORDING_IN_PROGRESS
Output mp4 hugeHigh-resolution display, long demoDefault resolution: source for v0.1; users can opt to 1080p / 720p in YAML; ffmpeg downscale handled in §4
Terminal overlay flickers / drops framesmacOS compositor under loadDocumented; user can switch to “split capture + post-stitch” mode
Wrong Chrome window picked (multi-window)User has 3 Chrome windows; we pick wrongYAML window.match accepts title_regex; default is “active window” which uses the focused one at recording start

Concrete work

In rough order:

  1. (amem-librarian) YAML script parser + schema validator — ~0.5d
  2. (amem-librarian) macOS window-id resolver via Quartz CGWindowListCopyWindowInfo (FFI through core-graphics crate) — ~0.5d
  3. (amem-librarian) screencapture driver: spawn, monitor, kill-early on completion — ~0.5d
  4. (amem-librarian) Step executor: bridge-driven navigate / click_selector / wait_for / extract_dom (mostly reuses RFC-001 bridge wrappers) — ~1d
  5. (amem-librarian) Caption / terminal overlay window (NSWindow with borderless Cocoa view) — ~1.5d (riskiest step; if compositing misbehaves, fall back to post-stitch mode and trim to ~0.5d)
  6. (amem-librarian) ffmpeg mp4 conversion + optional video-use handoff — ~0.5d
  7. (amem-librarian) amem_record_demo MCP tool surface + structured errors — ~0.5d
  8. (amem-librarian) Built-in cws-demo.yaml script + smoke test on the actual amem Clipper UI — ~1d (this is the dogfood — script will reveal UI bugs)
  9. (docs.amem.sh) Replace guide/self-recording.md content with pointers to this RFC + how-to for users — ~0.25d

Total: ~6.25d (5d if overlay step falls back to post-stitch).

Rejected alternatives

  • Keep chrome.tabCapture from offscreen documents. Cannot capture sidepanel; demo would have to fake the catalog UI. Disqualifying.
  • Use getDisplayMedia from the extension. User picker dialog every time; service worker lifecycle headaches; doubles the moving parts. See §7 Option B.
  • Run a generic OBS / ffmpeg pipe and let the user start/stop manually. Loses the “agent-driven scripted demo” capability that makes self- recording load-bearing. We’d be back to manual demo production.
  • Embed a video-use clone in the librarian. video-use is its own product with its own scope; reimplementing transcription / cut-filler in amem is out of scope. Optional handoff is the right boundary.
  • Ship Linux / Windows recording in v0.1. Tripled scope for an audience we don’t currently have. macOS-only is acceptable v0.1 posture.
  • Record from a JS-only stack (puppeteer + ffmpeg-screen). Loses the user’s Chrome profile and login state (per RFC-001 §3a / browser automation rules). Non-starter.
  • Skip captions / terminal overlays. Dropping captions makes the resulting mp4 unsuitable for CWS listings without manual editing, which defeats the whole “own best demo” pipeline. Worth the implementation cost.

Open questions

  • Overlay window rendering technology. Native NSWindow + Cocoa view (proposed) or a small SwiftUI helper app the librarian shells out to? Soft preference: Cocoa from Rust via the cocoa / objc2 crates; SwiftUI helper if FFI gets miserable.
  • Caption style. Plain translucent black bar with white text v0.1; do we offer themes / fonts? Soft preference: defer to v0.2; one hardcoded style in v0.1.
  • Whether to embed the YAML script in the mp4 metadata by default. Useful for reproducibility, mildly leaky for privacy if the user shares the mp4. Soft preference: opt-in (metadata: true in YAML).
  • video-use handoff API stability. v0.1 calls it as a binary subprocess; if video-use grows a Rust crate API later, switch over. Not blocking.
  • Recording lock granularity. Per-librarian (current proposal) or per-script-id? Soft preference: per-librarian (simpler; scripted recording is sequential by nature).
  • Should the sidepanel show a live preview of the recording? Tempting but adds complexity (need to read the in-progress mp4 or echo screencapture frames). Defer to v0.2.

Roll-out

  • Day 4 of v0.1 build: YAML parser + window resolver + screencapture driver land. Smoke test: record a 5-second blank capture.
  • Day 5: step executor + bridge wiring. Smoke test: 10-second scripted recording opens arxiv, clicks something, finishes.
  • Day 6: caption / terminal overlay. Risk day; fallback path ready.
  • Day 7 (CWS submission day): full cws-demo.yaml runs, video-use polishes, we have the listing video. Submit.
  • Post-v0.1: track real script usage. When the third user-script exists (per RFC-001 §6 rule-of-three trigger), re-evaluate whether YAML schema needs versioning, whether scripts need a UI, whether Linux support belongs in v0.2.
  • Future: cloud-render mode (a remote machine runs the script and returns the mp4) for users without macOS. Out of scope here; tracked in a future RFC.