Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

RFC-001 — Function-based v0.1 architecture

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-05-09
  • Related: RFC-002 (Clipper skills catalog UI), RFC-003 (recording orchestration), archived RFC-001 (bridge-first), SPEC.md § Mental model

TL;DR

Ship v0.1 as hardcoded MCP tools and Rust functions, not a skill engine. The three must-have features all have the same shape — agent calls a tool, librarian dispatches to a Rust function, function may bounce DOM-side work through the bridge to amem Clipper. There is no plugin runtime, no DSL, no sandbox. Skill engine is deferred to v0.5+, gated on the rule of three: ship when we have 5+ site adapters, 3+ recording scripts, or external user requests for installable skills.

7.5 days of hardcoded code now beats 13 days of skill-engine plumbing for zero current users. We promote to a real engine the day the third copy of “this looks the same as the previous one” arrives.

Motivation

The previous round of RFCs (archived 2026-05-09) circled around a generic “skills” concept — pluggable units the agent could discover and invoke. The abstraction was attractive on paper (catalog, sandbox, manifest, capability flags) and load-bearing on nothing: we have one user (the operator), zero external testers, zero shipped sites, and a CWS deadline.

Three forces pushed us to lock the architecture as function-based:

  1. No third instance. Out of the v0.1 must-haves, only the site-adapter family has any repetition (arxiv, github, hackernews — three sites, one shape). Recording is one-of-one; bridge-driven Chrome ops are one-of-one. The rule of three (Refactoring, Fowler) says abstract on the third occurrence — we have it for adapters and only for adapters, and even there the shape is “match URL → call extractor,” which is a Rust match block, not a runtime.
  2. Engineering cost is dispositive. A plausible v0.1 skill engine — manifest schema, loader, sandboxed JS runtime in the librarian, capability gating, registry, sidepanel discovery, error surfaces — is ~13 days. A function-based v0.1 is ~7.5 days (3 site adapters at 1d each, recording at 2d, bridge MCP wiring at 1.5d, polish at 1d). The 5.5-day delta is the entire CWS submission window.
  3. CWS positioning still works without an engine. The Chrome Web Store listing claims “the first agent-callable Chrome skills catalog” — the user-facing surface (RFC-002) renders three hardcoded skills as if they are an installable catalog. v0.1 ships the shape of the product, v0.2 ships the actual installability. The gap is honestly disclosed in-UI (“Custom skills coming v0.2”) so we’re not lying — we’re shipping the minimum that lets the rest of the story land.

The core insight: the operator (and Claude Code, when it’s the operator acting on the user’s behalf) does not care whether chrome_navigate is implemented as a Rust function or as a sandboxed JS skill. They care that the tool exists and works. Build the tools first. Generalise on the third occurrence.

Proposal

1. Component split — librarian (Rust) vs clipper (JS)

Two binaries, one bridge, one rule for splitting work:

       ┌────────────────────────────────────────────────────────┐
       │  Frontier model (Claude / GPT / Gemini)                 │
       └────────────────────────────────────────────────────────┘
                              │  MCP (stdio)
                              ▼
       ┌────────────────────────────────────────────────────────┐
       │  amem Librarian (Rust binary, formerly amem-sh)         │
       │  • All heavy work: capture pipeline, compile, fact-     │
       │    check, recording orchestration, site adapters,       │
       │    storage, OCR, screencapture                          │
       │  • Hosts MCP server                                     │
       │  • Hosts bridge server (loopback WS 7600)               │
       └────────────────────────────────────────────────────────┘
                              │  ws://127.0.0.1:7600
                              ▼
       ┌────────────────────────────────────────────────────────┐
       │  amem Clipper (Chrome MV3 extension, JS)                │
       │  • DOM-only work: read/write document, observe URL,     │
       │    capture visible tab, render sidepanel UI             │
       │  • Holds zero authoritative state (config lives in      │
       │    librarian's config.toml)                             │
       │  • No fetching, no parsing, no LLM calls,               │
       │    no file I/O — those live in the librarian            │
       └────────────────────────────────────────────────────────┘

Rule of thumb: if a step requires the file system, an LLM, ffmpeg, OCR, or a network fetch beyond the active tab, it belongs in the librarian. If it requires a DOM, the active page, or pixel-level capture of the active tab, it belongs in the clipper. Anything else (URL matching, string manipulation, scheduling) belongs in the librarian — the clipper is a sensor, not a brain (per archived RFC-001’s positioning, still in force).

2. Bridge protocol — action / params / id envelope

Single loopback WebSocket on 127.0.0.1:7600. All messages, both directions, share one envelope:

{
  "id":     "<uuid-v4>",
  "action": "<verb>",
  "params": { ... },
  "token":  "<bridge token from ~/.amem/bridge.token>"
}

Replies use the same shape, with "action": "<verb>_result" and "id" matching the request. Errors carry {"action":"error","id":"...","error":{ "code","message"}}. The same envelope handles librarian→clipper commands (e.g. chrome_navigate) and clipper→librarian events (e.g. auto_capture).

Initial verb set for v0.1:

DirectionVerbPurpose
L → Cchrome_navigateNavigate active tab to URL
L → Cchrome_clickClick DOM element by selector
L → Cchrome_extractRead DOM (selector → text/html/attrs)
L → Cchrome_screenshottabs.captureVisibleTab (best-effort, see §6)
L → Cchrome_waitWait for selector to appear / URL pattern match
L → Center_recording_overlayToggle “recording” sidepanel state for demos
C → Lauto_captureSidepanel-toggled URL match fired; please ingest
C → Linvoke_skillUser clicked a Run-button in the sidepanel
L ↔ Cping / pongLiveness

Security posture inherits from archived RFC-001 unchanged: loopback bind, origin allow-list (only our extension IDs), 32-byte token in ~/.amem/bridge.token (mode 0600), token rotates on librarian restart.

3. The three must-have features → tool + function map

Each feature is one MCP tool (or a small fixed set), each backed by one Rust function in the librarian. No registry, no plugins.

3a. Claude Code operates already-logged-in Chrome

MCP tools shipped:

chrome_navigate(url: string) -> { status, final_url }
chrome_click(selector: string, opts?: { wait_after_ms }) -> { ok }
chrome_extract(selector: string, kind: "text"|"html"|"attrs") -> { value }
chrome_wait(selector: string, timeout_ms: number) -> { ok }
chrome_screenshot(selector?: string) -> { png_path }

Implementation in librarian:

#![allow(unused)]
fn main() {
// crates/amem-librarian/src/tools/chrome.rs
pub async fn chrome_navigate(url: &str) -> Result<NavigateOut> {
    let req = BridgeReq::new("chrome_navigate", json!({ "url": url }));
    BRIDGE.send(req).await?.into()
}
}

The librarian is just a thin pass-through here; the clipper does the actual DOM work in chrome.scripting.executeScript. This is the only path — agents do not get raw access to the bridge or to the extension. The MCP boundary is the contract.

Why not use playwright / puppeteer / chrome-devtools-protocol? Because the operator’s authenticated Chrome (Gmail, LinkedIn, internal tools) is the whole point. CDP-based tooling either drops the user’s profile (Chrome 136+ --remote-debugging-port + --user-data-dir constraints) or breaks keychain-backed flows (Chrome for Testing’s ad-hoc signing). Going through amem Clipper, which lives inside the user’s real Chrome, is the only path that doesn’t lose login state. (See ~/.claude/CLAUDE.md § Browser Automation for the gory details.)

3b. URL-mentioned content auto-saves to wiki

MCP tool shipped:

amem_capture_url(url: string, mode?: "auto"|"force") -> {
  cite_key, amem_uri, status: "captured"|"already_have"|"unsupported"
}

Two trigger paths feeding the same function:

  1. Agent-side: any agent (Claude Code, Cursor) that mentions a URL in its reply calls amem_capture_url(url) directly via MCP. The librarian matches the URL against the hardcoded adapter table and dispatches.
  2. Browser-side: amem Clipper observes navigation events; if the URL matches a hardcoded adapter pattern and the user has the matching skill toggle on (RFC-002), clipper sends auto_capture over the bridge, which calls the same function.

Hardcoded adapters in crates/amem-librarian/src/adapters/:

#![allow(unused)]
fn main() {
// crates/amem-librarian/src/adapters/mod.rs
pub fn dispatch(url: &Url) -> Option<Box<dyn Adapter>> {
    match url.host_str()? {
        "arxiv.org" | "www.arxiv.org"          => Some(Box::new(arxiv::Arxiv)),
        "github.com"                           => Some(Box::new(github::Github)),
        "news.ycombinator.com"                 => Some(Box::new(hackernews::HackerNews)),
        _                                      => None,
    }
}

trait Adapter {
    async fn extract(&self, url: &Url) -> Result<CaptureRecord>;
}
}

arxiv.rs, github.rs, hackernews.rs each ~150–250 lines, each a straightforward fetch + parse. They share a tiny CaptureRecord struct; they do not share a runtime. When the fourth adapter ships, we revisit (see §6 “When to revisit”).

3c. Scripted Chrome-only recording

MCP tool shipped:

amem_record_demo(script_path?: string, script_inline?: string)
  -> { amem_uri: "amem://recording/<uuid>", mp4_path }

Full mechanics live in RFC-003. From this RFC’s point of view it is one more Rust function in the librarian — record_demo() — that:

  1. parses a YAML script,
  2. uses the bridge to drive the clipper through navigation/click/wait steps,
  3. simultaneously runs screencapture -v -l<chromeWindowId> (macOS window-level capture) so the recording covers the whole Chrome window including the clipper’s sidepanel UI,
  4. optionally hands the raw mp4 to video-use for post-processing.

3d. Catalog management (sidepanel needs to list things)

Two more MCP tools so the sidepanel and any agent can ask “what’s available”:

amem_list_skills() -> [{ id, name, description, kind: "auto"|"run", state }]
amem_invoke_skill(id: string, params?: object) -> { result }

For v0.1 these read from a hardcoded Vec<SkillCard> in the librarian. There is no manifest. There is no sandbox. amem_invoke_skill("cws-demo") literally calls record_demo(builtin_scripts::CWS_DEMO). The point is the tool surface: the moment we have a skill engine in v0.5+, these two tools’ signatures don’t change — only their implementation does. The agent contract is forward-compatible.

4. MCP tool inventory for v0.1

Total: 6 new tools ship in v0.1 (plus the four existing ones from Day 1: amem_capture, amem_compile, amem_cite, amem_recall).

ToolBacking functionNotes
chrome_navigatetools::chrome::navigateBridge pass-through
chrome_clicktools::chrome::clickBridge pass-through
chrome_extracttools::chrome::extractBridge pass-through
amem_capture_urladapters::dispatch + captureHardcoded site match
amem_invoke_skillskills::runHardcoded skill table
amem_list_skillsskills::listHardcoded skill table

amem_record_demo is exposed as a kind:"run" skill via amem_invoke_skill("cws-demo"), not as a separate top-level tool — keeping the agent-facing surface tighter and letting RFC-003 own the contract.

5. Why NOT skill engine (yet)

Five reasons, in priority order:

  1. Rule of three. Of the three v0.1 features, only one (site adapters) has even three instances. The other two are one-of-one. Abstracting now means designing for assumptions we have not yet earned.
  2. One operator. A skill engine optimises for external authors shipping skills. We have zero external authors. The first user benefitting from sandboxing-vs-trust would be us — and we trust our own code more than we trust a sandbox we just wrote.
  3. 13 vs 7.5 days. A real skill engine needs: manifest schema with versioning, loader with capability gating, sandboxed JS runtime (boa/quickjs) with controlled host bindings, error surfaces, registry, discovery, version pinning, update flow. None of this work helps the v0.1 user-facing demo.
  4. CWS deadline lives in this window. Chrome Web Store review can take 2–10 days. Submit narrow + working > submit broad + half-built. The 5.5-day delta is bigger than the review buffer.
  5. Forward-compatible API. amem_list_skills / amem_invoke_skill already shape the surface a future engine will use. We are not painting ourselves into a corner; we are shipping the same MCP signatures that v0.5+ will reuse.

6. When to revisit — the rule of three

Promote to a real engine when any of the following triggers fire:

TriggerWhat it meansWhat to revisit
5+ site adaptersThe match on url.host_str() has grown to 5+ arms with similar shapeExtract a SiteAdapter trait + manifest table
3+ recording scripts in productionYAML scripts have proven pattern: navigate, click, capture, repeatPromote YAML schema to versioned spec; consider script-side templating
1+ external user requestSomeone outside yiidtw/ asks “can I write my own skill”Engine becomes a user-facing feature, not internal cleanup
Custom skill in v0.2 plan firms upRFC-002 promises “Custom skills coming v0.2” — the moment that shipsEngine is the implementation

Until any trigger fires, the function-based approach is the correct endpoint, not a placeholder.

Privacy

Inherits the archived-RFC-001 posture unchanged:

  • All capture data lives in ~/.amem/, never uploaded.
  • Bridge is loopback only, token-authed, origin-checked.
  • The new MCP tools (chrome_*) execute in the user’s real Chrome under the user’s existing permissions — they cannot access tabs the user is not already authenticated to.
  • Recording (RFC-003) uses macOS window-level screencapture targeting the Chrome window’s window ID; desktop and other apps are not in frame.

One new consideration: chrome_extract returns DOM content to the librarian, which may surface to an agent over MCP. This is identical in sensitivity to today’s amem_capture(url), which already fetches and parses page content. Document the equivalence in the user guide.

Failure modes

ModeCauseMitigation
Bridge unreachable when agent calls chrome_*Librarian not running, or extension not connectedMCP tool returns structured error { code: "BRIDGE_UNAVAILABLE", install_hint }; agent surfaces install CTA
Selector not foundPage changed, login required, racechrome_click / chrome_extract return { ok: false, reason } rather than throw; agent retries with chrome_wait
Adapter doesn’t match URLURL outside hardcoded listamem_capture_url returns { status: "unsupported" }; agent can fall back to the generic amem_capture (existing Day 1 tool)
Skill ID typoAgent calls amem_invoke_skill("does-not-exist")Structured error with available_ids list (read from amem_list_skills)
Concurrent recording + chrome opsRecording and other MCP tools racing for the bridgeRecording acquires a recording-mode lock; concurrent chrome_* calls return { code: "RECORDING_IN_PROGRESS" }
Multi-tab / multi-window ambiguity“Active tab” is ambiguous on multi-window setupsDefault to focused-window’s active tab; expose tabId param later if it bites
Skill state drift between sidepanel and librarianUser toggles in sidepanel while CLI also togglesLibrarian is single source of truth (config.toml); sidepanel reads on every render via amem_list_skills

Concrete work

In rough order of dependency. Estimates are pessimistic-realistic.

  1. (amem-librarian) Bridge envelope + token + origin allow-list — ~1d
  2. (amem-librarian) MCP wrappers for chrome_navigate / _click / _extract / _wait / _screenshot — ~0.5d
  3. (amem-clipper) Bridge client + handlers for above verbs via chrome.scripting.executeScript — ~1.5d
  4. (amem-librarian) Adapters: arxiv.rs, github.rs, hackernews.rs
    • dispatch table + amem_capture_url MCP tool — ~2d
  5. (amem-librarian) Hardcoded skill catalog (SkillCard, the three v0.1 entries) + amem_list_skills / amem_invoke_skill MCP tools — ~0.5d
  6. (amem-clipper) Sidepanel “Skills” tab rendering catalog (full detail in RFC-002) — ~1d
  7. (amem-librarian) record_demo() skeleton stub returning not implemented — full implementation lives in RFC-003 — ~0.5d
  8. (docs.amem.sh) guide/skills.md documenting the v0.1 hardcoded skills + the v0.2 forward-compat note — ~0.5d

Total: ~7.5d. Compare to ~13d for an equivalent skill-engine v0.1.

Rejected alternatives

  • Ship a real skill engine in v0.1. Costs 5.5d we don’t have, optimises for a user we don’t yet have. See §5.
  • Ship skill engine manifest only, run hardcoded for v0.1. Tempting middle ground, but commits us to a manifest schema before we know what fields second-and-third skills will need. Premature schema lock.
  • Embed skills as JS in the clipper extension. Moves heavy work into the extension (LLM calls, file I/O), which violates the clipper-is-sensor positioning and runs into MV3 storage / CSP limits. The librarian must stay the brain.
  • Use playwright / puppeteer / Chrome DevTools Protocol for §3a. Loses the user’s Chrome profile (login state, passkeys, 2FA). The bridge-into-real-Chrome path is non-negotiable.
  • Skip the catalog UI, ship MCP tools only. Loses the CWS positioning (“first agent-callable Chrome skills catalog”). The catalog is the user-facing story; without it we’re just another Chrome extension.

Open questions

  • Multi-tab targeting. Should chrome_* tools default to the active tab in the focused window (proposed) or accept an explicit tabId? Soft preference: default-active for v0.1, expose tabId when the second use case asks.
  • Adapter timeouts. What does amem_capture_url do if arxiv is slow / down? Soft preference: 30s timeout, returns { status: "captured_partial", reason: "fetch_timeout" } so the agent can retry later.
  • MCP tool surface stability. The 6 new tool names are committed. We will add tools post-v0.1 but not rename or remove. The surface is the contract.
  • Should amem_list_skills include kind:"hidden" for skills used internally (e.g. by other tools) but not surfaced in the sidepanel? Open. v0.1 ships without; revisit when an internal skill emerges.

Roll-out

  • Day 1 (today): this RFC + RFC-002 + RFC-003 land. SUMMARY.md updated.
  • Day 2: bridge envelope + chrome_* MCP tools + clipper handlers.
  • Day 3: site adapters land. amem_capture_url ships.
  • Day 4: skill catalog + sidepanel UI (RFC-002 implementation).
  • Day 5: recording skeleton lands (RFC-003 implementation).
  • Day 6: end-to-end CWS demo recorded by amem itself, dogfooded.
  • Day 7: CWS submission. Listing copy claims “first agent-callable Chrome skills catalog” honestly — three skills shipped, custom-skills disclosure visible.
  • Post-v0.1: track the rule-of-three triggers in §6. When any fires, open RFC-00X for the skill engine.