RFC-001 — Function-based v0.1 architecture
- Status: Draft
- Authors: @yiidtw
- Created: 2026-05-09
- Related: RFC-002 (Clipper skills catalog UI), RFC-003 (recording orchestration), archived RFC-001 (bridge-first), SPEC.md § Mental model
TL;DR
Ship v0.1 as hardcoded MCP tools and Rust functions, not a skill engine. The three must-have features all have the same shape — agent calls a tool, librarian dispatches to a Rust function, function may bounce DOM-side work through the bridge to amem Clipper. There is no plugin runtime, no DSL, no sandbox. Skill engine is deferred to v0.5+, gated on the rule of three: ship when we have 5+ site adapters, 3+ recording scripts, or external user requests for installable skills.
7.5 days of hardcoded code now beats 13 days of skill-engine plumbing for zero current users. We promote to a real engine the day the third copy of “this looks the same as the previous one” arrives.
Motivation
The previous round of RFCs (archived 2026-05-09) circled around a generic “skills” concept — pluggable units the agent could discover and invoke. The abstraction was attractive on paper (catalog, sandbox, manifest, capability flags) and load-bearing on nothing: we have one user (the operator), zero external testers, zero shipped sites, and a CWS deadline.
Three forces pushed us to lock the architecture as function-based:
- No third instance. Out of the v0.1 must-haves, only the site-adapter
family has any repetition (arxiv, github, hackernews — three sites,
one shape). Recording is one-of-one; bridge-driven Chrome ops are
one-of-one. The rule of three (Refactoring, Fowler) says abstract on the
third occurrence — we have it for adapters and only for adapters, and
even there the shape is “match URL → call extractor,” which is a Rust
matchblock, not a runtime. - Engineering cost is dispositive. A plausible v0.1 skill engine — manifest schema, loader, sandboxed JS runtime in the librarian, capability gating, registry, sidepanel discovery, error surfaces — is ~13 days. A function-based v0.1 is ~7.5 days (3 site adapters at 1d each, recording at 2d, bridge MCP wiring at 1.5d, polish at 1d). The 5.5-day delta is the entire CWS submission window.
- CWS positioning still works without an engine. The Chrome Web Store listing claims “the first agent-callable Chrome skills catalog” — the user-facing surface (RFC-002) renders three hardcoded skills as if they are an installable catalog. v0.1 ships the shape of the product, v0.2 ships the actual installability. The gap is honestly disclosed in-UI (“Custom skills coming v0.2”) so we’re not lying — we’re shipping the minimum that lets the rest of the story land.
The core insight: the operator (and Claude Code, when it’s the operator
acting on the user’s behalf) does not care whether chrome_navigate is
implemented as a Rust function or as a sandboxed JS skill. They care that
the tool exists and works. Build the tools first. Generalise on the third
occurrence.
Proposal
1. Component split — librarian (Rust) vs clipper (JS)
Two binaries, one bridge, one rule for splitting work:
┌────────────────────────────────────────────────────────┐
│ Frontier model (Claude / GPT / Gemini) │
└────────────────────────────────────────────────────────┘
│ MCP (stdio)
▼
┌────────────────────────────────────────────────────────┐
│ amem Librarian (Rust binary, formerly amem-sh) │
│ • All heavy work: capture pipeline, compile, fact- │
│ check, recording orchestration, site adapters, │
│ storage, OCR, screencapture │
│ • Hosts MCP server │
│ • Hosts bridge server (loopback WS 7600) │
└────────────────────────────────────────────────────────┘
│ ws://127.0.0.1:7600
▼
┌────────────────────────────────────────────────────────┐
│ amem Clipper (Chrome MV3 extension, JS) │
│ • DOM-only work: read/write document, observe URL, │
│ capture visible tab, render sidepanel UI │
│ • Holds zero authoritative state (config lives in │
│ librarian's config.toml) │
│ • No fetching, no parsing, no LLM calls, │
│ no file I/O — those live in the librarian │
└────────────────────────────────────────────────────────┘
Rule of thumb: if a step requires the file system, an LLM, ffmpeg, OCR, or a network fetch beyond the active tab, it belongs in the librarian. If it requires a DOM, the active page, or pixel-level capture of the active tab, it belongs in the clipper. Anything else (URL matching, string manipulation, scheduling) belongs in the librarian — the clipper is a sensor, not a brain (per archived RFC-001’s positioning, still in force).
2. Bridge protocol — action / params / id envelope
Single loopback WebSocket on 127.0.0.1:7600. All messages, both directions,
share one envelope:
{
"id": "<uuid-v4>",
"action": "<verb>",
"params": { ... },
"token": "<bridge token from ~/.amem/bridge.token>"
}
Replies use the same shape, with "action": "<verb>_result" and "id"
matching the request. Errors carry {"action":"error","id":"...","error":{ "code","message"}}. The same envelope handles librarian→clipper commands
(e.g. chrome_navigate) and clipper→librarian events (e.g. auto_capture).
Initial verb set for v0.1:
| Direction | Verb | Purpose |
|---|---|---|
| L → C | chrome_navigate | Navigate active tab to URL |
| L → C | chrome_click | Click DOM element by selector |
| L → C | chrome_extract | Read DOM (selector → text/html/attrs) |
| L → C | chrome_screenshot | tabs.captureVisibleTab (best-effort, see §6) |
| L → C | chrome_wait | Wait for selector to appear / URL pattern match |
| L → C | enter_recording_overlay | Toggle “recording” sidepanel state for demos |
| C → L | auto_capture | Sidepanel-toggled URL match fired; please ingest |
| C → L | invoke_skill | User clicked a Run-button in the sidepanel |
| L ↔ C | ping / pong | Liveness |
Security posture inherits from archived RFC-001 unchanged: loopback bind,
origin allow-list (only our extension IDs), 32-byte token in
~/.amem/bridge.token (mode 0600), token rotates on librarian restart.
3. The three must-have features → tool + function map
Each feature is one MCP tool (or a small fixed set), each backed by one Rust function in the librarian. No registry, no plugins.
3a. Claude Code operates already-logged-in Chrome
MCP tools shipped:
chrome_navigate(url: string) -> { status, final_url }
chrome_click(selector: string, opts?: { wait_after_ms }) -> { ok }
chrome_extract(selector: string, kind: "text"|"html"|"attrs") -> { value }
chrome_wait(selector: string, timeout_ms: number) -> { ok }
chrome_screenshot(selector?: string) -> { png_path }
Implementation in librarian:
#![allow(unused)]
fn main() {
// crates/amem-librarian/src/tools/chrome.rs
pub async fn chrome_navigate(url: &str) -> Result<NavigateOut> {
let req = BridgeReq::new("chrome_navigate", json!({ "url": url }));
BRIDGE.send(req).await?.into()
}
}
The librarian is just a thin pass-through here; the clipper does the actual
DOM work in chrome.scripting.executeScript. This is the only path —
agents do not get raw access to the bridge or to the extension. The MCP
boundary is the contract.
Why not use playwright / puppeteer / chrome-devtools-protocol? Because the
operator’s authenticated Chrome (Gmail, LinkedIn, internal tools) is the
whole point. CDP-based tooling either drops the user’s profile (Chrome 136+
--remote-debugging-port + --user-data-dir constraints) or breaks
keychain-backed flows (Chrome for Testing’s ad-hoc signing). Going through
amem Clipper, which lives inside the user’s real Chrome, is the only path
that doesn’t lose login state. (See ~/.claude/CLAUDE.md § Browser
Automation for the gory details.)
3b. URL-mentioned content auto-saves to wiki
MCP tool shipped:
amem_capture_url(url: string, mode?: "auto"|"force") -> {
cite_key, amem_uri, status: "captured"|"already_have"|"unsupported"
}
Two trigger paths feeding the same function:
- Agent-side: any agent (Claude Code, Cursor) that mentions a URL in
its reply calls
amem_capture_url(url)directly via MCP. The librarian matches the URL against the hardcoded adapter table and dispatches. - Browser-side: amem Clipper observes navigation events; if the URL
matches a hardcoded adapter pattern and the user has the matching
skill toggle on (RFC-002), clipper sends
auto_captureover the bridge, which calls the same function.
Hardcoded adapters in crates/amem-librarian/src/adapters/:
#![allow(unused)]
fn main() {
// crates/amem-librarian/src/adapters/mod.rs
pub fn dispatch(url: &Url) -> Option<Box<dyn Adapter>> {
match url.host_str()? {
"arxiv.org" | "www.arxiv.org" => Some(Box::new(arxiv::Arxiv)),
"github.com" => Some(Box::new(github::Github)),
"news.ycombinator.com" => Some(Box::new(hackernews::HackerNews)),
_ => None,
}
}
trait Adapter {
async fn extract(&self, url: &Url) -> Result<CaptureRecord>;
}
}
arxiv.rs, github.rs, hackernews.rs each ~150–250 lines, each a
straightforward fetch + parse. They share a tiny CaptureRecord struct;
they do not share a runtime. When the fourth adapter ships, we revisit
(see §6 “When to revisit”).
3c. Scripted Chrome-only recording
MCP tool shipped:
amem_record_demo(script_path?: string, script_inline?: string)
-> { amem_uri: "amem://recording/<uuid>", mp4_path }
Full mechanics live in RFC-003. From this RFC’s point of view it is one
more Rust function in the librarian — record_demo() — that:
- parses a YAML script,
- uses the bridge to drive the clipper through navigation/click/wait steps,
- simultaneously runs
screencapture -v -l<chromeWindowId>(macOS window-level capture) so the recording covers the whole Chrome window including the clipper’s sidepanel UI, - optionally hands the raw mp4 to video-use for post-processing.
3d. Catalog management (sidepanel needs to list things)
Two more MCP tools so the sidepanel and any agent can ask “what’s available”:
amem_list_skills() -> [{ id, name, description, kind: "auto"|"run", state }]
amem_invoke_skill(id: string, params?: object) -> { result }
For v0.1 these read from a hardcoded Vec<SkillCard> in the librarian.
There is no manifest. There is no sandbox. amem_invoke_skill("cws-demo")
literally calls record_demo(builtin_scripts::CWS_DEMO). The point is the
tool surface: the moment we have a skill engine in v0.5+, these two
tools’ signatures don’t change — only their implementation does. The
agent contract is forward-compatible.
4. MCP tool inventory for v0.1
Total: 6 new tools ship in v0.1 (plus the four existing ones from Day
1: amem_capture, amem_compile, amem_cite, amem_recall).
| Tool | Backing function | Notes |
|---|---|---|
chrome_navigate | tools::chrome::navigate | Bridge pass-through |
chrome_click | tools::chrome::click | Bridge pass-through |
chrome_extract | tools::chrome::extract | Bridge pass-through |
amem_capture_url | adapters::dispatch + capture | Hardcoded site match |
amem_invoke_skill | skills::run | Hardcoded skill table |
amem_list_skills | skills::list | Hardcoded skill table |
amem_record_demo is exposed as a kind:"run" skill via
amem_invoke_skill("cws-demo"), not as a separate top-level tool — keeping
the agent-facing surface tighter and letting RFC-003 own the contract.
5. Why NOT skill engine (yet)
Five reasons, in priority order:
- Rule of three. Of the three v0.1 features, only one (site adapters) has even three instances. The other two are one-of-one. Abstracting now means designing for assumptions we have not yet earned.
- One operator. A skill engine optimises for external authors shipping skills. We have zero external authors. The first user benefitting from sandboxing-vs-trust would be us — and we trust our own code more than we trust a sandbox we just wrote.
- 13 vs 7.5 days. A real skill engine needs: manifest schema with versioning, loader with capability gating, sandboxed JS runtime (boa/quickjs) with controlled host bindings, error surfaces, registry, discovery, version pinning, update flow. None of this work helps the v0.1 user-facing demo.
- CWS deadline lives in this window. Chrome Web Store review can take 2–10 days. Submit narrow + working > submit broad + half-built. The 5.5-day delta is bigger than the review buffer.
- Forward-compatible API.
amem_list_skills/amem_invoke_skillalready shape the surface a future engine will use. We are not painting ourselves into a corner; we are shipping the same MCP signatures that v0.5+ will reuse.
6. When to revisit — the rule of three
Promote to a real engine when any of the following triggers fire:
| Trigger | What it means | What to revisit |
|---|---|---|
| 5+ site adapters | The match on url.host_str() has grown to 5+ arms with similar shape | Extract a SiteAdapter trait + manifest table |
| 3+ recording scripts in production | YAML scripts have proven pattern: navigate, click, capture, repeat | Promote YAML schema to versioned spec; consider script-side templating |
| 1+ external user request | Someone outside yiidtw/ asks “can I write my own skill” | Engine becomes a user-facing feature, not internal cleanup |
| Custom skill in v0.2 plan firms up | RFC-002 promises “Custom skills coming v0.2” — the moment that ships | Engine is the implementation |
Until any trigger fires, the function-based approach is the correct endpoint, not a placeholder.
Privacy
Inherits the archived-RFC-001 posture unchanged:
- All capture data lives in
~/.amem/, never uploaded. - Bridge is loopback only, token-authed, origin-checked.
- The new MCP tools (
chrome_*) execute in the user’s real Chrome under the user’s existing permissions — they cannot access tabs the user is not already authenticated to. - Recording (RFC-003) uses macOS window-level screencapture targeting the Chrome window’s window ID; desktop and other apps are not in frame.
One new consideration: chrome_extract returns DOM content to the
librarian, which may surface to an agent over MCP. This is identical in
sensitivity to today’s amem_capture(url), which already fetches and
parses page content. Document the equivalence in the user guide.
Failure modes
| Mode | Cause | Mitigation |
|---|---|---|
Bridge unreachable when agent calls chrome_* | Librarian not running, or extension not connected | MCP tool returns structured error { code: "BRIDGE_UNAVAILABLE", install_hint }; agent surfaces install CTA |
| Selector not found | Page changed, login required, race | chrome_click / chrome_extract return { ok: false, reason } rather than throw; agent retries with chrome_wait |
| Adapter doesn’t match URL | URL outside hardcoded list | amem_capture_url returns { status: "unsupported" }; agent can fall back to the generic amem_capture (existing Day 1 tool) |
| Skill ID typo | Agent calls amem_invoke_skill("does-not-exist") | Structured error with available_ids list (read from amem_list_skills) |
| Concurrent recording + chrome ops | Recording and other MCP tools racing for the bridge | Recording acquires a recording-mode lock; concurrent chrome_* calls return { code: "RECORDING_IN_PROGRESS" } |
| Multi-tab / multi-window ambiguity | “Active tab” is ambiguous on multi-window setups | Default to focused-window’s active tab; expose tabId param later if it bites |
| Skill state drift between sidepanel and librarian | User toggles in sidepanel while CLI also toggles | Librarian is single source of truth (config.toml); sidepanel reads on every render via amem_list_skills |
Concrete work
In rough order of dependency. Estimates are pessimistic-realistic.
- (
amem-librarian) Bridge envelope + token + origin allow-list — ~1d - (
amem-librarian) MCP wrappers forchrome_navigate/_click/_extract/_wait/_screenshot— ~0.5d - (
amem-clipper) Bridge client + handlers for above verbs viachrome.scripting.executeScript— ~1.5d - (
amem-librarian) Adapters:arxiv.rs,github.rs,hackernews.rs- dispatch table +
amem_capture_urlMCP tool — ~2d
- dispatch table +
- (
amem-librarian) Hardcoded skill catalog (SkillCard, the three v0.1 entries) +amem_list_skills/amem_invoke_skillMCP tools — ~0.5d - (
amem-clipper) Sidepanel “Skills” tab rendering catalog (full detail in RFC-002) — ~1d - (
amem-librarian)record_demo()skeleton stub returningnot implemented— full implementation lives in RFC-003 — ~0.5d - (
docs.amem.sh)guide/skills.mddocumenting the v0.1 hardcoded skills + the v0.2 forward-compat note — ~0.5d
Total: ~7.5d. Compare to ~13d for an equivalent skill-engine v0.1.
Rejected alternatives
- Ship a real skill engine in v0.1. Costs 5.5d we don’t have, optimises for a user we don’t yet have. See §5.
- Ship skill engine manifest only, run hardcoded for v0.1. Tempting middle ground, but commits us to a manifest schema before we know what fields second-and-third skills will need. Premature schema lock.
- Embed skills as JS in the clipper extension. Moves heavy work into the extension (LLM calls, file I/O), which violates the clipper-is-sensor positioning and runs into MV3 storage / CSP limits. The librarian must stay the brain.
- Use playwright / puppeteer / Chrome DevTools Protocol for §3a. Loses the user’s Chrome profile (login state, passkeys, 2FA). The bridge-into-real-Chrome path is non-negotiable.
- Skip the catalog UI, ship MCP tools only. Loses the CWS positioning (“first agent-callable Chrome skills catalog”). The catalog is the user-facing story; without it we’re just another Chrome extension.
Open questions
- Multi-tab targeting. Should
chrome_*tools default to the active tab in the focused window (proposed) or accept an explicittabId? Soft preference: default-active for v0.1, exposetabIdwhen the second use case asks. - Adapter timeouts. What does
amem_capture_urldo if arxiv is slow / down? Soft preference: 30s timeout, returns{ status: "captured_partial", reason: "fetch_timeout" }so the agent can retry later. - MCP tool surface stability. The 6 new tool names are committed. We will add tools post-v0.1 but not rename or remove. The surface is the contract.
- Should
amem_list_skillsincludekind:"hidden"for skills used internally (e.g. by other tools) but not surfaced in the sidepanel? Open. v0.1 ships without; revisit when an internal skill emerges.
Roll-out
- Day 1 (today): this RFC + RFC-002 + RFC-003 land. SUMMARY.md updated.
- Day 2: bridge envelope +
chrome_*MCP tools + clipper handlers. - Day 3: site adapters land.
amem_capture_urlships. - Day 4: skill catalog + sidepanel UI (RFC-002 implementation).
- Day 5: recording skeleton lands (RFC-003 implementation).
- Day 6: end-to-end CWS demo recorded by amem itself, dogfooded.
- Day 7: CWS submission. Listing copy claims “first agent-callable Chrome skills catalog” honestly — three skills shipped, custom-skills disclosure visible.
- Post-v0.1: track the rule-of-three triggers in §6. When any fires, open RFC-00X for the skill engine.