Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

amem.sh

Shared knowledge for agents and humans.

amem is a local-first knowledge engine. It captures references (arXiv papers, PDFs, web pages), compiles them into structured wiki notes with SHA-256 provenance, and serves them to AI agents over MCP. Everything runs on your Mac — no cloud.

Why amem

Frontier models write better papers when they can cite real sources. Humans maintain better notes when capture is friction-free. amem serves both audiences from one knowledge base:

  • For agents: MCP tools amem_capture, amem_compile, amem_cite, amem_recall — grounded citations with verifiable provenance
  • For humans: a CLI and a Chrome extension that drop into your existing workflow, storing everything as readable markdown under ~/.amem/

Relationship to aide.sh

aide orchestrates agents; amem gives them memory. The two are complementary:

LayerToolConcern
Orchestrationaide.shDispatch, budgets, teams
Knowledgeamem.shCapture, compile, recall

aide can use amem as its memory sync method ([sync.memory] method = "amem" in aide.toml).

Installation

Prerequisites

  • macOS (primary target) or Linux
  • Rust toolchain (cargo)
  • Ollama running locally — required for the compile step
  • pdftotext (from poppler) — brew install poppler

Install the CLI

The CLI is in private beta and is not publicly distributed yet. If you have access to the source:

cargo install --path .

The Chrome extension works without it — captures stop at page title and URL, which is the part that needs no daemon.

Verify

amem --version
amem --help

Register with Claude Code

cargo install does not auto-register the MCP server. Run this once:

amem mcp install

This registers amem under your user scope (~/.claude.json). After restarting Claude Code, /mcp will show amem · ✔ connected, giving any agent access to amem_capture, amem_compile, amem_cite, amem_recall.

The one-liner shell installer (install.sh) runs this automatically if claude is on PATH.

To undo: amem mcp uninstall.

Pull an Ollama model

ollama pull llama3.1

Override the default model with AMEM_OLLAMA_MODEL=<name> if needed.

Next

Quick Start

Capture a paper, compile it into a wiki note, and recall it.

1. Capture

amem capture https://arxiv.org/abs/1706.03762

Downloads the PDF to ~/.amem/raw/ and prints a cite key (e.g., vaswani2017attention).

2. Compile

amem compile vaswani2017attention

Parses the PDF, chunks it with SHA-256 provenance, runs Ollama paraphrase passes, and writes ~/.amem/wiki/{ts}_vaswani2017attention.md.

3. Recall

amem recall "attention mechanism"

Grep-searches your wiki and returns excerpts with cite keys.

4. Cite

amem cite vaswani2017attention --format bibtex

Prints a formatted citation.

5. Hook it up to Claude

claude mcp add amem -- amem mcp serve

Then ask Claude: “Cite vaswani2017attention in APA format using the amem MCP server.”

Next

Concepts

Knowledge, not chat logs

Most “second brain” tools store a stream of captures — articles, notes, clippings — and hope you can find them later. amem does something different: it compiles captures into wiki notes with verbatim quotes and verifiable provenance. The wiki is the product; the raw captures are just source material.

This follows Andrej Karpathy’s recommendation of maintaining a personal wiki as the substrate for long-term thinking.

Dual audience

amem serves both agents (over MCP) and humans (via CLI + extension) from the same store. An agent citing a paper sees the same markdown a human sees. There’s no hidden “agent memory” that drifts from what’s on disk.

Provenance by construction

Every chunk carries a SHA-256 hash of its source. If the original file changes, amem verify detects drift. Citations stay grounded.

Offline-first

Zero cloud dependencies by default. Ollama runs locally. The only network calls are to fetch papers from their public URLs (arXiv, DOI resolvers). You can operate amem entirely air-gapped after initial capture.

Three interfaces

  • CLI — amem capture, amem compile, amem recall, amem cite
  • MCP server — amem_capture, amem_compile, amem_cite, amem_recall for agents
  • Chrome extension — one-click web page capture + self-recorded demos

MCP Tools

amem exposes four MCP tools over stdio. Start the server with amem mcp serve.

amem_capture

Download a paper and generate a cite key.

Input: url (string) — arXiv URL, DOI, PDF URL, or local file path

Output: cite_key (string), raw_path (string)

amem_compile

Parse + chunk + paraphrase a captured source into a wiki note.

Input: cite_key (string)

Output: wiki_path (string), chunk_count (number)

amem_cite

Format a citation in a supported style.

Input: cite_key (string), format (string, one of bibtex / apa / mla / chicago / ieee)

Output: citation (string)

amem_recall

Search the wiki for matching chunks.

Input: query (string), limit (number, optional, default 10)

Output: list of {cite_key, excerpt, sha256, score}

Register with Claude

claude mcp add amem -- amem mcp serve

Register with other MCP clients

Any MCP client that supports stdio transport works. Point it at the amem mcp serve command.

Storage Layout

Everything lives under ~/.amem/.

~/.amem/
├── raw/                          # original captures
│   ├── vaswani2017attention.pdf
│   └── ...
├── wiki/                         # compiled notes
│   ├── 20260418_vaswani2017attention.md
│   └── ...
└── index.md                      # auto-maintained TOC

raw/

Original files exactly as downloaded. amem verify re-hashes against these.

wiki/

Compiled markdown notes. One file per cite key. Filename prefix is the compile timestamp so recompiles preserve history (you can git init ~/.amem/wiki && git commit if you want versioning).

index.md

Auto-regenerated on every amem compile — a flat list of all cite keys with paths. Never edit by hand.

Custom location

Override with AMEM_HOME=<path> (planned — not yet wired up).

Citation Formats

amem cite <cite_key> --format <fmt> supports:

FormatFlagUse
BibTeXbibtexLaTeX papers
APAapaSocial sciences
MLAmlaHumanities
ChicagochicagoHistory, arts
IEEEieeeEngineering

Default: bibtex.

Example

$ amem cite vaswani2017attention --format apa
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N.,
Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need.

Metadata sources

  • arXiv API (for arXiv papers)
  • CrossRef (for DOIs)
  • Extracted from PDF metadata as fallback

amem Clipper (Chrome Extension)

amem Clipper is the browser-side companion to the native amem CLI — a Chrome MV3 side-panel extension for capturing web pages into your amem knowledge base. Positioning and install flow are defined by RFC-001.

Status

Bridge mode (Day 2 — in progress): extension talks to amem locally over WebSocket on port 7600. Extension captures → bridge forwards → amem stores.

Standalone mode (Day 3 — planned): when amem isn’t running, the extension falls back to Google Drive backup. Requires a new OAuth client (not shipped yet).

Install

The extension will be published on the Chrome Web Store as a new listing under the same developer account as crossmem (the v0 prototype). crossmem remains published and unchanged.

Published URL will appear here after approval.

Self-recording workflow

The extension ships with a self-recording skeleton: chrome.runtime.sendMessage({cmd: 'start_recording'}) triggers a tabCapture session via the offscreen document, producing amem-recording-<iso>.webm in Downloads.

This is how amem produces its own demo videos — the extension records itself being used. See Self-Recording Workflow.

Self-Recording Workflow

One of amem’s foundational features: amem is its own best demo. The extension records itself being used, producing the marketing video and Chrome Web Store listing assets automatically.

How it works

  1. An orchestrator (human or agent) sends {cmd: 'start_recording'} to the extension’s background service worker.
  2. The background worker opens an offscreen.html document with a MediaRecorder consuming tabCapture.
  3. The agent then drives the extension UI through the bridge — clicking the side panel, capturing pages, showing the wiki.
  4. {cmd: 'stop_recording'} finalizes the .webm and saves to Downloads as amem-recording-<iso>.webm.

Why this matters

  • Demos stay in sync with reality — the recording shows the current UI, not a stale screenshot
  • Chrome Web Store listings can be refreshed in minutes
  • Agents learn to operate amem by watching their own recordings

Status

The recording skeleton is shipped. The agent-driven recording scenarios are Day 2 work. The complete workflow ships alongside the Chrome Web Store submission.

amem capture

Download a paper or PDF into ~/.amem/raw/ and generate a cite key.

Usage

amem capture <url> [--doi <doi>] [--cite-key <key>]

Arguments

  • <url> — arXiv URL, DOI, PDF URL, or local file path
  • --doi <doi> — override DOI lookup (useful when a PDF URL doesn’t resolve)
  • --cite-key <key> — override the auto-generated cite key

Examples

amem capture https://arxiv.org/abs/1706.03762
amem capture https://example.com/paper.pdf
amem capture 10.1038/nature14539 --doi 10.1038/nature14539
amem capture ./local-paper.pdf --cite-key smith2024local

Output

Prints the cite key on success. Exits non-zero on failure.

See also

amem compile

Parse a captured source, chunk with SHA-256 provenance, run Ollama paraphrase, and emit a wiki note.

Usage

amem compile <cite_key>

Environment

  • AMEM_OLLAMA_MODEL — override the paraphrase model (default: llama3.1)

Output

Writes ~/.amem/wiki/{timestamp}_{cite_key}.md and updates ~/.amem/index.md.

Example

$ amem compile vaswani2017attention
Chunked into 47 chunks. Ollama paraphrase 47/47 ok.
Wrote ~/.amem/wiki/20260418_vaswani2017attention.md

See also

amem recall

Grep-search compiled wiki notes and return excerpts.

Usage

amem recall <query> [--limit N]

Arguments

  • <query> — free-text search term
  • --limit N — max results (default 10)

Example

$ amem recall "attention mechanism" --limit 3
vaswani2017attention (sha256:a1b2c3…): "…the Transformer, based solely on attention mechanisms…"
bahdanau2014neural  (sha256:d4e5f6…): "…a neural machine translation model learns to pay attention to…"

See also

  • MCP Tools — same backend exposed to agents

amem cite

Print a formatted citation for a captured source.

Usage

amem cite <cite_key> [--format bibtex|apa|mla|chicago|ieee]

Default format: bibtex.

See also

amem mcp

Subcommands for the MCP stdio server and its Claude Code registration.

amem mcp serve

Start the MCP stdio server. Exposes amem_capture, amem_compile, amem_cite, amem_recall to any MCP client.

amem mcp serve

amem mcp install

Register the current amem binary with Claude Code at user scope. Idempotent — safe to re-run; it removes any stale registration first. Requires claude on PATH.

amem mcp install

install.sh calls this automatically if Claude Code is installed.

amem mcp uninstall

Remove the registration from Claude Code.

amem mcp uninstall

See also

Philosophy

Wiki as substrate

Karpathy argues that a personal wiki — verbatim quotes, paraphrase, cross-links — is the durable substrate of long-term thinking. Apps come and go; markdown with provenance stays readable for decades.

amem treats the wiki as the product, not a byproduct. Capture is just staging; compile is the commitment.

Offline-first, always

Cloud dependencies are a liability. amem runs entirely on a user’s Mac: Ollama for LLM, pdftotext for PDF parsing, local disk for storage. The only network calls fetch papers from their canonical URLs.

This matters for three reasons:

  1. Sovereignty — your knowledge base doesn’t evaporate when a vendor pivots
  2. Speed — local I/O beats round-trips
  3. Reproducibility — SHA-256 provenance only means something if the source sits next to the hash

Dual audience

MCP for agents, CLI + extension for humans — same store, one source of truth. If an agent cites a chunk, a human can read that exact chunk by grep-ing ~/.amem/wiki/.

Complement, don’t compete

aide orchestrates agents. amem gives them memory. Neither duplicates the other. [sync.memory] method = "amem" in an aide.toml is the contract.

No legacy

amem is a clean v1 rewrite of crossmem. No code was ported verbatim; every module was reconsidered against these principles. crossmem remains published as v0 for historical continuity.

MCP Protocol

amem implements the Model Context Protocol over stdio using the rmcp Rust crate.

Transport

  • Stdio only (Day 1)
  • HTTP/WebSocket bridge for the Chrome extension is Day 2 work

Server identification

  • server_name: amem
  • server_version: matches the crate version

Tools

See MCP Tools for input/output schemas.

Error conventions

  • Malformed inputs → InvalidParams
  • Missing cite_key → ResourceNotFound
  • Ollama unreachable → Internal with hint start ollama
  • File IO errors → Internal with the OS error

Session model

Each amem mcp serve process is one session. No state shared between sessions beyond what’s on disk under ~/.amem/.

Relationship to aide.sh

amem and aide are complementary layers of the same system.

aide — orchestration

aide.sh dispatches work between agents. An Aidefile defines each agent’s persona, budget, triggers, and skills. aide dispatch <agent> <task> spawns an isolated Claude Code process and returns a bounded summary.

amem — knowledge

amem captures, compiles, and serves references. An agent running under aide can call amem_recall over MCP to ground its reasoning in real sources.

Integration point

In aide.toml:

[sync.memory]
method = "amem"
conflict = "causal"

When this is set, aide uses amem as its cross-agent memory substrate. Agents share a wiki; a fact learned by one agent is available to all.

Division of labor

Concernaideamem
Which agent runs?✓—
How much budget?✓—
What does the agent read?—✓
Where are citations stored?—✓
How are secrets gated?✓—
What’s the SHA-256 of chunk 7?—✓

Using one without the other

  • amem alone: CLI + extension + MCP. Useful for human-led research and single-agent setups.
  • aide alone: works fine, just without shared knowledge. Agents reason from their own seed context.
  • Both together: the full stack.

formace-00 — Zotero reference corpus (ops)

Set up 2026-08-16. Backend for issue #24 (Zotero data dir as the librarian’s reference corpus, fine-grained ref check).

Where things live

WhatPathDisk
Zotero 7 program (140.12.0esr)/mnt/storage1/users/ydwu/opt/zoterobig disk (19T)
Zotero data dir (zotero.sqlite, storage/*.pdf)/mnt/storage1/users/ydwu/zotero-databig disk
Zotero profile (prefs)~/.zotero/zotero/5sun4qjy.defaultsystem disk
systemd user unit~/.config/systemd/user/zotero.servicesystem disk
amem corpus (source of the import)~/.amem/raw (157 files, 78 with metadata)system disk

Nothing of size sits on / — it runs ~82% full (319G free of 1.8T).

Service

systemctl --user status zotero        # active/enabled
systemctl --user restart zotero
journalctl --user -u zotero -n 50

loginctl enable-linger ydwu is set, so the unit starts at boot without a login session. Zotero is a GUI app, so the unit runs it under xvfb-run.

Library contents

78 papers + 78 PDF attachments, imported from ~/.amem/raw/*.meta.json by generating BibTeX with file = {…} fields and POSTing to http://127.0.0.1:23119/connector/import. Cite keys match amem’s (belrose2023eliciting etc.), so Zotero items and amem wiki nodes line up.

Read the library:

curl -s "http://127.0.0.1:23119/api/users/0/items?limit=200"
curl -s "http://127.0.0.1:23119/api/users/0/collections"

The local read API is off by default; it is enabled via extensions.zotero.httpServer.localAPI.enabled in the profile’s prefs.js.

Security — this is a shared machine

formace-00 has 25+ accounts (a class of students). Three exposures were found and closed on setup; keep them closed:

  1. Zotero’s HTTP server on 127.0.0.1:23119 has no authentication. That is fine on a personal desktop and a data leak here — on a multi-user box localhost is not a security boundary, so any local account could read the whole library. The unit gates the port by uid via an iptables owner match, added in ExecStartPre and removed in ExecStopPost (so it needs no iptables-persistent and cannot outlive the service).
  2. zotero-data was created world-readable (drwxrwxr-x) → now drwx------.
  3. ~/.amem was also world-readable (drwxrwsr-x) → now drwx------. The 157 raw PDFs and every wiki node were readable by every account.

New directories under ~ inherit a group-writable umask on this host — check permissions after creating anything that holds corpus or credentials.

Gotchas hit during setup

  • pkill -f "opt/zotero/zotero" kills your own SSH command, because the pattern matches the command line running it. The script dies silently mid-way. Use pkill -x zotero-bin instead.
  • Zotero’s Linux download is served as .tar.bz2 but is actually XZ; tar -xf auto-detects, an explicit -j fails.
  • First launch must be given a display (xvfb-run) to create zotero.sqlite; it never exits on its own, so run it under timeout when you only want initialisation.
  • Do not read zotero.sqlite while Zotero holds it. Prefer a better-bibtex .bib auto-export (plain text, no lock) as the metadata source, or copy the sqlite first.

Chunked reference layer (built 2026-08-16)

src/refindex.rs in amem-librarian. Paragraph-level chunks, each hashed, with BM25 retrieval — no model, no GPU, so the pipeline is verifiable end to end before embeddings go on top.

export PATH="$HOME/.cargo/bin:$PATH"          # cargo is NOT on the default ssh PATH
AM=/mnt/storage1/users/ydwu/cargo-target/release/amem
$AM refindex build --bib /mnt/storage1/users/ydwu/zotero-data/amem-corpus.bib
$AM refcheck "<a sentence from your draft>" [--cite-key rafailov2023direct] [--limit 5]

State on formace-00: 78 refs → 12,751 chunks, 13 s, index at ~/.amem/refindex/chunks.jsonl.

What it is for: a hit carries cite_key#p<page>c<n>, the page, a SHA-256 and the query terms that actually matched, so “this paper does not support that sentence” becomes visible rather than a hunch. Measured on the Mac corpus, a DPO claim scores 32.2 against rafailov2023direct and 5.9 when misattributed to vaswani2017attention.

Content addressing immediately paid for itself — identical hashes across different keys exposed duplicate papers in the corpus:

identical chunkssame paper filed twice as
197wei2022chain = wei2022cot
156dpo_2023 = rafailov2023direct
97merrill2023expressive = merrill2024expressive

Worth deduping before the corpus grows; 12,751 chunks carry only 12,129 unique hashes.

Cross-machine ref check (built 2026-08-16)

amem refserve — a separate service from clipper_bridge on purpose. That one drives the user’s real logged-in Chrome, so exposing it off-host would hand a remote caller the browser. This one is read-only over the chunk corpus, so it is the only piece with any business on a non-loopback address.

serviceamem-refserve.service (user unit, linger already on)
bind100.86.146.58:7602 — the tailnet address, not 0.0.0.0, so the lab LAN cannot see it
authbearer token, constant-time compare; token in vault as AMEM_REF_TOKEN, on the box at ~/.config/amem/refserve.env (mode 600)
guardrailthe server refuses to start on a non-loopback bind with no token, or a token under 16 chars

From the Mac:

export AMEM_REF_TOKEN=$(bash ~/.claude/skills/vault/run.sh get AMEM_REF_TOKEN)
amem refcheck "<sentence from your draft>" --remote http://100.86.146.58:7602 \
    [--cite-key rafailov2023direct] [--limit 5]

Verified from the Mac: no token → 401, wrong token → 401, valid token → 12,751 chunks searched. --cite-key vaswani2017attention on a compositionality claim correctly returns NO SUPPORT, which is the wrong-citation signal working across the network.

GET /status is deliberately unauthenticated and deliberately boring ({ok, chunks, refs}) so a health probe reveals nothing about the corpus.

Bibliography chunks are excluded

Reference lists mention every topic in the field while asserting none of them, and they score well — a claim phrased like a paper’s title matches that title in every bibliography citing it. Observed live: querying the CoT paper’s own title returned three other papers’ reference lists above any real text. Chunks are now flagged by signal-counting (arxiv:/doi:/bare URLs weighted double, et al., In Proceedings, and 2022./2022b. year stamps) and skipped at search time. 16.5% of the formace corpus flags; after the fix the same query returns the actual paper’s title page first.

Dense retrieval (built 2026-08-16)

BM25 alone only finds passages sharing words with the claim. Measured: “transformers struggle to compose multiple reasoning steps” never surfaced dziri2023faith, the paper that argues exactly that, because it says “compositionality” and “error propagation”. With vectors, all three top hits are that paper — including its “Error Propagations: The Theoretical Limits” section.

ollama pull nomic-embed-text            # 768-d, ~274 MB
$AM refindex embed                      # 12,994 vectors in 4m11s on the RTX 6000

Vectors live at ~/.amem/refindex/vectors.bin (raw LE f32, fixed stride, ~38 MB) with vectors.meta.json carrying a fingerprint of the corpus they were built from. A mismatch is refused, not warned about — vectors out of step with the chunks would attribute quotes to the wrong papers. Re-run refindex embed after any refindex build.

Fusion is reciprocal-rank, not a weighted score sum: BM25 is unbounded and corpus-dependent while cosine is in [-1,1], so weighting needs calibration that drifts as the corpus grows.

Read the components, not the fused rank. RRF orders well but says nothing about whether a passage is support, so hits carry both:

bm25cosine
genuine support (dziri2023faith#p9c5)16.60.82
unrelated paper (liu2024spinquant#p5c4)1.80.52

Metadata verification — CrossRef and DataCite (built 2026-08-16)

amem refverify answers the question refcheck cannot: does this citation exist, and is the metadata right? (refcheck answers what it says.)

Which registry matters. arXiv registers preprints through DataCite (10.48550/arXiv.*); CrossRef 404s them. Verified by hand: CrossRef returns 404 for 10.48550/arXiv.1706.03762 while DataCite returns “Attention Is All You Need”, 2017, publisher arXiv. Routing everything to CrossRef made all 7 local preprints report ABSENT — indistinguishable from fabricated references. With prefix routing: 7/7, and 74/75 on formace.

An arXiv id is turned into an exact DOI, so most lookups are direct resolution rather than fuzzy search. Short titles additionally need the first author’s surname to corroborate: “Attention Is All You Need” is five tokens and matched a 2025 book chapter at 71% Jaccard, reported as a year error when it was the wrong paper entirely.

Result on the 78-ref corpus — every duplicate content hashing suspected, confirmed by DOI:

10.48550/arxiv.2201.11903  ←  wei2022chain = wei2022cot
10.48550/arxiv.2305.18290  ←  dpo_2023 = rafailov2023direct
10.48550/arxiv.2310.07923  ←  merrill2023expressive = merrill2024expressive

Note on the last one: both records carry year 2023 and DataCite agrees, so only the cite key string merrill2024expressive is wrong — the metadata never was.

Not done yet

  • better-bibtex plugin — not installed; stable cite keys currently come from amem’s own metadata, not from Zotero. The .bib currently indexed is generated from ~/.amem/raw/*.meta.json, not exported by Zotero.
  • Dedupe the three duplicate cite-key groups, now confirmed by DOI.
  • Re-embedding is manual — refindex build invalidates the vectors by fingerprint (correctly), but nothing re-runs refindex embed for you.
  • The clipper bridge is still loopback-only, deliberately — see the cross-machine section. Only refserve is exposed.
  • GUI curation — no VNC on the box. Adding/annotating items by hand needs either a VNC server or doing curation on the Mac and syncing.

OpenWiki / llm-wiki-agent interop research for amem-librarian

Research date: 2026-08-10. Read-only study. Both repos cloned to /private/tmp/claude-501/-Users-ydwu-claude-projects-amem-hq/09fb9b1e-7475-4775-aad3-0eeb124d78aa/scratchpad/.


0. Headline finding

OpenWiki does not have a format of its own. It emits Google’s Open Knowledge Format (OKF). The interop question is therefore not “should amem talk to OpenWiki” but “should amem emit OKF”. Those are different decisions with different risk profiles, and the second one is the one worth making.

Second finding: OpenWiki shipped a personal mode with Gmail / X / Slack / Notion / Hacker News / web-search connectors writing to ~/.openwiki/wiki. That is the same product shape as amem — sensors feeding a local agent wiki. OpenWiki is now a competitor, not just an interop target. Details in §6.


1. Repos found

OpenWikillm-wiki-agent
URLhttps://github.com/langchain-ai/openwikihttps://github.com/SamurAIGPT/llm-wiki-agent
AuthorBrace Sproul (Head of Applied AI, LangChain) + LangChainAnil Chandra Naidu Matcha (SamurAIGPT)
Versionnpm openwiki 0.3.1no release; pyproject says 0.1.0
LanguageTypeScript (Node ≥22), React/Ink CLIPython 3.10–3.13
LicenseMITMIT
First commit2026-06-222023-04-21 (repurposed later)
Last commit2026-08-072026-08-03
Commits total255102
Commits last 30d16411
Contributors60+; top 3 = Colin Francis 139, bracesproul 63, Brace Sproul 5110; top 3 = Anil 32(+9+3), watsonk1998 23, Tony Lin 17
Bus factor~2–3 (LangChain-backed, corporate)~2 (single maintainer + one heavy contributor)

The ai.miraheze.org background page’s claim of >10k stars is consistent with the repo’s activity, though star count is not verifiable from a clone.

Engine

OpenWiki is a DeepAgents documentation agent (deepagents + @langchain/*). Model providers are pluggable and numerous — OpenAI (default gpt-5.6-terra), Anthropic, Gemini/Vertex, Bedrock, Copilot, OpenRouter, Baseten, Fireworks, Nebius, NVIDIA, Ollama, LM Studio, any OpenAI-compatible gateway. Keys live in ~/.openwiki/.env. BYO key; no hosted LangChain service.

llm-wiki-agent has two execution paths: standalone Python scripts calling litellm (LLM_MODEL env, default claude-3-5-sonnet-latest), or — the headline mode — no API key at all, just Claude Code / Codex / Gemini CLI reading CLAUDE.md / AGENTS.md / GEMINI.md and driving the workflow with its own file tools.


2. Storage layout

amem (current, on disk)

~/.amem/
  raw/          -> symlink to /Volumes/YDExtSSD/amem/raw
  recordings/   -> symlink to external SSD
  wiki/         34 flat .md files, no subdirectories

Flat namespace, node id encoded in filename with a colon: url:d5ffad083d4a.md, arxiv:1706.03762.md, x:AnthropicAI.md. Namespace counts: url: 30, arxiv: 1, x: 1, plus 2 legacy files.

~/.amem/index.md does not exist, although SPEC.md §“Storage layout” lists it as “auto-maintained”. No log.md. No config.toml either.

amem has two coexisting node schemas in one directory, which matters for any mapping work:

Schema A — legacy PDF/arxiv compile pipeline (1776567380_vaswani2017attention.md):

cite_key: vaswani2017attention
title: "Attention Is All You Need"
authors: [ ... ]
year: 2017
arxiv_id: "1706.03762"
raw: "/Users/ydwu/.amem/raw/1776567380_vaswani2017attention.pdf"
pdf_sha256: "bdfaa68d…"
chunks: 15

Body: # Title → ## Citations (APA/MLA/Chicago/IEEE/BibTeX) → ## Chunks with per-chunk ### p1c1 (p.1) + `sha256:3b3e5b…`.

Schema B — clipper node (the format the task description gives, 32 of 34 files):

id: url:d5ffad083d4a
type: "page"
title: "ForMACE Lab"
url: https://…
host: "formace-lab.zulipchat.com"
first_seen_at: 2026-07-28T09:48:19.062Z
captured_at: [ …, … ]      # list, grows on re-capture
tldr: "…"
tags: [amem-clipper]

Body: # Title → > tldr → **Source**: <url> → ## Content (verbatim extracted text) → ## Links → ## Connections (empty in v0.2; the file literally carries <!-- amem-clipper v0.2 leaves this empty; v0.3 populates with [[<node_id>]] wikilinks. -->).

Only Schema A carries SHA-256. The clipper nodes — 32 of 34 — have no hash at all, so the “provenance / SHA-256 excerpts” differentiator currently exists only on the legacy arxiv/PDF path.

OpenWiki

Two modes, two roots (src/config/openwiki-home.ts, src/okf/index-sync.ts:66):

code mode:      <repo>/openwiki/          (virtual root /openwiki)
personal mode:  ~/.openwiki/wiki/         (virtual root /)

~/.openwiki/                    mode 0700, ACL-restricted on Windows
  wiki/                         the knowledge bundle
  connectors/<id>/
    config.json                 never contains raw secrets, only env var names
    state.json                  {version, lastRunAt, latestIds, runs[]}
    raw/                        raw dumps + manifests
    logs/
  conversation_history/
  skills/                       bundled SKILL.md dirs, synced on each run
  .env                          provider keys
  install-id                    telemetry install id

Nested directories are first-class: /sources/gmail.md, /topics/ai-research.md. Every directory gets its own generated index.md.

llm-wiki-agent

Repo-rooted, not $HOME-rooted (tools/_utils.py):

raw/          immutable source documents, never modified
wiki/
  index.md    catalog, updated every ingest
  log.md      append-only chronological record
  overview.md living synthesis across all sources
  sources/    one page per source document      (kebab-case.md)
  entities/   people, companies, projects       (TitleCase.md)
  concepts/   ideas, frameworks, methods        (TitleCase.md)
  syntheses/  saved query answers
graph/        graph.json, graph.html, .cache.json, .refresh_cache.json
tools/        ingest/query/heal/health/lint/build_graph/refresh/…

One file per role, not per source: a single ingest fans out into one source page plus N entity pages plus M concept pages.


3. File format

OKF — the actual standard

Spec: GoogleCloudPlatform/knowledge-catalog/okf/SPEC.md. The published spec is now v0.2. OpenWiki still emits and hardcodes v0.1 (src/okf/index-sync.ts renderIndex() writes okf_version: "0.1"; src/agent/prompts/personal.ts:97 instructs “MUST follow the Google Knowledge Catalog OKF v0.1 schema”).

OKF v0.2 frontmatter:

FieldStatusShape
typeREQUIREDshort string, uncontrolled vocabulary (“BigQuery Table”, “Playbook”, “Reference”)
titlerecommendedstring
descriptionrecommendedone-sentence summary, retrieval-optimized
resourcerecommendedcanonical URI of underlying asset
tagsrecommendedlist of strings
sourcesoptionallist of {resource(req), id, title, author, usage_count, last_modified}
usage_windowoptional{from, to}
generatedoptional{by: <actor>, at: <ISO8601>}
verifiedoptionallist of {by: <actor>, at: <ISO8601>}
statusoptionaldraft | stable | deprecated (default stable)
stale_afteroptionalYYYY-MM-DD
extensionsallowedany additional keys

Actor convention (§7): <producer>/<version> for tools, human:<id> for people, process:<id> for automation. Consumers classify trust by the human: prefix.

Structural rules:

  • Reserved files index.md and log.md are not concepts; everything else .md is a concept document.
  • okf_version may appear only in the bundle-root index.md — the sole place frontmatter is allowed in an index.
  • Cross-links are standard markdown links, not wikilinks: [label](/tables/customers.md) bundle-absolute, or [label](./other.md) relative.
  • Conformance (§11): parseable YAML + non-empty type + reserved-file rules. Consumers MUST NOT reject a bundle for unknown keys, unknown type values, broken links, or missing index.md.
  • Extensions (§10): “Consumers SHOULD preserve unknown keys when round-tripping and MUST NOT reject documents with unrecognized fields.”

v0.1 → v0.2 breaking changes (§13.1): timestamp → generated.at; body # Citations section → frontmatter sources.

OpenWiki’s validator (src/okf/frontmatter.ts) enforces exactly the v0.1 subset: type required; type|title|description|resource|timestamp must be non-empty strings when present; tags must be a list of non-empty strings. It knows two of its own extension fields, openwiki_generated and openwiki_translation_pending.

llm-wiki-agent

Schema lives in CLAUDE.md (also AGENTS.md, GEMINI.md) — the agent reads it as instructions. It is prose convention, not a validator; nothing in tools/ enforces frontmatter shape.

title: "Page Title"
type: source | entity | concept | synthesis   # closed vocabulary
tags: []
sources: []          # list of source slugs (NOT OKF's structured sources)
last_updated: YYYY-MM-DD

Source pages additionally use date: YYYY-MM-DD and source_file: raw/….

Body convention for a source page: ## Summary → ## Key Claims → ## Key Quotes → ## Connections → ## Contradictions. Domain templates exist for diary/journal and meeting notes.

Links are [[WikiLinks]] — tools/_utils.py:extract_wikilinks() is re.findall(r"\[\[([^\]]+)\]\]", content). Resolution is by filename stem, case-insensitive, ignoring directories (tools/ingest.py:validate_ingest, tools/lint.py:page_name_to_path). Obsidian-compatible by design, with a documented vault-symlink pattern in the README.

wiki/index.md is a hand-structured catalog with fixed sections (Overview / Sources / Entities / Concepts / Syntheses); new entries are string-spliced under the right heading by tools/ingest.py:update_index().

wiki/log.md: append-only, ## [YYYY-MM-DD] <operation> | <title>, deliberately grep-parseable (grep "^## \[" wiki/log.md | tail -10). Operations: ingest, query, health, lint, graph.


4. Update loop, dedupe, provenance

OpenWiki

  1. openwiki personal --init / --update, or openwiki ingest <connector|all>.
  2. Connector tools (deterministic, no model) fetch and write ~/.openwiki/connectors/<id>/raw/ plus a manifest, and update state.json with lastRunAt / latestIds / a runs[] summary. Incrementality is per-connector watermarks — latestIds, commit SHAs, Notion object ids + last-edited timestamps + content hashes.
  3. migrateWikiToOkf() runs before the agent, normalizing every concept page’s frontmatter so the agent sees a conformant wiki.
  4. The DeepAgents agent reads raw dumps and writes/edits wiki pages.
  5. synchronizeWikiIndexes() regenerates every directory’s index.md deterministically.
  6. validateWikiInternalLinks() stamps broken links inline.
  7. Mermaid fences validated; failures downgraded to text fences with a reason comment.
  8. Post-run snapshot; unchanged runs record no new metadata (“no-op runs are free”).

Dedupe is agent-mediated — the model decides whether to edit an existing concept page or create a new one. There is no content-hash dedupe at the wiki layer and no stable node-identity field.

Provenance: source URI in resource; connector raw dumps retained on disk; state.json run history. No hashes in frontmatter. No verbatim archive of the source in the wiki page — pages are agent-synthesized prose. OKF v0.2’s sources[] / generated / verified families are not emitted, because OpenWiki targets v0.1.

llm-wiki-agent

python tools/ingest.py <file> (or “ingest raw/x.md” to Claude Code):

  1. Auto-convert non-.md via markitdown (20+ formats incl. pdf/docx/pptx/ xlsx/epub/ipynb and wav/mp3 audio transcription).
  2. sha256(source_content, truncate=16) computed and printed.
  3. Build context = index.md + overview.md + 5 most-recently-modified source pages (that’s the contradiction-detection window).
  4. One LLM call returns a strict JSON envelope: title, slug, source_page, index_entry, overview_update, entity_pages[], concept_pages[], contradictions[], log_entry.
  5. Write pages, splice index, append log.
  6. Post-ingest validation: broken [[wikilinks]] + pages missing from index.md, both printed.

Dedupe/staleness: tools/refresh.py re-hashes each source_file and compares to graph/.refresh_cache.json; only changed sources are re-ingested (--force overrides). tools/build_graph.py caches by SHA-256 so only changed pages get re-processed by the semantic pass.

Provenance: source_file: points at the immutable raw/ document, date:, sources: [] slug list, plus the hash cache. Hashes are truncated to 16 chars and stored in a cache file, not in the page frontmatter — so a page alone is not verifiable.


5. Entry points and MCP posture

OpenWikillm-wiki-agentamem
Installnpm i -g openwikigit clone + pip install -r requirements.txtRust CLI + MCP server
Serves MCPNoNo (zero MCP references in the repo)Yes — amem_capture/compile/cite/recall/…
Consumes MCPYes — src/connectors/mcp-client.ts, full JSON-RPC over stdio + HTTP, openwiki_list_mcp_tools / openwiki_call_mcp_tool, backend: "mcp-http" | "mcp-stdio"Non/a
Agent integrationwrites AGENTS.md/CLAUDE.md blocks between <!-- OPENWIKI:START/END --> markersis driven by Claude Code/Codex/Gemini via CLAUDE.md + .claude/commands/*.md slash commandsMCP tools
API keyrequired (any of ~15 providers)optional — none needed in Claude Code moden/a
Skillsbundles skills/mermaid-diagrams, skills/write-connector; syncBundledSkills() installs into ~/.openwiki/skills.claude/commands/wiki-{ingest,query,lint,graph}.mdRFC-007 distillation (not yet shipped)

This asymmetry is the interop lever. OpenWiki is an MCP client with a generic mcp-stdio / mcp-http connector backend and an McpReadOnlyOperation config. amem is an MCP server. An OpenWiki user can therefore configure amem as a connector today, with zero code in either project — no format conversion involved. That is a distribution channel, not a schema problem.


6. Competitive read (unrequested but load-bearing)

OpenWiki personal mode overlaps amem’s core pitch: local markdown wiki, private (~/.openwiki, mode 0700), BYO key, sensors pulling from the user’s tools, agent-maintained, MIT. With LangChain’s distribution and 164 commits/month it will not stay behind.

Where amem still differs, on the evidence:

  • Real browser sensor. OpenWiki’s web reach is Tavily search and public APIs. It cannot see a logged-in page. amem Clipper runs in the user’s real Chrome session — Zulip DMs, X timeline, paywalled reading. url:d5ffad083d4a (a Zulip login-gated page) is exactly the capture OpenWiki structurally cannot make.
  • iOS share-sheet sensor. No equivalent.
  • Verbatim archive. amem keeps ## Content — the actual extracted text. OpenWiki keeps only agent-written synthesis in the wiki (raw dumps live separately under connectors/*/raw/ and are not the wiki page).
  • Hash-level provenance + citation formatting. OpenWiki has neither; llm-wiki-agent has truncated hashes in a side cache.
  • Fact-check. Neither project has any grounding/verification tool. amem_factcheck (RFC-003) has no counterpart in either codebase.
  • Skill distillation. OpenWiki ships skills to its own agent; neither project compiles user knowledge into reusable agent skills. RFC-007 is still unique.

Where amem is behind: index generation, link inference, staleness policy, audit log, schema self-healing, multi-format ingest, graph visualization. Section 9 lists what to take.


7. Field-by-field schema mapping

amem clipper node (Schema B) as the source of truth.

amem fieldOKF v0.2 targetllm-wiki-agent targetVerdict
id (url:d5ffad083d4a)none — keep as extension amem_id; optionally sources[].idfilename stemlossy. OKF has no document-identity field; identity is the path. Colons in filenames also need sanitizing.
type: "page"type (required)type: sourcelossless but degraded. amem’s type encodes media kind; OKF’s encodes concept kind; llm-wiki’s encodes role. "page" is legal OKF but carries almost no routing signal — remap to Web Page / Paper / Profile.
titletitletitlelossless
urlresourcesource_file (local path)lossless to OKF. lossy to llm-wiki, which expects a repo-relative raw path, not a URL.
hostnone — extension amem_hostnonemissing both; trivially re-derivable from resource.
first_seen_atnone — extension, or sources[].last_modified (date-only)date (date-only)lossy: both flatten ISO-8601 datetime to a date, and neither has “first seen” semantics.
captured_at[] (list)nonenonelossy — worst case. Re-capture history is a list; OKF’s nearest list-valued time field is verified[], which means something else. Encode as extension amem_captured_at.
tldrdescription## Summary body sectionlossless to OKF (rename). Structural move for llm-wiki.
tagstagstagslossless
## Content (verbatim)no standard sectionno standard sectionlossy. Both target formats expect synthesized prose. Survives only as an unrecognized body section — which conformant consumers must tolerate, but no consumer will understand.
## Links——lossy, low value
## Connections [[node_id]][label](/node_id.md) markdown links[[PageName]]lossless to llm-wiki (same syntax, resolved by stem). Requires rewrite for OKF — wikilinks are not OKF edges. Currently empty at v0.2, so this is a forward-looking cost, not a migration cost.

Legacy Schema A (arxiv/PDF nodes):

amem fieldOKF v0.2llm-wiki-agentVerdict
cite_keysources[].idsource sluglossless-ish
authors[]sources[].author (single actor)—lossy — OKF’s author is one actor per source entry, not a list.
year, arxiv_idextension fields—lossy
raw (local path)sources[].resource (bundle-relative or absolute)source_filelossless
pdf_sha256none — extensioncache onlymissing from both standards. amem-unique.
per-chunk sha256:nonenonemissing from both. amem-unique; the substrate amem_cite depends on.
## Citations (APA/BibTeX/…)v0.1 had # Citations; v0.2 retired it in favour of sources—churn risk: the one body section OKF standardized was removed in the very next version.
chunks: 15extension—lossy

Fields amem should adopt from OKF, not just map to:

OKF v0.2 fieldWhy amem wants it
verified: [{by, at}]The standard’s designated slot for verification events. amem_factcheck output belongs here: verified: [{by: "amem/factcheck@0.3", at: …}]. amem would be the first producer populating the trust family with real verification rather than self-attestation.
generated: {by, at}Distinguishes clipper-captured from agent-synthesized nodes; human: prefix convention gives free provenance classing.
status, stale_afteramem has captured_at[] but no staleness policy; stale_after is a re-capture trigger.
sources[]Structured multi-source provenance — a node captured from 3 URLs is currently inexpressible in amem’s single url field.

8. Interop options

Effort: S. Recommendation: reject.

Mechanically near-free, but it is not a read: OpenWiki mutates the directory it is pointed at.

  • migrateWikiToOkf() (src/okf/index-sync.ts) walks every .md and calls normalizeConceptContent(), which rewrites any page lacking a usable type — replacing its frontmatter with a minimal derived block and stamping openwiki_generated: true. amem’s Schema A nodes have no type field at all, so every legacy arxiv/PDF node would have its cite_key, authors, year, arxiv_id, raw, pdf_sha256, and chunks deleted on first run. Only openwiki_translation_pending is on the preservation list.
  • synchronizeWikiIndexes() writes an index.md into every directory.
  • validateWikiInternalLinks() rewrites files to stamp broken links.
  • The agent then edits page bodies as prose — destroying ## Content verbatim text, which invalidates every chunk sha256.

Also: personal mode’s wiki root is hardcoded to ~/.openwiki/wiki (openWikiLocalWikiDir), so “pointing” means symlinking that path at ~/.amem/wiki — full write exposure, no read-only mode.

llm-wiki-agent is gentler (its WIKI_DIR is repo-relative and its tools are opt-in), and its [[wikilink]] syntax already matches amem. But it expects wiki/{sources,entities,concepts,syntheses}/ subdirectories and a hand-structured index.md; a flat directory of url:*.md files yields “unindexed page” warnings for all 34 nodes and an empty graph.

Secondary blocker for both: colons in filenames. url:d5ffad083d4a.md is illegal on Windows, ambiguous in some markdown link resolvers, and gets percent-encoded to url%3Ad5ffad083d4a by OpenWiki’s index generator (encodeURIComponent).

Effort: M (~2–4 days for a conformant v0.2 bundle).

Target OKF v0.2 the specification, and note in the docs that OpenWiki currently reads v0.1. Write to a fresh directory (~/.amem/export/okf/), never in place.

Transform:

  1. tldr → description; url → resource; title, tags pass through.
  2. type: "page" → a real concept type derived from the id namespace (url:→Web Page, arxiv:→Paper, x:→Social Profile).
  3. Emit generated: {by: "amem-clipper/0.2", at: <latest captured_at>}.
  4. Emit sources: [{resource: <url>, id: <node_id>, last_modified: <first_seen_at date>}]; for Schema A add author and the raw path, and carry pdf_sha256 as an extension.
  5. Preserve everything unmappable under amem_* extension keys (amem_id, amem_host, amem_captured_at, amem_pdf_sha256, amem_chunks) — §10 guarantees consumers must not reject them.
  6. Sanitize filenames: url:abc.md → web-page/abc.md, i.e. use the namespace as a directory. Kills the colon and gives OKF the nested structure it expects.
  7. Rewrite ## Connections [[id]] → [title](/web-page/abc.md).
  8. Generate root index.md with okf_version: "0.2" and per-directory indexes — reuse the algorithm in §9.1.
  9. Emit log.md from captured_at history.

What breaks / what you accept: the export is a projection, not the node. Chunk-level sha256 blocks and ## Content have no OKF home and survive only as extension body — legal, but no OKF consumer will act on them. Re-importing an OKF bundle edited elsewhere is explicitly out of scope (see (c)).

Cost is bounded because it is pure output — nothing in amem’s write path changes, nothing in amem’s schema is held hostage to OKF’s version churn, and if OKF v0.3 breaks again you edit one exporter.

(c) Bidirectional sync

Effort: L. Recommendation: reject, and say so in the docs.

Five independent reasons, any one sufficient:

  1. No stable identity in OKF. Identity is the file path. amem’s id is the join key and has no standard home, so a page renamed by the other tool is unmatchable on the way back.
  2. The other writer is an LLM. OpenWiki’s agent rewrites body prose non-deterministically. Round-tripping mutates ## Content, which invalidates every chunk sha256 — the exact substrate amem_cite and amem_factcheck stand on. Sync would corrupt the differentiator.
  3. Lossy in both directions. §7 shows captured_at[], chunk hashes, and verbatim content have no OKF representation; a round trip cannot restore what the projection dropped.
  4. No merge model. Two writers, no vector clocks, no CRDT, no conflict surface. Last-writer-wins over a knowledge base is silent data loss.
  5. Version skew is already real. OpenWiki writes v0.1 while the spec is v0.2, so a bidirectional bridge must translate between two versions of a moving format on every hop.

(d) Ship amem as an OpenWiki MCP connector — the cheap win

Effort: S (documentation only).

OpenWiki already speaks MCP as a client (src/connectors/mcp-client.ts, backend: "mcp-stdio" | "mcp-http", readOnlyOperations). amem already serves MCP. So an ~/.openwiki/connectors/amem/config.json pointing at amem mcp serve with readOnlyOperations: [{type: "tool", name: "amem_recall"}] makes amem a source OpenWiki reads from — no format conversion, no exporter, no schema commitment.

Strategically this is the right shape: amem is the sensor-and-provenance layer; OpenWiki becomes one more consumer of amem_recall. It also inverts the competitive dynamic — instead of amem exporting into LangChain’s format, LangChain’s tool queries amem’s API.


9. Ideas worth stealing

9.1 Deterministic per-directory index generation — openwiki/src/okf/index-sync.ts:synchronizeWikiIndexes() / synchronizeDirectory() / renderIndex(). Zero LLM calls; reads title + description from each page’s frontmatter; sorts by href; writes only when the rendered content differs from what’s on disk, so scheduled runs don’t churn git. amem’s SPEC promises ~/.amem/index.md and it does not exist — this is a ~100-line port and amem already has the two fields it needs (title, tldr).

9.2 Normalize-before-agent, never reject — openwiki/src/okf/frontmatter.ts:normalizeConceptContent() + migrateWikiToOkf(). A non-conformant page is repaired deterministically (derive title from first H1, else filename) and stamped openwiki_generated: true; the agent later sees that flag and enriches it. PRESERVED_EXTENSION_FIELDS carries control markers across the rebuild. Directly applicable to amem’s own Schema A/Schema B split and to the v0.2→v0.3 ## Connections migration: repair silently, mark for enrichment, never fail a run.

9.3 Broken-link stamping instead of failing — openwiki/src/agent/wiki-link-validator.ts:validateWikiInternalLinks(). Broken links get an inline <!-- openwiki: broken internal link … --> comment; existing stamps are cleared at the start of each pass so they never accumulate, and a later run finds the comment and repairs the href. A self-healing loop that degrades instead of erroring.

9.4 Degrade-and-repair for generated artifacts — same pattern for Mermaid (src/mermaid/validate.ts, README §Diagrams): an invalid diagram becomes a plain text fence with a reason comment rather than a broken render; the next update finds it and fixes it. Quality recovers over successive runs instead of requiring a correct one-shot.

9.5 The health/lint cost boundary — llm-wiki-agent/tools/health.py vs tools/lint.py, boundary documented as a table in CLAUDE.md. health is deterministic, zero LLM calls, free, run every session (empty/stub files, index sync, log coverage). lint is semantic, costs tokens, run every 10–15 ingests (orphans, broken links, contradictions, gaps). “Run health first — linting an empty file wastes tokens.” This is the correct shape for an amem doctor, and the explicit run-order rule is the valuable half.

9.6 Two-pass graph with typed edges — llm-wiki-agent/tools/build_graph.py. Pass 1 parses [[wikilinks]] → EXTRACTED edges (deterministic). Pass 2 asks the model for implicit relationships → INFERRED (with confidence) or AMBIGUOUS. Louvain community detection clusters topics; SHA-256 cache means only changed pages are re-inferred; output is a self-contained graph.html. amem’s ## Connections is empty at v0.2 — this is a ready-made v0.3 design, and the EXTRACTED/INFERRED distinction keeps model guesses auditable rather than laundering them into the same link namespace.

9.7 Contradiction flagging at ingest time — llm-wiki-agent/tools/ingest.py, the contradictions[] field of the JSON envelope, checked against index.md + overview.md + the 5 most recent source pages. Cheap precursor to amem_factcheck: catch conflicts when writing, not at query time. Their own README frames it as the RAG differentiator (“Contradictions surface at query time (maybe)” vs “Flagged at ingest time”).

9.8 Append-only, grep-parseable log — llm-wiki-agent/wiki/log.md, ## [YYYY-MM-DD] <operation> | <title>, designed so grep "^## \[" wiki/log.md | tail -10 is the read API. log.md is also an OKF reserved filename, so adopting it is free conformance. amem has no audit trail today.

9.9 Refresh-by-hash staleness — llm-wiki-agent/tools/refresh.py re-hashes each source_file, compares against graph/.refresh_cache.json, re-ingests only what changed. amem already stores captured_at[] but has no policy that consumes it; combine with OKF stale_after for a re-capture trigger.

9.10 Multi-format ingest via markitdown — llm-wiki-agent/tools/ingest.py:convert_to_md() gets pdf/docx/pptx/xlsx/ html/epub/ipynb and wav/mp3 transcription from one dependency. amem’s Pockist share-sheet would inherit a large format surface for very little code.

9.11 Secrets by reference, never by value — openwiki/src/connectors/: connector config.json stores env var names; values live only in ~/.openwiki/.env; ~/.openwiki is mode 0700 with Windows ACL restriction (src/platform/windows-acl.ts). Matches RFC-007’s “no secrets, use vault references” guardrail and shows the config-file shape.

9.12 .openwikiignore as a read boundary — gitignore syntax; when active, it filters filesystem discovery and restricts shell execute. The README is careful about what it does not promise (“does not guarantee a topic is never mentioned, since the agent may still infer an ignored area from other allowed evidence”). Good model for an .amemignore on the clipper, and good copy for honest scoping language.


10. Risks

Schema churn — high. OKF v0.1 → v0.2 shipped two breaking changes (timestamp→generated.at, # Citations→sources) inside roughly a month, and one of them retired the only body section the format had standardized. OpenWiki has not caught up: it validates and emits v0.1 today. Anything amem builds against “OpenWiki’s format” is targeting a stale snapshot of a moving spec. Mitigating factors: type is the sole required field, and §11 forbids consumers from rejecting unknown keys — so a minimal export (type + title + description + resource + tags) is very unlikely to break, while the rich provenance families are where churn will bite.

Velocity churn — high for OpenWiki. 164 commits in 30 days, 255 total since 2026-06-22, with active refactors landing in the exact modules that matter here (#611 reorganized the CLI, #513 restructured the repo into domain directories, #371 added the link validator). Any code-level coupling will rot fast. Format-level coupling will not.

Bus factor. OpenWiki ~2–3 with LangChain behind it — low abandonment risk, high direction risk (it is a company’s product and will follow the company’s roadmap). llm-wiki-agent ~2, single-maintainer, 11 commits in the last month and most of them cosmetic star-history chores — treat as a design reference, not a dependency.

License — no obstacle. Both MIT. amem’s repos are private/proprietary; MIT permits use and derivation with attribution, so reading their code for ideas is fine and vendoring a file is fine with the notice retained. Implementing OKF creates no license relationship at all: a data format is not copyrightable, and OKF is published by Google as an open spec. The clean rule: implement the format, do not vendor the code.

Coupling cost. Option (b) is one output-only module with no upstream dependency — its failure mode is “the export is stale”, which is cheap. Options (a) and (c) put a competitor’s LLM agent inside amem’s write path; their failure mode is silent corruption of the provenance data that amem_factcheck is supposed to stand on. That asymmetry is the whole decision.

Strategic risk. Emitting OKF is also a small act of standard adoption in LangChain-and-Google’s direction. It is worth it — the format is genuinely better specified than anything amem would invent, and verified/sources give the fact-check story a standard vocabulary — but amem’s internal schema should stay amem’s, with OKF as an export target only.


11. Recommendation

Do (b) + (d): a one-way amem export --okf targeting OKF v0.2, plus a documented recipe for registering amem as an OpenWiki MCP connector.

Rationale:

  • The real standard is OKF, not OpenWiki. Build against the spec; treat OpenWiki as one consumer that happens to lag at v0.1.
  • One-way export is bounded, reversible, and touches nothing in amem’s write path. If OKF v0.3 breaks, one module changes.
  • OKF’s verified: [{by, at}] and sources[] give amem_factcheck a standard vocabulary to publish into. amem would be the first producer filling the trust family with actual verification rather than self-attestation — that is a positioning asset, not just plumbing.
  • The MCP connector path costs a documentation page and inverts the dependency: LangChain’s tool queries amem, rather than amem exporting into LangChain’s world.

Order of work: steal 9.1 (index generation) and 9.5 (health/lint split) first — both are pure wins independent of any interop decision, and 9.1 closes a gap between SPEC.md and reality. Then the exporter. Then the MCP connector doc.

What amem should NOT do:

  1. Do not let OpenWiki write to ~/.amem/wiki/. migrateWikiToOkf() strips frontmatter from any page without a type — that is every legacy arxiv/PDF node, including pdf_sha256 and chunks.
  2. Do not build bidirectional sync. §8(c).
  3. Do not adopt OKF as amem’s internal schema. It has no field for captured_at[], no field for chunk hashes, and no document identity — the three things amem’s provenance model is built on. Export to it; don’t live in it.
  4. Do not target OpenWiki’s v0.1 dialect. Emit v0.2 and let OpenWiki catch up; v0.1 consumers tolerate the extra keys by §11 conformance anyway.
  5. Do not vendor their code. Reimplement the patterns; the value is in the design, and both codebases are moving too fast to track.
  6. Do not chase connector parity (Gmail/Slack/Notion/X). That is LangChain’s strength and a treadmill. amem’s edge is the logged-in browser session and the iOS share sheet — sensors OpenWiki structurally cannot build — plus fact-check and skill distillation, which neither project has attempted.

amem 競品掃描 — 最接近的對手是誰

研究日期:2026-08-11。唯讀研究,未修改任何 project repo。 方法:先讀 amem 自家 SPEC/RFC/design 與 ~/.amem/wiki/ 實際產出,再做廣掃,再對前三名做原始碼層級驗證。 llm_wiki 與 llm_wiki_skill 已 clone 至本 scratchpad 逐檔查核。


0. 結論先講

最接近的競品是 LLM Wiki(nashsu/llm_wiki)。

它不是「同類產品」,而是同一張架構圖的另一個實作:Chrome MV3 clipper 抓當前分頁 → 本機 Rust daemon 用 LLM 兩段式摘要 → 寫成 Obsidian 相容的 markdown wiki → 內建 MCP server 讓 frontier model 取用並附引用。amem 的 SPEC「Mental model」四層,它四層都有。

更關鍵的是兩者同源。amem SPEC 的 storage layout 寫「wiki/ # compiled wiki notes (Karpathy style, agent-friendly chunks)」;llm_wiki 的 README 首段就寫「This project is based on Karpathy’s LLM Wiki pattern」。兩個產品在讀同一份 gist。這不是巧合式撞車,是同一個公開設計模式的兩個實作,而對方已經做到 16,173 stars。

前一份 OpenWiki 研究漏掉它,是因為那份研究從「interop 對象」出發,掃的是 LangChain 生態。llm_wiki 不在那條線上,它在 Karpathy gist 那條線上——也就是 amem 自己站的那條線。


1. 為什麼是它:逐項對照 amem 的 Mental model

以下每一列都經我親自 clone 原始碼查核,非讀行銷文案。

amem SPEC 的層amem 現況llm_wiki 現況證據
Clipper(Chrome MV3 web sensor)有有,MV3,activeTab+scripting+Readability.js+Turndown.js,Alt+Shift+Lextension/manifest.json、extension/clipper-core.js
本機 daemon(Rust)amem-librarian(Rust)有,Tauri v2 Rust backend,本機 HTTP API 127.0.0.1:19827(clip)與 :19828(API)src-tauri/、mcp-server/README.md
AI 摘要編譯成 wiki有有,two-step chain-of-thought ingest(先分析再生成)README「3. Two-Step Chain-of-Thought Ingest」
Obsidian 相容 markdown有(扁平 ~/.amem/wiki/)有,而且更完整——自動產生 .obsidian/ 設定README:79「the wiki directory works as an Obsidian vault」、README:370
MCP 供 frontier model 取用有(amem_recall/amem_cite/amem_ground…)有,bundled MCP server,10 個 toolmcp-server/src/index.ts
附引用的 grounded recallamem_ground 回傳 hits + inline_md有,[1] [2] 編號引用 + Cited references panel + 引用持久化README:229、README:242-243
來源可追溯frontmatter url有,每頁 frontmatter sources: [] 指回 raw 檔README:132
Raw 原件保留~/.amem/raw/有,raw/sources/ 明訂 immutableREADME 目錄結構
index.md / log.mdSPEC 承諾但檔案不存在有,wiki/index.md + wiki/log.md 皆已實作README 目錄結構
蒸餾成 Claude skill(RFC-007)未實作部分——見 §3,這格要講精確llm_wiki_skill/SKILL.md
iOS sensoramem-pockist(TestFlight)無全 repo 無 mobile
fact-checkRFC-003 規劃中無全 repo 無 claim verification

成熟度:16,173 stars、1,918 forks,建立於 2026-04-08(僅 4 個月),最新 v0.6.8(2026-08-08),最後 push 2026-08-10,GPL-3.0(LICENSE 檔為 GPL v3 全文,作者 Yong Su;GitHub API 因檔頭多一行版權宣告而回 NOASSERTION)。桌面三平台 macOS / Windows / Linux。作者 nash_su。

四個月 16k stars,代表這個方向的市場需求已被驗證,也代表時間視窗正在關閉。


2. Top 6 競品對照表

軸線取「真正決定重疊度」的九項。「登入頁」欄位特別標註推論與明文的差別。

LLM WikiSiYuanObsidian Clipper(+社群 MCP)Karakeepbasic-memorySurfSense
Capture 面Chrome MV3,抓當前分頁 DOM官方 Chrome/Edge extensionChrome/Firefox/Safari(含 iOS/iPadOS)Chrome/Firefox/Safari + iOS/Android app無Chrome extension
登入頁結構上可(讀 live DOM),官方未明文同上,未明文同上,未明文同上,未明文—舊 README 宣稱可抓 authenticated,現行 docs 已無此頁(404),無法驗證
抓自己的 AI 對話無專用支援無無官方 template無—無
Mobile share sheet無iOS/Android/HarmonyOS appSafari extension,非 share sheet有無無
儲存本機 markdown 檔.sy JSON(非 plain md)本機 markdown 檔SQLite/Postgres + assets本機 markdown 檔server 端 DB
Obsidian 相容是,自動產 .obsidian/否(僅單向匯入/匯出)原生否是否
誰消費agent(MCP)+ 桌面 UIagent(MCP)+ UI人(vault)/agent 靠第三方 MCPagent(MCP)+ UIagent 為主agent(MCP)+ web UI
MCPserve(10 tools,內建)serve + consume(官方,30+ tools)官方無;社群 mcp-obsidian 4,287★serve(官方 @karakeep/mcp)serve(原生,20+ tools)serve
Provenancesources: [] + raw immutable + [1] 引用clipper 寫入原始 URL + 時間戳frontmatter 有 source URL/author/datemonolith 全頁存檔(verbatim 最強)無 source 追蹤引用式回答
內容 hash僅圖片去重用 sha2,非 provenance無無無無無
Fact-check無無(web_search 是查資料非驗證)無無無無
Skill 迴路消費 skill(/skill)+ 官方 access skill消費 skill(SKILL.md 機制)無無無無
LicenseGPL-3.0AGPL-3.0MITAGPL-3.0AGPL-3.0未在 README 明示
平台mac/Win/Linux 桌面桌面+行動+Docker瀏覽器自架+行動CLI/MCPDocker 自架+雲
商業模式免費 OSSFree / $64 買斷 / $148 年App 個人免費;Sync $4/mo免費自架本機免費;雲端 $15/mo自架免費;雲 PAYG
成熟度16,173★,4 個月,v0.6.8 (2026-08-08)45,733★,v3.7.3 (2026-07-21)4,993★,1.7.1 (2026-07-22)28,249★,v0.33.1 (2026-08-01)3,627★,v0.22.1 (2026-06-13)15,876★,v0.0.36 (2026-08-06),自陳未達 production

star / release 數據以 GitHub REST API 於 2026-08-11 查核。

表格讀出來的三件事:

  1. MCP 不再是差異點。 六家有五家 serve MCP,SiYuan 甚至雙向。amem 把「MCP 存取層」當賣點已經失效。
  2. 「本機 markdown + agent 可讀」也不再是差異點。 llm_wiki、basic-memory、Obsidian 生態都做到了。更要命的是 Claude Code 自己的 auto memory 就寫在 ~/.claude/projects/<project>/memory/ 的明文 markdown,是免費預設值。
  3. 內容 hash 與 fact-check 兩格全業界皆空。 這是 amem 唯一真正無人佔領的象限——但見 §4,amem 自己也還沒真的佔住。

3. 對 LLM Wiki 的攻防拆解

3.1 幾乎完全重疊的部分

  • capture 機制一模一樣。 兩者都是 MV3 extension 在使用者真實 session 讀 live DOM,經 Readability 抽取後 POST 到 localhost 的本機 daemon。amem 用 :7601,llm_wiki 用 :19827。連「登入頁能不能抓」的答案都一樣:結構上可以,因為 content script 跑在使用者已登入的分頁裡。
  • compile 目標一模一樣。 LLM 摘要 → Obsidian 相容 markdown wiki → wikilink 知識圖譜。
  • agent 取用方式一模一樣。 本機 HTTP API 包一層 MCP server 給 Claude Code / Codex。
  • 知識來源三層架構一模一樣。 raw immutable → wiki → schema,因為都照 Karpathy gist。

3.2 amem 真正領先的地方

(a)iOS sensor。 llm_wiki 完全沒有行動端。amem-pockist 已上 TestFlight。share sheet 是行動端唯一低摩擦的捕捉入口,這格對方短期補不上(要做原生 app + 一套同步)。

(b)登入頁與自有 AI 對話是「明講的產品主張」而非副作用。 兩邊技術上都能抓登入頁,但 llm_wiki 從未把它寫成賣點,也沒有 per-site recipe。amem 的 design memo(2026-07-31)已經把 per-site recipe 定為要建的東西,且 ~/.amem/wiki/ 裡真的躺著 Gmail 搜尋結果頁與 Zulip 登入牆後的 capture。把「別人抓不到的頁面」做成明確能力,是 llm_wiki 沒佔的位置。

(c)chunk 級 SHA-256 + 格式化 citation。 llm_wiki 的 sha2 依 src-tauri/Cargo.toml 註解明寫是「for the dedup cache (Phase 3) — same image hash」,只用於圖片去重;agent workspace 追蹤用的是非密碼學的 DefaultHasher。amem 的 cite.rs 有 text_sha256 per chunk、pdf_sha256,並輸出 APA/MLA/Chicago/IEEE/BibTeX。全業界只有 amem 做這件事——但只在 PDF/arxiv 路徑上,見 §4。

(d)fact-check 是規劃中的產品核心。 llm_wiki 只有 Read Sources Only(限縮回答來源)與人工 Review 佇列,沒有對外部 trust list 驗證 claim 的機制。全掃描 18+ 個標的皆無。

(e)分發成熟度(潛在)。 llm_wiki 的 extension 沒有上 Chrome Web Store,README 只教 chrome://extensions → Developer mode → Load unpacked。對非開發者是硬門檻。但這一格 amem 目前也還沒兌現——amem-clipper(id jgknnaaaobkdggmjhhbklagidilmadii)CWS 查詢顯示 crx 0.3.0、無 published version、in review。要贏這格得先真的上架。

3.3 LLM Wiki 領先的地方

(a)規模與動能。 16,173★ / 1,918 forks / 4 個月 / 每兩天一個 build。amem 是單人私有 repo。這是最大的落差,且會自我強化。

(b)多格式 ingest 遠勝。 PDF(含 MinerU)、DOCX、PPTX、XLSX、EPUB/MOBI、影音、圖片 vision caption。amem 只有 web clip 與 PDF。

(c)wiki 維護機制成熟。 index.md / log.md / overview.md / Lint / Review 佇列 / cascade 刪除(含 dead wikilink 清理)/ 4-signal 知識圖譜 + Louvain 分群。amem 的 ## Connections 至今是空的(節點裡還留著 <!-- amem-clipper v0.2 leaves this empty --> 註解),SPEC 承諾的 index.md 也不存在。

(d)Deep Research 迴路。 偵測知識缺口 → 自動 web search(Tavily/SerpApi/SearXNG)→ 合成新 wiki 頁 → 自動 ingest。amem 沒有主動補洞能力。

(e)purpose.md。 讓使用者宣告「這個 wiki 為什麼存在」,LLM 每次 ingest/query 都讀。這是把使用者意圖變成可執行 context 的簡單好設計,amem 沒有對應物。

(f)scenario templates。 Research / Reading / Personal Growth / Business 各自預設 purpose.md 與 schema.md,降低冷啟動成本。

3.4 一個必須修正的 amem 內部說法

RFC-007 寫:「Every clipper competitor (Obsidian Web Clipper, Readwise, Notion) stops at storage; nobody closes the loop into agent capability. Step 3 is the moat.」

這句話現在有一半是錯的,需要改。

llm_wiki 已經把迴路收到 agent capability:MCP server + 官方發布的 llm_wiki_skill,npx skills add 一鍵裝進 Claude Code / Codex。所以「沒人閉環到 agent capability」不成立。

但精確地看,amem 的主張仍然活著,只是範圍小得多。 我逐字讀了 llm_wiki_skill/SKILL.md:它是人手寫的、唯讀的存取型 skill,內容是「怎麼呼叫 127.0.0.1:19828 的 API」,frontmatter 的 description 甚至明訂「DO NOT trigger on generic ‘search my notes’」。它讓 agent 讀 wiki,不是把 wiki 內容編譯成新 skill。llm_wiki 的「Agent Skills」功能同樣是掃描並啟用既有 SKILL.md——消費 skill,不生產 skill。

所以 RFC-007 該改成這樣:閉環到 agent 存取(MCP + access skill)已是紅海;把重複出現的程序性 capture 叢集自動編譯成新的 SKILL.md,目前仍無人做。moat 不是「step 3」整段,而是 step 3 裡的「自動生成」那半段。這個修正很重要,因為前者站不住,後者站得住。

3.5 amem 該怎麼做

定位語言

  • 停用「local-first markdown wiki for agents」這類描述。llm_wiki、basic-memory、Obsidian 生態、Claude Code 內建 memory 全都符合這句話。
  • 改押兩件別人結構上做不到或沒做的事:(1)別人抓不到的頁面(登入牆後、自己的 AI 對話、行動端 share sheet);(2)記憶可被稽核(chunk hash + verbatim + 格式化引用 + fact-check)。
  • 「sensor」這個字要繼續用,而且要講滿。競品的入口是「你手動匯入的文件」,amem 的入口是「你本來就在讀的東西」。這是敘事上的真差異。

table-stakes 但目前缺席,必須補

  1. ~/.amem/index.md 與 log.md。SPEC 已承諾、對手已實作、與任何 interop 決策無關。前一份 OpenWiki 研究的 §9.1 已經給了可直接移植的演算法。
  2. ## Connections 真的填上 wikilink。空的 Connections 讓 wiki 退化成一堆孤立檔案,知識圖譜是這個品類的基本盤。
  3. Lint / health 這類 wiki 維護迴路。
  4. 把 amem-clipper 真的推上 CWS。這是對 llm_wiki 少數幾個可立即兌現的優勢,卡在 review 就等於沒有。

不要做的事

  • 不要追多格式 ingest 與知識圖譜視覺化的 parity。那是 llm_wiki 有 60 個貢獻者級動能的地方,單人追不上,且不是 amem 贏的理由。
  • 不要把 MCP 當賣點寫在 landing page 第一屏。

4. amem 最沒防守的一塊

「verbatim archive + SHA-256 provenance」這個差異化主張,在實際出貨的路徑上不存在。

這是本次掃描對 amem 最不利、也最該立刻處理的發現。三個獨立證據:

  1. ~/.amem/wiki/ 34 個檔案中,32 個沒有任何內容 hash。 只有兩個 legacy PDF 節點(1776567380_vaswani2017attention.md、1776825486_belrose2023eliciting.md)帶 pdf_sha256 與 per-chunk sha256:。所有 url:* clipper 節點都沒有。
  2. clipper 路徑的 SHA-256 是拿來算 ID,不是算內容。 amem-librarian/src/clipper_bridge.rs 的 sha12() 把正規化後的 URL 做 SHA-256 再截前 12 個 hex,產出 url:d5ffad083d4a 這種節點 id。內容本身從未被 hash。
  3. amem_ground 不是 fact-check。 讀 src/mcp/mod.rs:172 的 tool description,它是「search the local wiki for the topic and return JSON hits」——對自有 wiki 的檢索加引用。SPEC 定義的核心差異化 amem_factcheck(recall + 撥出去打 trust list + 驗證)仍是 RFC-003 規劃中。

也就是說:對外講的兩個核心差異(可驗證的 provenance、fact-check),一個只存在於非主力的 PDF 路徑,一個還沒開始。而主力路徑正是 clipper——它產出 32/34 的節點,是整個產品的入口。

附帶的品質風險:clipper 節點的 ## Content 雖然是 verbatim,但實測是未經整理的 DOM 傾印。url:e55e42200ccb.md(Gmail)裡是大段錯亂的表格標記與 ![](images/cleardot.gif)。llm_wiki 用 Readability + Turndown 做過一輪清洗。「verbatim archive」若拿不出乾淨可引用的文字,對 fact-check 也撐不起來。

優先修的順序很清楚:先讓 clipper 節點帶 chunk 級 SHA-256 與乾淨的 verbatim 文字,再談 fact-check。 沒有前者,amem_factcheck 沒有可驗證的底層資料可站。這是整個 amem 論述的地基,而地基目前只鋪在 2/34 的面積上。


5. 亞軍與其他值得記錄的

SiYuan(45,733★, AGPL-3.0) — 功能面最完整的對手:官方 clipper(寫入原始 URL + 時間戳)、官方 MCP 雙向、AI Agent、SKILL.md 機制、三大行動平台官方 app。輸在資料主權敘事:.sy 是 JSON 不是 plain markdown,且非 Obsidian 相容,要匯出才拿得回來。amem 在「檔案就是你的」這點上贏它。

Obsidian Web Clipper(4,993★, MIT)+ 社群 mcp-obsidian(4,287★) — 「組裝出來的 amem」。官方 clipper 已有依網站自動套用 template 的機制,也就是 per-site recipe 概念已經存在於出貨產品裡。缺的是中間層:沒有常駐 daemon、沒有 AI 摘要編譯、沒有 provenance 層,且 capture 與 agent 存取由互不相關的專案提供。amem 的機會正是那個中間層。

Karakeep(28,249★, AGPL-3.0) — provenance 的另一種解法,且在「保真」這軸上其實比 amem 強:用 monolith 做全頁存檔對抗 link rot,加 yt-dlp 影片存檔。有官方 MCP、瀏覽器 extension、iOS/Android app。輸在儲存是 DB 不是 markdown,且無 Obsidian、無 agent 導向的 wiki 編譯。若 amem 要講「verbatim archive」,Karakeep 是必須被比較的對象。

basic-memory(3,627★, AGPL-3.0) — MCP 原生 + 本機 markdown + Obsidian 相容,agent-first 程度最高。完全沒有 capture 面,知識靠對話產生。它證明了「MCP + markdown」這半邊已被佔滿。

SurfSense(15,876★) — 曾以「儲存 authenticated 頁面」為 extension 主打,但現行 README 與 docs 已無該頁(/docs/browser-extension 回 404),無法從一手來源證實現況。官方自陳「not yet production-ready」。

Reor(8,573★,2026-03-07 已封存) — 值得引用的死亡案例。它有 local-first + 本機 LLM + 本機 vector DB 的完整組合,死在沒有 capture surface,匯入要人手動拷貝 markdown 進資料夾。這反向支持 amem 押 sensor 的判斷。

平台風險(非競品但更緊迫) — Claude Code auto memory 預設開啟,寫在 ~/.claude/projects/<project>/memory/ 的明文 markdown,使用者可直接編輯。「本機 markdown、agent 自寫、可稽核」已是 Anthropic 免費內建。它唯一缺的是 provenance——只記 modified 時間戳,不記來源。這再次指向同一個結論:amem 只能靠 provenance 與 sensor 活,不能靠儲存格式活。


6. 建議的定位句

amem 記錄你真正讀過的東西——包括登入牆後面那些——並且每一句話都能追回原文。

拆解:「你真正讀過的」= sensor(對手要你手動匯入);「登入牆後面」= 結構性優勢,且 llm_wiki 沒佔;「追回原文」= provenance 與 fact-check,全業界空白。整句刻意不提 markdown、不提 local-first、不提 MCP——這三個詞現在都無法把 amem 和任何人區分開。

但這句話目前只有前半是真的。 後半要成立,得先讓 clipper 節點帶上內容 hash 與乾淨的 verbatim 文字。在那之前,這是一句願景,不是一句產品描述。


7. 未能驗證的項目

  1. 登入頁擷取能力:llm_wiki、SiYuan、Obsidian Clipper、Karakeep 官方文件皆未明文陳述。本報告全部標為「結構上可推論」,因為 MV3 content script 在使用者已登入的分頁執行。
  2. SurfSense extension 現況:/docs/browser-extension 404,surfsense_browser_extension/README.md 只有 Plasmo 建置指令。「儲存 authenticated 頁面」的說法來自舊版 README 與 fork 快照,無法確認是否仍為現行能力。
  3. llm_wiki 的 GitHub license 欄位回 NOASSERTION,但 LICENSE 檔內容為完整 GPL v3 全文加一行作者版權宣告。以檔案內容為準。
  4. llm_wiki 是否計畫做 mobile 或 fact-check:repo 內無 roadmap 佐證,僅能陳述「目前沒有」。

amem Clipper + Librarian — 使用者痛點 vs 產品現況

研究日期:2026-08-17。唯讀研究,未修改任何 repo,未發任何 issue/PR。

方法:先親自讀 amem 原始碼與 ~/.amem/ 實際產出,再挖四類一手來源。 一手來源為 GitHub issue tracker(10 個 repo)、HN Algolia API、forum.obsidian.md Discourse API、Firefox AMO ratings API。前一份競品掃描(2026-08-11)的結論視為已知, 不重推。


0. 結論先講

amem 對「登入頁擷取」這個品類最大痛點答得很好,對「存下來的東西是否可信」答得最差。

第二點是致命的,因為它正是 amem 對外的定位句。SPEC 說 fact-check 是核心, 定位句說「每一句話都能追回原文」。實際出貨的 clipper 路徑做不到: 14 個節點在 6000 字截斷、沒有任何內容 hash、沒有 raw 備份、重剪直接覆蓋舊檔。

第三點是 RFC-007 的 moat 已經消失。不是變弱,是消失。Anthropic 官方文件明寫 產生 SKILL.md 不需要工具,Claude Code 二進位檔本身就會自動寫 SKILL.md。


1. amem 現況查核(我親自讀碼,非讀 SPEC)

這張表是後面所有判斷的地基。每一列都有檔名行號。

項目SPEC/定位的說法程式碼實際行為證據
內容 hash「每一句話都能追回原文」clipper 節點完全沒有 hash 欄位amem-librarian/src/clipper_bridge.rs:570-597 render_wiki_md 的 frontmatter 只有 id/type/title/url/host/first_seen_at/captured_at/tldr/tags
hash 涵蓋率—2/35 個 wiki 節點有 hash,都是 legacy PDF~/.amem/wiki/ 實測;30 個 url:* 節點與 1 個 clip:* 節點皆無
verbatim 全文「raw immutable → wiki」三層body 截斷在 6000 字,尾巴接 …(truncated)clipper_bridge.rs:588
截斷實際比例—14/30 個 clipper 節點已被截斷grep -l '…(truncated)' ~/.amem/wiki/*.md
raw 層~/.amem/raw/ 保留原件clipper 路徑不寫 raw。raw/ 只有 14 個 arxiv PDF 與 metabackground.js 只 POST /save_wiki_node,無 raw 寫入路徑
重剪語意去重、防 re-clip spam保留 first_seen_at 與 captured_at 清單,然後整檔覆蓋clipper_bridge.rs:359-397,std::fs::write 覆寫;captured_at 上限 20 筆
recall 檢索amem_recall / amem_groundtoken 計數 grep,非 BM25、非向量src/query/grep.rs:1 檔頭自述「MVP search… Token-count scoring」;src/mcp/mod.rs:167,175 走 helpers → grep
BM25 + 向量已建已建,但只服務 Zotero 參考文獻語料,不服務 wikisrc/refindex.rs(948 行)、src/embed.rs(519 行),入口是 refindex build / refcheck
Connectionsv0.3 填 wikilink31/35 個節點仍是字面 placeholder 註解<!-- amem-clipper v0.2 leaves this empty -->
index.md / log.mdSPEC 承諾不存在~/.amem/ 與 ~/.amem/wiki/ 均無
AI 對話自動存manifest 宣傳「Rich extraction on Claude / ChatGPT / Gemini」三個 autosave content script 都是 8 行 stub,只有 console.debugextension/content-scripts/{claude,chatgpt,gemini}-autosave.js 檔頭自述「Day 1-2: stub only」
AI 對話手動擷取同上真的可用。三個 extractor 各 200-230 行,含 artifact 與 binary 判別extension/extractors/claude.js 等;background.js:77-83 依 host 派送
離線容錯—無佇列、無重試。daemon 掛掉則該次擷取遺失background.js 唯一 alarm 是 amem-bridge-reconnect,只服務 bridge 輪詢
權限面—<all_urls> + 12 個 permission,含 tabCapture、downloads、scriptingextension/manifest.json
CWS 上架前次掃描判定「卡在 review」已公開上線。v0.3.0,更新 2026-08-13,3 users,0 則評論CWS 公開頁實測;chrome-extension skill 的 status 判讀為誤報

兩件要更正前次掃描的事。

第一,amem-clipper 已經上架了。前次掃描說「卡在 review 就等於沒有」, 那是 CWS API 的 uploadState: NOT_FOUND 造成的誤判。公開頁面查得到 v0.3.0。

第二,「verbatim archive」比前次掃描的描述更糟。前次說是「未整理的 DOM 傾印」, 問題是髒。實際上更嚴重:47% 的節點被截斷,而且沒有 raw 備份可以還原。 截斷後的 markdown 就是唯一一份副本。


2. Part A — 痛點排序

排序依據是「頻率 × amem 目前答得多差」。答得好的排在後面。

#痛點頻率amem 現況落差
1靜默的部分擷取與保真度崩壞7 repo,~90 comment,~70 reaction失敗最大
2存了找不回來4 repo + HN dominant失敗大
3權限過寬引發的供應鏈恐懼安裝決策點,有卸載實例失敗,且無謂大
4AI memory 錯誤、過期、無法驗證2026 最熱,HN 180 pts部分中
5離線時擷取不進去~92 reaction忽略中
6讀了不回頭的墳場HN 十年 dominant,tracker 僅 41 reaction忽略中(但被高估)
7登入頁與付費牆擷取全調查最高 reaction mass結構性解決無
8重複擷取6 repo,兩位維護者宣告做不到解決,但用資料遺失換反轉風險
9本機 LLM 連線摩擦5 repo,~80 comment部分小
10上架被下架的存亡風險有死亡案例已上架無

痛點 1 — 靜默的部分擷取與保真度崩壞(amem 答得最差)

陳述:擷取失敗最常見的形態不是報錯,是安靜地存了一個不完整或錯誤的東西。

一手證據:

  • kepano/defuddle#352(0 comment,2026-07-28,OPEN)— getElementSelector() 未 escape 冒號 id,React streaming SSR 觸發後 Defuddle 自己拋錯,靜默退回抓整個 <body>。這是「靜默部分擷取」的機制本體。 https://github.com/kepano/defuddle/issues/352
  • obsidianmd/obsidian-clipper#196(4 comment,2024-11-21)— jwhitley 回報存到了 頁面上從不可見的隱私聲明,可見內文一個字都沒存。他自己說「只發生在某些頁面, 我找不出規律」。https://github.com/obsidianmd/obsidian-clipper/issues/196
  • obsidianmd/obsidian-clipper#37(19 comment,32 reaction,CLOSED)— 「Download pictures to local」是該 repo 史上最高 reaction 的 closed issue。 遠端圖片連結會隨原站一起腐爛。https://github.com/obsidianmd/obsidian-clipper/issues/37
  • karakeep-app/karakeep#1652(16 reaction,OPEN)+ #1522(11 reaction)+ #594/#999/#1306 — 同一個 archive.org 中介需求被獨立提出五次,合計 27+ reaction。https://github.com/karakeep-app/karakeep/issues/1652
  • deathau/markdownload#366(7 comment,OPEN)— 「Only half of an article is extracted」,多位獨立回報者。一位指出根因是 Readability.js,另一位量化 「砍掉大約 16 段」。https://github.com/deathau/markdownload/issues/366
  • gildas-lormeau/SingleFile#1744(2 comment,OPEN)— Perplexity 與 ChatGPT 這類 對話視窗擷取被截斷。維護者:「It’s not an easy problem to deal with in a generic way.」https://github.com/gildas-lormeau/SingleFile/issues/1744
  • HN 44095939(bayindirh,2025-05-26)— 把需求講到最準:「Because I want the version I have seen. Not the edited/updated one.」 https://news.ycombinator.com/item?id=44095939

AMO 逐字評論獨立佐證同一件事。SingleFile 1033 則評分中 52 則低星,前三主題是 capture 不完整(4)、掉圖片與互動元素(3)、慢(2)。MarkDownload 有一則 3 星 (2023-06-17, alex):「it only converts what is visable on the screen rather than the whole page」。https://addons.mozilla.org/en-US/firefox/addon/single-file/reviews/

amem 現況:失敗,而且是最壞的一種失敗。

三個獨立缺陷疊在一起:

  1. 6000 字截斷,14/30 節點已中彈(clipper_bridge.rs:588)。
  2. 沒有 raw 備份,截斷後無法還原。
  3. 重剪整檔覆蓋,舊內容直接消失,沒有 diff、沒有版本、沒有 hash 可以偵測變化。

第三點值得單獨看。amem 用 URL 正規化把重剪折疊回同一個檔案,這解決了痛點 8。 但實作方式是覆寫,所以它同時製造了痛點 1。使用者兩週後重剪一個被改過的頁面, 第一版就永久消失了,而且系統不會告訴他。

對照 basicmachines-co/basic-memory#124(6 comment,4 reaction,OPEN)裡那句話: 「5 minutes into using it I realized I HAVE to have a mechanism for checkpoints… the LLM could easily mess up the memory in one prompt, and you want to be able to go back.」https://github.com/basicmachines-co/basic-memory/issues/124

這條落差最傷,因為 amem 的定位句押的就是這一格。全業界沒人做 chunk 級 hash, 這是真空象限。但 amem 目前只在 2/35 的面積上鋪了地基,而那 2 個還是不走 clipper 的 legacy PDF 節點。


痛點 2 — 存了找不回來(amem 答得差)

陳述:使用者的抱怨分兩種,檢索失敗與忘記自己存過。第二種更常被講。

一手證據:

  • HN 22105561(279 pts / 268 comment,2020-01-21)— hooande 一句話定義了它: 「My problem with bookmarks isn’t managing them, but remembering that they exist.」同 thread 的 JohnFen 補上為什麼「再 Google 一次」不算替代方案: 「Most of the things I bookmark are things that were really hard to find in the first place.」https://news.ycombinator.com/item?id=22105561
  • HN 44927588(jtqq,2025-08-16)— 自建 Emacs + org-roam + Elfeed + Wallabag 全套, 結論仍是「Retrievability is the biggest pain point right now」。下一步計畫是加 向量資料庫。https://news.ycombinator.com/item?id=44927588
  • karakeep-app/karakeep#2489(0 comment,2026-02-16,OPEN)— 「Full Text search STILL unavaliable… this seems as the most obvious use case」。六個月零回覆。 https://github.com/karakeep-app/karakeep/issues/2489
  • karakeep-app/karakeep#1955(4 comment,OPEN)— 部分關鍵字搜尋回傳不完整結果。 https://github.com/karakeep-app/karakeep/issues/1955
  • basicmachines-co/basic-memory#951(5 comment,OPEN,維護者自撰)— 有硬數字。 LoCoMo 1,986 個 query,single-hop R@5 basic-memory 0.486 vs mem0 0.592。 281 個獨有 miss 裡 146 是真 retrieval miss,主因是跨對話實體混淆。 https://github.com/basicmachines-co/basic-memory/issues/951
  • nashsu/llm_wiki#5(15 comment,OPEN)— 50GB+ 語料庫每次開啟 loading 極久, 且作者給不出建議上限。#397(3 reaction)— 6566 頁重建索引要 2 小時以上。 https://github.com/nashsu/llm_wiki/issues/397

amem 現況:失敗,而且失敗的方式很諷刺。

amem_recall 與 amem_ground 走 src/query/grep.rs,檔頭自述是「MVP search… Token-count scoring」。它把每個 wiki 檔整份讀進記憶體再數 token 命中次數。 35 個檔案可以,3,500 個不行。

諷刺的地方在於:BM25 加密集向量檢索已經寫好了。src/refindex.rs 948 行、 src/embed.rs 519 行,含 RRF 融合、bibliography chunk 過濾、向量指紋校驗。 但它服務的是 Zotero 的 78 篇 PDF 語料,不是 wiki。

也就是說,好的檢索建在對的技術上、錯的語料上。而 wiki 才是 clipper 的產物, 是產品入口。

~/.amem/ 沒有 index.md 也沒有 log.md,## Connections 在 31/35 個節點裡是空的。 所以「忘記自己存過」這半邊的痛點,amem 連最便宜的解法都沒有。


痛點 3 — 權限過寬引發的供應鏈恐懼(amem 答得差,且是無謂的差)

陳述:反對的不是 telemetry,是 <all_urls> 加 scripting 這個後門面積。使用者 會因此卸載,並且會寫下來。

一手證據:

  • karakeep-app/karakeep#2782(9 comment,2026-05-10,CLOSED)— 全調查最好的 單一物證。使用者 psla:「This permission is a no-go from me.」接著:「As much as I trust you and the project, supply chain attacks are a thing, and I generally don’t allow extensions that request full control over the page. Have you considered using optional permissions?」維護者 24 小時內改成 optional_permissions 並發 1.2.11。https://github.com/karakeep-app/karakeep/issues/2782
  • karakeep-app/karakeep#2790(2026-05-12,CLOSED)— 另一位獨立提出,附威脅模型: 「its a great backdoor if repo ever gets compromised」,並標 all_urls 🔴 HIGH、 scripting 🔴 HIGH。https://github.com/karakeep-app/karakeep/issues/2790
  • gildas-lormeau/SingleFile#1030(7 comment,OPEN 四年)— 「if I install this extension on Firefox, it will inject a script into every page I load even if I don’t intend to save the web page」。維護者答:必須永遠注入,無法關閉。 https://github.com/gildas-lormeau/SingleFile/issues/1030
  • obsidianmd/obsidian-clipper#165(2024-11-14)— brian-burton 因 0.9.6 新增下載權限 而卸載,並在 template 欄位寫「Sorry, I can’t provide a template because I’ve uninstalled Web Clipper」。https://github.com/obsidianmd/obsidian-clipper/issues/165
  • AMO 逐字,Instapaper 1 星(2025-08-30, dev421):「The permissions required by this add-on are ridiculous. “Access your data for all websites”!? You only need the current tab URL!」https://addons.mozilla.org/en-US/firefox/addon/instapaper-official/reviews/

amem 現況:失敗,而且有一半的權限沒有換到任何功能。

manifest 有 <all_urls> 加 12 個 permission,含 tabCapture、downloads、 scripting、tabGroups。這正是讓 psla 拒裝 karakeep 的權限輪廓。

更糟的是四個 always-injected content script 裡有三個是 8 行 stub。 claude-autosave.js、chatgpt-autosave.js、gemini-autosave.js 都只做一件事: 設一個 flag 然後 console.debug。它們對 claude.ai、chatgpt.com、gemini.google.com 永遠注入,換到零功能。

這是純負債。使用者付出「這個 extension 在我的 AI 對話頁上跑程式」的信任成本, 產品沒有拿到任何東西。真正的 AI 對話擷取在 extractors/ 裡,是按需 executeScript 派送的,不需要宣告 content script。

補一個對 amem 有利的觀察。維護者本人有時是隱私鷹派,而使用者站在他那邊。 kepano 在 #112 以安全理由永久拒絕雙向整合:「I don’t think most Obsidian users would trust giving their browser full access to their Obsidian vault.」同 thread 的 writtenfool 附議:「I would not want, at anytime, for a web app to have access to my local app.」amem 的 daemon 是本機的,這是可以講的故事,但前提是權限面要乾淨。


痛點 4 — AI memory 錯誤、過期、無法驗證(amem 部分答)

陳述:2026 年最熱的一條。抱怨不是記不住,是自動記下了錯的東西再拿去誤導使用者。

一手證據:

  • HN 48776232「Memorizing session transcripts isn’t useful」(180 pts / 159 comment, 2026-07-03)。Fabricio20 最完整:「I specifically disabled claude memory in a project because it kept writing down thigns to memory that didn’t need to be in memory, including severly wrong statements that then would confuse it later.」 他還指出自動記憶被自動重新啟用,要同時關 autoMemoryEnabled 與 autoDreamEnabled。https://news.ycombinator.com/item?id=48776232
  • 同 thread,mastax 指出污染的正回饋:「If you allow any low value things into memory, Claude will notice that established pattern and start trying to add low value memories」。semiquaver:「Claude’s own memories severely mislead it」。
  • HN 47531882(ravikirany22,2026-03-26)— 唯一有量化的一筆:「We’ve been auditing TypeScript repos and finding 10-84% of symbol references in AI config files are stale… it’s getting a confident lie.」 https://news.ycombinator.com/item?id=47531882
  • nashsu/llm_wiki#458(2 reaction,OPEN)— 「Concerns About LLM-Based Knowledge Bases: Hallucinations and the Need for Faithful Original Text Retrieval」: 「When I query specific regulations or raw tables, the output is often hallucinated… frequently diverges significantly from the original content.」 https://github.com/nashsu/llm_wiki/issues/458
  • nashsu/llm_wiki#229(5 reaction,OPEN)— 要一個修正入口。使用者指出唯一的修復 路徑是手改 markdown,而這與「Wiki 頁面全部由 LLM 維護」的核心設計理念相悖。 https://github.com/nashsu/llm_wiki/issues/229

這裡有一個對 amem 極重要的相關性,是 GitHub 挖掘得到的最有價值推論之一。 幻覺抱怨只出現在 llm_wiki 一家。 其他八個 repo 的 LLM 只寫 tag 與 summary, 不寫知識本體,所以沒有幻覺抱怨。llm_wiki 是唯一讓 LLM 撰寫 wiki 的,也是唯一 被抱怨幻覺的。

amem 的 clipper 節點把 LLM 產物(tldr)與原文(Content)並置。這個結構比 llm_wiki 安全。但這個安全性依賴原文真的在旁邊,而 47% 的節點原文被截斷了。

amem 現況:部分答,且答案正在漏氣。

有的部分:frontmatter 帶 url、first_seen_at、captured_at 清單。這比 Claude Code 內建 memory 多(後者只記 modified 時間戳,不記來源)。tldr 與 Content 分離 也對。

漏的部分:沒有內容 hash,所以無法偵測來源頁面變了。amem_ground 的 tool description 自述是「search the local wiki for the topic and return JSON hits」 (src/mcp/mod.rs:172),這是對自有 wiki 的檢索加引用,不是驗證。 SPEC 定義的 amem_factcheck 仍在 RFC-003 規劃中。

refverify(CrossRef + DataCite)與 refcheck(BM25 + 向量)是真的驗證, 而且做得好。它們只服務 PDF 語料。


痛點 5 — 離線時擷取不進去(amem 忽略)

陳述:擷取的那一刻經常是離線的,沒有工具會排隊。

這是本次調查最意外的一條,前次掃描與 design memo 都沒有列。

一手證據:

  • karakeep-app/karakeep#1077(9 comment,52 reaction,CLOSED)— 「Offline cache on Mobile app」。 https://github.com/karakeep-app/karakeep/issues/1077
  • karakeep-app/karakeep#2401(17 comment,20 reaction,2026-01-14)— 自架者離開 LAN 之後,分享選單一直轉。 https://github.com/karakeep-app/karakeep/issues/2401
  • karakeep-app/karakeep#274(3 comment,20 reaction)— 「Cache Hoards while away from server」。合計約 92 reaction。 https://github.com/karakeep-app/karakeep/issues/274
  • obsidianmd/obsidian-clipper#828(0 comment,2026-05-05,OPEN 三個月無人回)— 同一個形狀的靜默資料遺失:「In my case I lost ~10 job-prospect clips this morning before realizing none of them had hit disk. There is no recovery path other than redoing the captures from browser history.」 https://github.com/obsidianmd/obsidian-clipper/issues/828

amem 現況:忽略,而且暴露面比 obsidian-clipper 更大。

background.js 唯一的 alarm 是 amem-bridge-reconnect,只服務 bridge 輪詢, 不是擷取佇列。沒有 retry、沒有 navigator.onLine 判斷、沒有 pending capture 儲存。

amem 的架構讓這件事比競品更常發生。obsidian-clipper 只需要 Obsidian 這個 app 存在。amem 需要 amem mcp serve 這個 daemon 正在跑。daemon 沒跑、剛重開機、 或 crash 了,使用者按下擷取就是失敗。而擷取失敗的當下,使用者已經離開那個頁面了。


痛點 6 — 讀了不回頭的墳場(amem 忽略,但這條被高估了)

陳述:痛感真實且橫跨十年,但它不是使用者會為之投票的痛點。

HN 證據極厚,判定 dominant:

  • HN 44066646(rossant,2025-05-22,母 thread 1222 pts / 761 comment)— 「I just exported my data and found 13,000 unread articles out of a total of 34,000.」https://news.ycombinator.com/item?id=44066646
  • HN 46880996(al_borland,2026-02-04)— 「They are where my good intentions go to die.」https://news.ycombinator.com/item?id=46880996
  • HN 16306040 與 HN 44925963(pixelmonkey,相隔七年半講同一句)— 「I sometimes describe Instapaper as ‘/dev/null for web content’.」 https://news.ycombinator.com/item?id=44925963
  • HN 36146108(bsnnkv,2023-05-31)— 「both note apps and read it later queues/apps are where ideas go to die.」https://news.ycombinator.com/item?id=36146108

但 tracker 的 reaction 質量說了另一件事。

resurfacing 相關的三個 issue 合計 41 reaction(karakeep #435 10、#705 16、 #863 15)。同一個 repo 裡,單一個登入擷取 issue(#172)就有 57 reaction, 離線佇列合計 92 reaction。

而且十個 repo 裡沒有任何一個有「我的存檔是墳場」這種 issue。它只以間接形式 出現,例如 karakeep#863:「we keep bookmarking the pages or videos, but we do not have time to fully read it.」

這對 design memo 是一個修正。 memo 把 #3「write-only graveyard」與 #6「friction」 當成「真正殺死 clipper 的兩個痛」。HN 支持墳場的存在,但 tracker 說使用者不為它 投票,他們為「存不進去」和「找不到」投票。

amem 現況:忽略。 沒有 resurfacing、沒有 index.md、沒有隨機回顧, ## Connections 是空的。但依據上面的證據,這應該排在補完痛點 1 與 2 之後, 不該當頭條。


痛點 7 — 登入頁與付費牆擷取(amem 結構性解決)

陳述:server 端 crawler 永遠看不到登入使用者看到的東西。每個專案都獨立重新發現 唯一解是「從使用者自己的瀏覽器擷取」。這是全調查最高 reaction mass 的主題。

一手證據:

  • karakeep-app/karakeep#172(52 comment,57 reaction,CLOSED)— karakeep 史上最高 reaction 的 issue,「Local Scraper (Use browser auth)」。 結局是改成相容 SingleFile extension 的 REST endpoint。server 端 crawler 輸給了 瀏覽器 extension。 https://github.com/karakeep-app/karakeep/issues/172
  • karakeep-app/karakeep#414(66 comment,40 reaction,CLOSED)— 存到 cookie 同意視窗而非文章。有使用者的本機 Llama3.2 去摘要了付費牆公告。 https://github.com/karakeep-app/karakeep/issues/414
  • karakeep-app/karakeep#2814(4 comment,2026-05-16,OPEN)— 最關鍵的一筆。 karakeep 已經出貨 client-side crawling,付費牆仍然失敗:「Expected: Full article text is archived. Actual: Only the article preview/teaser is saved.」 他的 workaround 是用 Obsidian Web Clipper 抓,再寫 Python 腳本同步進 karakeep。 https://github.com/karakeep-app/karakeep/issues/2814
  • karakeep-app/karakeep#2885(2026-06-13,維護者本人開)— 「As part of a recent reddit crackdown on crawlers, the endpoint that we were using… is now completely blocked.」https://github.com/karakeep-app/karakeep/issues/2885
  • 2026 年的新退化不只 Reddit:#2952 archive.is 又要 captcha(2026-07)、 #2423 Cloudflare 封鎖(2026-01)、#2381 JS/cookie 攔截頁被存下來(2026-01)。

amem 現況:結構性解決,這是最該押的一格。

content script 跑在使用者已登入的分頁裡,沒有 cookie 轉移問題,沒有 server 再抓一次 的問題。~/.amem/wiki/ 裡真的躺著 Gmail 與 openreview 登入後才看得到的 capture。

趨勢對 amem 有利。bot wall 在 2026 年系統性收緊,server-side crawler 這條路正在 失效,而「已登入的真實瀏覽器」正在變成唯一可行路徑。

倫理阻力方面,一手來源查到的是零。使用者一律把它當純能力缺口。 最常見的自我正當化是「我付錢看的」(#2236,bronikowski:「when I read something I paid for」)。

但這一格已經有人在賣了。 見 §3 的 Web2MD。


痛點 8 — 重複擷取(amem 解決,但用資料遺失換)

陳述:沒有工具會在你重存之前告訴你「你已經存過了」。而且有兩位維護者公開宣告 這件事做不到。

  • obsidianmd/obsidian-clipper#112(8 comment,16 reaction,OPEN)— kepano 是架構性拒絕:「For now there are no plans to create a two-way integration… This is intentional.… I think it poses a security risk.」 #323 與 #521 兩個同樣的請求都被 closed as duplicate。三次請求,一次永久拒絕。 使用者 CarcajadaArtificial:「I just want to not end up with duplicate web clippings. That’s all.」https://github.com/obsidianmd/obsidian-clipper/issues/112
  • gildas-lormeau/SingleFile#1642(OPEN)— 維護者:「This is not really possible from a technical point of view, for privacy reasons. Extensions cannot scan folders on the filesytem.」#956 同樣答「A database is required… it seems complicated to implement reliably」。 https://github.com/gildas-lormeau/SingleFile/issues/1642
  • karakeep-app/karakeep#486(10 comment,7 reaction,CLOSED)— 十家裡唯一出貨 「已存過」指示器的,花了 13 個月。 https://github.com/karakeep-app/karakeep/issues/486
  • karakeep-app/karakeep#864(3 reaction,OPEN)— 去重不處理結尾斜線, example.com/a 與 example.com/a/ 算兩筆。 https://github.com/karakeep-app/karakeep/issues/864

amem 現況:解決了,而且解得比業界好,但代價是痛點 1。

node_id_from_url(clipper_bridge.rs:425-471)做的事比 karakeep 多: 剝除 11 個追蹤參數、小寫 host、去尾斜線、arxiv/github/HN/x.com 各有專屬 id 規則。karakeep #864 抱怨的尾斜線問題,amem 在 url_canonical:521 已經處理。 #633 抱怨的追蹤參數,amem 在 :500-503 已經處理。

這是一個乾淨的勝場。兩位維護者公開說做不到的事,amem 因為有本機 daemon 而做得到。

但實作用覆寫達成去重。 舊內容消失,沒有版本、沒有 hash。所以 amem 把「重複 spam」換成了「靜默資料遺失」。這兩個痛點在 amem 身上是同一行程式碼的兩面 (clipper_bridge.rs:381 的 std::fs::write)。


痛點 9 — 本機 LLM 連線摩擦(amem 部分答)

陳述:「接你自己的模型」是 AI memory 工具的第一大死路。錯誤訊息很籠統, timeout 不可設定。

  • karakeep-app/karakeep#185(20 comment,OPEN 27 個月)— 標題就是 「How to verify hoarder app is working with the local ollama」。 https://github.com/karakeep-app/karakeep/issues/185
  • karakeep-app/karakeep#424(19 comment,OPEN)— 使用者拿到的 log 是 inference job failed: TypeError: fetch failed,沒有 host、沒有 status。 維護者:「Unfortunately the logging does not show where the issue happens.」 https://github.com/karakeep-app/karakeep/issues/424
  • MODSetter/SurfSense 有十個 open 的本機 LLM 連線 issue(#1379、#1616、 #1550、#587、#517、#1464、#1405、#1394、#1493、#1518)。 https://github.com/MODSetter/SurfSense/issues/1379
  • obsidianmd/obsidian-clipper#515(2 reaction,OPEN)— 「Custom Ollama provider requires non-existent API key」,UI 要一個該 provider 根本沒有的憑證。 https://github.com/obsidianmd/obsidian-clipper/issues/515
  • karakeep 有五個獨立 open issue 都是同一件事:本機模型比硬寫的 timeout 慢 (#2770、#2679、#2994、#1129、#1806)。

amem 現況:部分答,而且方向對。

clipper_bridge.rs 的 summarize 有三條路徑:summarize_via_claude(:256)、 summarize_via_codex(:286)、summarize_via_ollama(:314)。有 CLI fallback 是 對的設計,使用者不必先裝 ollama 才能用。這比 SurfSense 與 karakeep 好。

未查證的部分:三條路徑全掛時的錯誤是否對使用者可讀。依據痛點 5 的分析, 擷取失敗沒有佇列,所以錯誤處理路徑值得單獨測。


痛點 10 — 上架被下架的存亡風險(amem 已上架)

陳述:clipper 的死亡證明是 Google 簽的,不是 bug 數量。

  • deathau/markdownload#378(9 reaction,OPEN)— 該 repo 最高 reaction 的 open issue,標題是「This extension is no longer available because it doesn’t follow best practices for Chrome extensions」。同一件事被獨立提報三次 (#357 3 reaction、#393 1 reaction),合計 13 reaction。3,997 star 的專案 死於下架。https://github.com/deathau/markdownload/issues/378
  • #393 裡有陌生人貼未審核的 fork,下一位留言者回「now the deployments website is unable, the zip is not found」。使用者被推向已經死掉的未簽名 fork。
  • HN 46880866(2026-02-04)— 「most of the things I come across are dead and gone, or seem abandoned somehow」。2026 年有三個 Show HN 把死亡寫進標題: 「because the others keep dying」(48745735)、「because Mozilla killed it」 (46956985)、「built after Pocket shut down」(48449568)。 https://news.ycombinator.com/item?id=46880866

amem 現況:已上架,前次掃描這一格判錯了。

CWS 公開頁查得到 amem Clipper v0.3.0,更新 2026-08-13,3 users,0 則評論。 llm_wiki 的 extension 至今仍要 chrome://extensions load unpacked。這一格 amem 贏。

順帶一個 null result:「load unpacked / Developer Mode」在十個 repo 裡是零筆 issue。 這不是真實使用者痛點。design memo 用 Developer Mode 摩擦當作不做 chrome.userScripts 的理由,那個結論仍然對(安全面與品牌面成立),但摩擦論據 本身沒有一手支持。


3. Part B — 前次掃描漏掉的競品

只列前次掃描的搜尋角度會漏掉的。星數與日期我親自用 gh api 於 2026-08-17 複驗。

3.1 Web2MD — 架構逐項複製,且已在賣

https://web2md.org/

今天出貨: Chrome extension、16 個站台專用 extractor(含 Reddit、YouTube、 GitHub、arXiv)、npx web2md CLI、npx web2md-mcp-server MCP server、 GPT-4 與 Claude token 計數、watch mode。

Agent Bridge 是核心賣點,官網原文:「Agent Bridge uses your actual Chrome with your cookies and login state — Reddit can’t tell the difference from normal browsing」。

這是 amem 痛點 7 那格優勢的逐字複製,而且已經包裝成產品在賣。 定價 Pro $4.17/月(原價標 $15)。免費層每日 3 次轉換。非開源。

單一勝出軸:CLI 批次同步。 amem 是一頁一頁點,Web2MD 有 web2md sync。

需注意:前次掃描漏掉它,是因為它不在 GitHub 上,搜 repo 搜不到。它也是 web2md.org SEO 內容農場的擁有者,該站產出偽裝成「honest review」的競品評測。 評論挖掘的子代理明確標記了這一點並拒絕採用其內容。

3.2 Letta — SKILL.md 自動生成,全自動且已出貨

https://github.com/letta-ai/letta-code — 3,017 star,v0.30.25 發佈於 2026-08-17(今天)。

我親自讀了 src/agent/subagents/builtin/reflection.md,逐字引用:

name: reflection description: Background agent that reflects on recent conversations to update memory and maintain skills

Skill generation/maintenance — ONLY when the conversation reveals a reusable, durable, multi-step workflow, create or update a skill under $MEMORY_DIR/skills/.

Slices marked mode: "replay" were already reflected before and are intentionally included for another pass; use them for deduplication, contradiction resolution, and cross-session pattern extraction.

最後一段特別重要。它殺掉的不只是「自動生成 SKILL.md」,還包括「跨 session 縱向 recurrence detection」這個備援縫。Letta 明文在做跨 session 模式抽取。

它還有 amem 沒設計的東西:skill 生命週期回收。操作集是 update / extend / deprecate / split / create,寫進 MemFS git repo,有版本。

單一勝出軸:同一個 artifact,更完整的生命週期,已在有資金的產品內出貨。

它缺的是 capture surface。它的語料是對話,不是網頁。

3.3 OpenAI Computer History — 廠商層的 sensor → recurrence → skill

https://learn.chatgpt.com/docs/customization/computer-history

官方文件原文:「Computer History turns your activity across apps and websites into memories and a timeline that ChatGPT and Codex can reference.」以及 「When Computer History notices repeatable work, a timeline entry can suggest a skill or automation.」

這是「跨網站活動 → 偵測重複 → 建議 skill」,由模型廠商出貨。

限制就是 amem 剩下的縫:僅 macOS 桌面版、Pro/Business/Enterprise、預設關閉、 EEA/瑞士/英國不可用。而且它讀 interaction event(點擊、輸入、app 切換), 明確不讀頁面內容本身。

單一勝出軸:通路加作業系統級 sensor。

3.4 其他架構撞車的

名稱出貨成熟度(gh api 複驗)單一勝出軸
RowboatMac/Win/Linux 桌面 app,Apache-2.017,290 star,pushed 2026-08-17論述撞車。README 寫「living Obsidian-style backlinked knowledge graph」「All data is stored locally as plain Markdown」
claude-memCLI,v13.15.290,938 star,pushed 2026-08-17「捕獲即記憶」的心智位置佔有量。無 skill 生成、無網頁擷取
openhumandmg/exe/deb/AppImage36,323 star,created 2026-02-18發版速度近乎每日
open-knowledge桌面 app + CLI,GPL-3.03,480 star,created 2026-06-03自述「AI-native markdown IDE and LLM wiki」,前端成熟度
screenpipe桌面20,980 star,YC S26捕獲面總量最大,隨時可加瀏覽器 lane

Rowboat 有一點要更正子代理的判讀。它的內建瀏覽器隔離於使用者主瀏覽器, README 明寫使用者必須在裡面重新登入。這是 amem 的優勢,不是 Rowboat 的。

3.5 格式層風險

GoogleCloudPlatform/knowledge-catalog — 8,667 star,created 2026-05-04,pushed 2026-08-15。Google 把「LLM 編譯的 markdown wiki」變成規格。子代理報告 v0.2 加入 trust signals,這直接踩進 amem 的 citation-grounding 差異化。該 v0.2 細節我未親自複驗。建議獨立評估 OKF 相容性。

3.6 負面結果(掃過、確認沒有)

這些空白是 amem 剩餘的可辯護空間。

  • MCP 生態沒有 clipper。 punkpeye/awesome-mcp-servers(92,462 star)全文搜 clipper / web clip / clip page 零筆。modelcontextprotocol/servers README 無任何 clipper server。
  • 大型 memory 玩家都沒有 capture surface。 graphiti、letta、memU、cognee、 txtai、memoripy、memento-mcp 的 repo 樹內都沒有 browser extension 目錄。 mem0ai/mem0-chrome-extension 已 archived,最後 push 2026-03-23。 最大玩家退出了瀏覽器擷取賽道。
  • browser-control MCP 沒有一個寫 KB。 chrome-devtools-mcp(49,285)、 browser-use(109,480)、playwright-mcp(36,195)、Skyvern、stagehand 全部純自動化、 零持久化。
  • 「從我剪過的東西裡找跨篇反覆出現的模式」沒有人做。 搜過 chrome extension SKILL.md、web clipper skill agent、capture to skill、 bookmarks to agent skill、recurring patterns into skills 全部零相關結果。
  • YC 四個 2026 batch 共 638 家,零家做「瀏覽器 extension + 本機 markdown wiki」。 (子代理用 YC 自家 API 分頁,我未複驗)

4. RFC-007 的 moat 判決:GONE

RFC-007 寫「Every clipper competitor stops at storage; nobody closes the loop into agent capability. Step 3 is the moat.」

前次掃描已經把它修正成「moat 不是 step 3 整段,而是 step 3 裡的自動生成那半段」。 那個修正現在也不成立了。 三份證據,全部我親自複驗。

證據一:Anthropic 官方說這件事不需要工具。

https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices 逐字:

Claude models understand the Skill format and structure natively. You don’t need special system prompts or a “writing skills” skill to get Claude to help create Skills. Simply ask Claude to create a Skill and it generates properly structured SKILL.md content with appropriate frontmatter and body content.

護城河不能建在供應商文件標註「此處不需要工具」的動作上。

證據二:Claude Code 二進位檔本身就會自動寫 SKILL.md。

我在本機 /Users/ydwu/.local/share/claude/versions/2.1.233 裡抓到字串, 逐字(去掉 UTF-16 間隔):

If this repo has no project verify skill (.claude/skills/verify/SKILL.md), that is a reason to run /verify, not to skip it: the run creates that file, saving the working build-and-drive recipe for future sessions.

同一份二進位檔裡 run-skill-generator 出現 7 次,註冊為 bundled skill, menuDescription 是「Create a skill that knows how to run this project’s app」。

/verify 寫檔是跑 verify 的副作用,使用者沒有要求產生 skill。 這就是無人在迴路的自動蒸餾,而且它是內建的、免費的。

證據三:Letta 已全自動出貨同型機制。 見 §3.2 的逐字引用。

證據四:「X → SKILL.md」已是 commoditized 類別。 星數我逐一用 gh api 複驗:

專案Star做什麼
microsoft/SkillOpt16,074把 SKILL.md 當可訓練參數優化,含 nightly self-evolution
yusufkaraaslan/Skill_Seekers14,77318 種來源(docs/repo/PDF/EPUB/YouTube)→ SKILL.md
bergside/design-md-chrome2,663Chrome extension 剪一頁 → 產 SKILL.md

最後一列要特別看。它已經佔住「Chrome extension 剪頁 → 吐 SKILL.md」這個一模一樣 的手勢。

還剩什麼

一條窄縫,是 feature 不是 moat:沒有人從「跨站累積數週的 clipped web pages」這個 語料出發做蒸餾。

三家的語料都不是網頁。Anthropic 的 capture surface 是螢幕錄影、當下對話、本機 repo。 Letta 的是對話 transcript。OpenAI Computer History 讀 interaction event,明確不讀頁面 內容。design-md-chrome 是一頁換一檔,沒有語料庫概念。

這條縫真實存在,但它撐不起「moat」這個詞。它撐得起一個功能。

建議:不要再把「產生 SKILL.md」寫進對外定位。


5. 建議:第一個該改的東西

先讓 clipper 節點停止破壞證據,再談任何新功能。

理由是這一個改動同時關掉排名第 1、第 2、第 4、第 8 四個痛點的落差, 而且它是 amem 唯一真空象限(chunk 級 hash provenance)的地基。

具體是三件小事,都在 clipper_bridge.rs 一個檔案裡:

  1. 移除 6000 字截斷,或把完整原文寫進 ~/.amem/raw/。 目前截斷後無副本可還原 (:588)。
  2. frontmatter 加 content_sha256,並對 chunk 逐段 hash。 cite.rs 已有 text_sha256,PDF 路徑在用,clipper 路徑沒接上。
  3. 重剪時若 hash 改變,另存版本而不是覆寫。 目前 :381 直接 std::fs::write。

做完這三件,「amem 記錄你真正讀過的東西,而且每一句話都能追回原文」才從願景變成 產品描述。在那之前,這句話的後半是不實陳述。

第二順位是把 refindex 的 BM25 加向量檢索接到 wiki 語料上。程式碼已經寫好了 (refindex.rs 948 行、embed.rs 519 行),只是指向錯的語料。這關掉痛點 2。

第三順位是刪掉三個 stub content script。這是零成本降低權限面(痛點 3), 它們目前換到零功能。


6. 取樣缺口與未查證項目

誠實揭露,這些不要當成已經查過。

  1. Reddit 完全未取樣。 www.reddit.com、old.reddit.com、api.reddit.com 三個 domain 都被 harness 阻擋。r/ObsidianMD、r/PKMS、r/selfhosted、r/DataHoarder、 r/LocalLLaMA、r/ClaudeAI 全部沒有覆蓋。自架與 DataHoarder 族群的觀點缺口很大。
  2. Chrome Web Store 的 1-3 星評論全部拿不到。 CWS 從瀏覽器核心層禁止 content script 注入 chromewebstore.google.com(bridge 回 "The extensions gallery cannot be scripted."),這條路永久不可行,不是 bug。 WebFetch 只回傳預設「最相關」排序,幾乎全 5 星。受影響最大的是 Recall (getrecall.ai),它是 amem 最直接的 AI-summary 競品,負評完全沒有資料。
  3. AI 摘要品質主題嚴重取樣不足。 AMO 唯一有 AI 的樣本是 Raindrop 的 2 則。 不要從這 2 則推論任何結論。
  4. AMO 樣本有 Firefox 結構偏誤。 31 則登入/cookie 抱怨有相當比例是 Firefox 跨站 cookie 預設造成的,不能直接外推到 Chrome。
  5. 「AI 摘要在筆記庫裡幻覺」找不到第一手案例。 forum.obsidian.md 搜 「AI summary hallucinate」零筆。obsidian-clipper 的 Interpreter 相關 issue 全部是 整合層故障,沒有一筆抱怨輸出內容錯誤。含意:使用者對 agent memory 的幻覺很敏感, 對 clip 時的摘要幻覺還沒有痛感,可能因為原文還在旁邊。
  6. 付費牆擷取的倫理/ToS 抱怨:零筆。 這是 null result,不是沒查。
  7. notes/wiki → 自動生成 skills:完全沒有人要求。 使用者的行為是從 session 事後蒸餾 skill(HN 47543139,627 pts / 265 comment,ccosky:「Anytime I do something as a one-off that I know I’ll do in the future, at the end of the session I’ll ask Claude to write a new skill based on what it did」)。 方向與 RFC-007 的假設相反。
  8. 未親自複驗: OKF v0.2 的 trust signals 細節、YC batch 歸屬與家數、 Minibase 與 LLMnesia 的使用者數、Gemini CLI 那個「require recurrence evidence before extracting skills」PR 的編號。
  9. amem 三條 summarize 路徑全掛時的錯誤可讀性未測。 依痛點 5 的分析值得單獨測。

RFC-001 — Function-based v0.1 architecture

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-05-09
  • Related: RFC-002 (Clipper skills catalog UI), RFC-003 (recording orchestration), archived RFC-001 (bridge-first), SPEC.md § Mental model

TL;DR

Ship v0.1 as hardcoded MCP tools and Rust functions, not a skill engine. The three must-have features all have the same shape — agent calls a tool, librarian dispatches to a Rust function, function may bounce DOM-side work through the bridge to amem Clipper. There is no plugin runtime, no DSL, no sandbox. Skill engine is deferred to v0.5+, gated on the rule of three: ship when we have 5+ site adapters, 3+ recording scripts, or external user requests for installable skills.

7.5 days of hardcoded code now beats 13 days of skill-engine plumbing for zero current users. We promote to a real engine the day the third copy of “this looks the same as the previous one” arrives.

Motivation

The previous round of RFCs (archived 2026-05-09) circled around a generic “skills” concept — pluggable units the agent could discover and invoke. The abstraction was attractive on paper (catalog, sandbox, manifest, capability flags) and load-bearing on nothing: we have one user (the operator), zero external testers, zero shipped sites, and a CWS deadline.

Three forces pushed us to lock the architecture as function-based:

  1. No third instance. Out of the v0.1 must-haves, only the site-adapter family has any repetition (arxiv, github, hackernews — three sites, one shape). Recording is one-of-one; bridge-driven Chrome ops are one-of-one. The rule of three (Refactoring, Fowler) says abstract on the third occurrence — we have it for adapters and only for adapters, and even there the shape is “match URL → call extractor,” which is a Rust match block, not a runtime.
  2. Engineering cost is dispositive. A plausible v0.1 skill engine — manifest schema, loader, sandboxed JS runtime in the librarian, capability gating, registry, sidepanel discovery, error surfaces — is ~13 days. A function-based v0.1 is ~7.5 days (3 site adapters at 1d each, recording at 2d, bridge MCP wiring at 1.5d, polish at 1d). The 5.5-day delta is the entire CWS submission window.
  3. CWS positioning still works without an engine. The Chrome Web Store listing claims “the first agent-callable Chrome skills catalog” — the user-facing surface (RFC-002) renders three hardcoded skills as if they are an installable catalog. v0.1 ships the shape of the product, v0.2 ships the actual installability. The gap is honestly disclosed in-UI (“Custom skills coming v0.2”) so we’re not lying — we’re shipping the minimum that lets the rest of the story land.

The core insight: the operator (and Claude Code, when it’s the operator acting on the user’s behalf) does not care whether chrome_navigate is implemented as a Rust function or as a sandboxed JS skill. They care that the tool exists and works. Build the tools first. Generalise on the third occurrence.

Proposal

1. Component split — librarian (Rust) vs clipper (JS)

Two binaries, one bridge, one rule for splitting work:

       ┌────────────────────────────────────────────────────────┐
       │  Frontier model (Claude / GPT / Gemini)                 │
       └────────────────────────────────────────────────────────┘
                              │  MCP (stdio)
                              ▼
       ┌────────────────────────────────────────────────────────┐
       │  amem Librarian (Rust binary, formerly amem-sh)         │
       │  • All heavy work: capture pipeline, compile, fact-     │
       │    check, recording orchestration, site adapters,       │
       │    storage, OCR, screencapture                          │
       │  • Hosts MCP server                                     │
       │  • Hosts bridge server (loopback WS 7600)               │
       └────────────────────────────────────────────────────────┘
                              │  ws://127.0.0.1:7600
                              ▼
       ┌────────────────────────────────────────────────────────┐
       │  amem Clipper (Chrome MV3 extension, JS)                │
       │  • DOM-only work: read/write document, observe URL,     │
       │    capture visible tab, render sidepanel UI             │
       │  • Holds zero authoritative state (config lives in      │
       │    librarian's config.toml)                             │
       │  • No fetching, no parsing, no LLM calls,               │
       │    no file I/O — those live in the librarian            │
       └────────────────────────────────────────────────────────┘

Rule of thumb: if a step requires the file system, an LLM, ffmpeg, OCR, or a network fetch beyond the active tab, it belongs in the librarian. If it requires a DOM, the active page, or pixel-level capture of the active tab, it belongs in the clipper. Anything else (URL matching, string manipulation, scheduling) belongs in the librarian — the clipper is a sensor, not a brain (per archived RFC-001’s positioning, still in force).

2. Bridge protocol — action / params / id envelope

Single loopback WebSocket on 127.0.0.1:7600. All messages, both directions, share one envelope:

{
  "id":     "<uuid-v4>",
  "action": "<verb>",
  "params": { ... },
  "token":  "<bridge token from ~/.amem/bridge.token>"
}

Replies use the same shape, with "action": "<verb>_result" and "id" matching the request. Errors carry {"action":"error","id":"...","error":{ "code","message"}}. The same envelope handles librarian→clipper commands (e.g. chrome_navigate) and clipper→librarian events (e.g. auto_capture).

Initial verb set for v0.1:

DirectionVerbPurpose
L → Cchrome_navigateNavigate active tab to URL
L → Cchrome_clickClick DOM element by selector
L → Cchrome_extractRead DOM (selector → text/html/attrs)
L → Cchrome_screenshottabs.captureVisibleTab (best-effort, see §6)
L → Cchrome_waitWait for selector to appear / URL pattern match
L → Center_recording_overlayToggle “recording” sidepanel state for demos
C → Lauto_captureSidepanel-toggled URL match fired; please ingest
C → Linvoke_skillUser clicked a Run-button in the sidepanel
L ↔ Cping / pongLiveness

Security posture inherits from archived RFC-001 unchanged: loopback bind, origin allow-list (only our extension IDs), 32-byte token in ~/.amem/bridge.token (mode 0600), token rotates on librarian restart.

3. The three must-have features → tool + function map

Each feature is one MCP tool (or a small fixed set), each backed by one Rust function in the librarian. No registry, no plugins.

3a. Claude Code operates already-logged-in Chrome

MCP tools shipped:

chrome_navigate(url: string) -> { status, final_url }
chrome_click(selector: string, opts?: { wait_after_ms }) -> { ok }
chrome_extract(selector: string, kind: "text"|"html"|"attrs") -> { value }
chrome_wait(selector: string, timeout_ms: number) -> { ok }
chrome_screenshot(selector?: string) -> { png_path }

Implementation in librarian:

#![allow(unused)]
fn main() {
// crates/amem-librarian/src/tools/chrome.rs
pub async fn chrome_navigate(url: &str) -> Result<NavigateOut> {
    let req = BridgeReq::new("chrome_navigate", json!({ "url": url }));
    BRIDGE.send(req).await?.into()
}
}

The librarian is just a thin pass-through here; the clipper does the actual DOM work in chrome.scripting.executeScript. This is the only path — agents do not get raw access to the bridge or to the extension. The MCP boundary is the contract.

Why not use playwright / puppeteer / chrome-devtools-protocol? Because the operator’s authenticated Chrome (Gmail, LinkedIn, internal tools) is the whole point. CDP-based tooling either drops the user’s profile (Chrome 136+ --remote-debugging-port + --user-data-dir constraints) or breaks keychain-backed flows (Chrome for Testing’s ad-hoc signing). Going through amem Clipper, which lives inside the user’s real Chrome, is the only path that doesn’t lose login state. (See ~/.claude/CLAUDE.md § Browser Automation for the gory details.)

3b. URL-mentioned content auto-saves to wiki

MCP tool shipped:

amem_capture_url(url: string, mode?: "auto"|"force") -> {
  cite_key, amem_uri, status: "captured"|"already_have"|"unsupported"
}

Two trigger paths feeding the same function:

  1. Agent-side: any agent (Claude Code, Cursor) that mentions a URL in its reply calls amem_capture_url(url) directly via MCP. The librarian matches the URL against the hardcoded adapter table and dispatches.
  2. Browser-side: amem Clipper observes navigation events; if the URL matches a hardcoded adapter pattern and the user has the matching skill toggle on (RFC-002), clipper sends auto_capture over the bridge, which calls the same function.

Hardcoded adapters in crates/amem-librarian/src/adapters/:

#![allow(unused)]
fn main() {
// crates/amem-librarian/src/adapters/mod.rs
pub fn dispatch(url: &Url) -> Option<Box<dyn Adapter>> {
    match url.host_str()? {
        "arxiv.org" | "www.arxiv.org"          => Some(Box::new(arxiv::Arxiv)),
        "github.com"                           => Some(Box::new(github::Github)),
        "news.ycombinator.com"                 => Some(Box::new(hackernews::HackerNews)),
        _                                      => None,
    }
}

trait Adapter {
    async fn extract(&self, url: &Url) -> Result<CaptureRecord>;
}
}

arxiv.rs, github.rs, hackernews.rs each ~150–250 lines, each a straightforward fetch + parse. They share a tiny CaptureRecord struct; they do not share a runtime. When the fourth adapter ships, we revisit (see §6 “When to revisit”).

3c. Scripted Chrome-only recording

MCP tool shipped:

amem_record_demo(script_path?: string, script_inline?: string)
  -> { amem_uri: "amem://recording/<uuid>", mp4_path }

Full mechanics live in RFC-003. From this RFC’s point of view it is one more Rust function in the librarian — record_demo() — that:

  1. parses a YAML script,
  2. uses the bridge to drive the clipper through navigation/click/wait steps,
  3. simultaneously runs screencapture -v -l<chromeWindowId> (macOS window-level capture) so the recording covers the whole Chrome window including the clipper’s sidepanel UI,
  4. optionally hands the raw mp4 to video-use for post-processing.

3d. Catalog management (sidepanel needs to list things)

Two more MCP tools so the sidepanel and any agent can ask “what’s available”:

amem_list_skills() -> [{ id, name, description, kind: "auto"|"run", state }]
amem_invoke_skill(id: string, params?: object) -> { result }

For v0.1 these read from a hardcoded Vec<SkillCard> in the librarian. There is no manifest. There is no sandbox. amem_invoke_skill("cws-demo") literally calls record_demo(builtin_scripts::CWS_DEMO). The point is the tool surface: the moment we have a skill engine in v0.5+, these two tools’ signatures don’t change — only their implementation does. The agent contract is forward-compatible.

4. MCP tool inventory for v0.1

Total: 6 new tools ship in v0.1 (plus the four existing ones from Day 1: amem_capture, amem_compile, amem_cite, amem_recall).

ToolBacking functionNotes
chrome_navigatetools::chrome::navigateBridge pass-through
chrome_clicktools::chrome::clickBridge pass-through
chrome_extracttools::chrome::extractBridge pass-through
amem_capture_urladapters::dispatch + captureHardcoded site match
amem_invoke_skillskills::runHardcoded skill table
amem_list_skillsskills::listHardcoded skill table

amem_record_demo is exposed as a kind:"run" skill via amem_invoke_skill("cws-demo"), not as a separate top-level tool — keeping the agent-facing surface tighter and letting RFC-003 own the contract.

5. Why NOT skill engine (yet)

Five reasons, in priority order:

  1. Rule of three. Of the three v0.1 features, only one (site adapters) has even three instances. The other two are one-of-one. Abstracting now means designing for assumptions we have not yet earned.
  2. One operator. A skill engine optimises for external authors shipping skills. We have zero external authors. The first user benefitting from sandboxing-vs-trust would be us — and we trust our own code more than we trust a sandbox we just wrote.
  3. 13 vs 7.5 days. A real skill engine needs: manifest schema with versioning, loader with capability gating, sandboxed JS runtime (boa/quickjs) with controlled host bindings, error surfaces, registry, discovery, version pinning, update flow. None of this work helps the v0.1 user-facing demo.
  4. CWS deadline lives in this window. Chrome Web Store review can take 2–10 days. Submit narrow + working > submit broad + half-built. The 5.5-day delta is bigger than the review buffer.
  5. Forward-compatible API. amem_list_skills / amem_invoke_skill already shape the surface a future engine will use. We are not painting ourselves into a corner; we are shipping the same MCP signatures that v0.5+ will reuse.

6. When to revisit — the rule of three

Promote to a real engine when any of the following triggers fire:

TriggerWhat it meansWhat to revisit
5+ site adaptersThe match on url.host_str() has grown to 5+ arms with similar shapeExtract a SiteAdapter trait + manifest table
3+ recording scripts in productionYAML scripts have proven pattern: navigate, click, capture, repeatPromote YAML schema to versioned spec; consider script-side templating
1+ external user requestSomeone outside yiidtw/ asks “can I write my own skill”Engine becomes a user-facing feature, not internal cleanup
Custom skill in v0.2 plan firms upRFC-002 promises “Custom skills coming v0.2” — the moment that shipsEngine is the implementation

Until any trigger fires, the function-based approach is the correct endpoint, not a placeholder.

Privacy

Inherits the archived-RFC-001 posture unchanged:

  • All capture data lives in ~/.amem/, never uploaded.
  • Bridge is loopback only, token-authed, origin-checked.
  • The new MCP tools (chrome_*) execute in the user’s real Chrome under the user’s existing permissions — they cannot access tabs the user is not already authenticated to.
  • Recording (RFC-003) uses macOS window-level screencapture targeting the Chrome window’s window ID; desktop and other apps are not in frame.

One new consideration: chrome_extract returns DOM content to the librarian, which may surface to an agent over MCP. This is identical in sensitivity to today’s amem_capture(url), which already fetches and parses page content. Document the equivalence in the user guide.

Failure modes

ModeCauseMitigation
Bridge unreachable when agent calls chrome_*Librarian not running, or extension not connectedMCP tool returns structured error { code: "BRIDGE_UNAVAILABLE", install_hint }; agent surfaces install CTA
Selector not foundPage changed, login required, racechrome_click / chrome_extract return { ok: false, reason } rather than throw; agent retries with chrome_wait
Adapter doesn’t match URLURL outside hardcoded listamem_capture_url returns { status: "unsupported" }; agent can fall back to the generic amem_capture (existing Day 1 tool)
Skill ID typoAgent calls amem_invoke_skill("does-not-exist")Structured error with available_ids list (read from amem_list_skills)
Concurrent recording + chrome opsRecording and other MCP tools racing for the bridgeRecording acquires a recording-mode lock; concurrent chrome_* calls return { code: "RECORDING_IN_PROGRESS" }
Multi-tab / multi-window ambiguity“Active tab” is ambiguous on multi-window setupsDefault to focused-window’s active tab; expose tabId param later if it bites
Skill state drift between sidepanel and librarianUser toggles in sidepanel while CLI also togglesLibrarian is single source of truth (config.toml); sidepanel reads on every render via amem_list_skills

Concrete work

In rough order of dependency. Estimates are pessimistic-realistic.

  1. (amem-librarian) Bridge envelope + token + origin allow-list — ~1d
  2. (amem-librarian) MCP wrappers for chrome_navigate / _click / _extract / _wait / _screenshot — ~0.5d
  3. (amem-clipper) Bridge client + handlers for above verbs via chrome.scripting.executeScript — ~1.5d
  4. (amem-librarian) Adapters: arxiv.rs, github.rs, hackernews.rs
    • dispatch table + amem_capture_url MCP tool — ~2d
  5. (amem-librarian) Hardcoded skill catalog (SkillCard, the three v0.1 entries) + amem_list_skills / amem_invoke_skill MCP tools — ~0.5d
  6. (amem-clipper) Sidepanel “Skills” tab rendering catalog (full detail in RFC-002) — ~1d
  7. (amem-librarian) record_demo() skeleton stub returning not implemented — full implementation lives in RFC-003 — ~0.5d
  8. (docs.amem.sh) guide/skills.md documenting the v0.1 hardcoded skills + the v0.2 forward-compat note — ~0.5d

Total: ~7.5d. Compare to ~13d for an equivalent skill-engine v0.1.

Rejected alternatives

  • Ship a real skill engine in v0.1. Costs 5.5d we don’t have, optimises for a user we don’t yet have. See §5.
  • Ship skill engine manifest only, run hardcoded for v0.1. Tempting middle ground, but commits us to a manifest schema before we know what fields second-and-third skills will need. Premature schema lock.
  • Embed skills as JS in the clipper extension. Moves heavy work into the extension (LLM calls, file I/O), which violates the clipper-is-sensor positioning and runs into MV3 storage / CSP limits. The librarian must stay the brain.
  • Use playwright / puppeteer / Chrome DevTools Protocol for §3a. Loses the user’s Chrome profile (login state, passkeys, 2FA). The bridge-into-real-Chrome path is non-negotiable.
  • Skip the catalog UI, ship MCP tools only. Loses the CWS positioning (“first agent-callable Chrome skills catalog”). The catalog is the user-facing story; without it we’re just another Chrome extension.

Open questions

  • Multi-tab targeting. Should chrome_* tools default to the active tab in the focused window (proposed) or accept an explicit tabId? Soft preference: default-active for v0.1, expose tabId when the second use case asks.
  • Adapter timeouts. What does amem_capture_url do if arxiv is slow / down? Soft preference: 30s timeout, returns { status: "captured_partial", reason: "fetch_timeout" } so the agent can retry later.
  • MCP tool surface stability. The 6 new tool names are committed. We will add tools post-v0.1 but not rename or remove. The surface is the contract.
  • Should amem_list_skills include kind:"hidden" for skills used internally (e.g. by other tools) but not surfaced in the sidepanel? Open. v0.1 ships without; revisit when an internal skill emerges.

Roll-out

  • Day 1 (today): this RFC + RFC-002 + RFC-003 land. SUMMARY.md updated.
  • Day 2: bridge envelope + chrome_* MCP tools + clipper handlers.
  • Day 3: site adapters land. amem_capture_url ships.
  • Day 4: skill catalog + sidepanel UI (RFC-002 implementation).
  • Day 5: recording skeleton lands (RFC-003 implementation).
  • Day 6: end-to-end CWS demo recorded by amem itself, dogfooded.
  • Day 7: CWS submission. Listing copy claims “first agent-callable Chrome skills catalog” honestly — three skills shipped, custom-skills disclosure visible.
  • Post-v0.1: track the rule-of-three triggers in §6. When any fires, open RFC-00X for the skill engine.

RFC-002 — amem Clipper skills catalog UI (v0.1)

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-05-09
  • Related: RFC-001 (function-based v0.1), RFC-003 (recording orchestration), archived RFC-001 (bridge-first), SPEC.md § Mental model

TL;DR

The amem Clipper sidepanel grows a second tab — Skills — that renders the v0.1 catalog as if it were an installable marketplace. v0.1 ships three hardcoded skills behind that UI: Auto-capture arxiv, LinkedIn inbox glance, CWS demo recording. The skill cards look identical to what v0.2’s actually-installable skills will look like, with one honest difference: a “Custom skills coming v0.2” disclosure underneath the catalog. CWS positioning leans on this UI: the listing claims “the first agent-callable Chrome skills catalog,” which is true — there is a catalog, each entry is agent-callable via MCP, and the catalog will accept custom entries in v0.2. We are shipping the shape of the product before the generality.

Motivation

RFC-001 commits to function-based v0.1 — hardcoded MCP tools in the librarian, no skill engine. That decision optimises engineering throughput to the CWS deadline. It does not, by itself, give us a CWS listing or a narrative.

The sidepanel UI is where the strategic positioning lives:

  • Without a catalog UI, we are submitting “another Chrome extension that captures pages and integrates with an MCP server.” Crowded category; forgettable listing.
  • With a catalog UI, the listing reads: “amem Clipper turns your browser into an agent-callable skills runtime. Three skills shipped — arxiv auto-capture, LinkedIn inbox glance, demo recording. Custom skills coming v0.2.” This is the differentiator. Nobody else is shipping Chrome skills the agent can list and invoke over MCP.

The catalog UI also pulls future weight: the v0.2 install flow is identical in structure to the v0.1 hardcoded path (skill card → toggle → bridge command). Users who learn the v0.1 model carry the mental model forward unchanged.

Proposal

1. Sidepanel tab strip

The sidepanel grows from a single capture surface to a two-tab strip:

┌────────────────────────────────────┐
│  amem Clipper                  ⚙   │
├────────────────────────────────────┤
│  [ Captures ]   Skills              │   ← tabs, [bracketed] = active
├────────────────────────────────────┤
│                                    │
│   (tab content)                    │
│                                    │
└────────────────────────────────────┘
  • Captures — the existing capture log (recent items, filters, search). Default tab on first launch.
  • Skills — the new catalog tab. The shape this RFC is about.

State is stored in chrome.storage.local; users land on whichever tab they last closed.

2. Skill card schema (UI-level, not manifest)

Each skill renders as one card. The schema below is the render contract between the librarian (which owns the catalog truth) and the sidepanel (which renders it). It is not a manifest on disk in v0.1; it is the shape returned by amem_list_skills over the bridge.

type SkillCard = {
  id:          string;          // stable, e.g. "arxiv-autocapture"
  name:        string;          // "Auto-capture arxiv"
  description: string;          // one-liner, ~80 chars
  icon:        string;          // emoji or built-in icon name
  kind:        "auto" | "run";  // toggle vs button
  state:       SkillState;
  badges?:     string[];        // e.g. ["builtin", "v0.1"]
};

type SkillState =
  | { kind: "auto"; enabled: boolean }                    // for auto skills
  | { kind: "run";  busy: boolean; last_run?: string }    // for run skills
  | { kind: "error"; message: string };

Card layout:

┌─────────────────────────────────────────────────────┐
│  📄  Auto-capture arxiv                  [ ●─── ]   │
│      Captures every arxiv abstract you open.        │
│      builtin · v0.1                                  │
└─────────────────────────────────────────────────────┘
  • Top-right control depends on kind:
    • kind:"auto" → toggle switch, bound to the enabled boolean
    • kind:"run" → “Run” button, disabled while busy=true, last-run timestamp under the button
    • kind:"error" → red text + “Retry” button
  • Badges render as small tags at the bottom of the card.
  • Cards are not reorderable in v0.1 (catalog order is hardcoded). v0.2 may add user pinning.

3. The three v0.1 skills

3a. Auto-capture arxiv (id: "arxiv-autocapture", kind: "auto")

Toggle-on means: any time the user navigates to an arxiv URL matching arxiv.org/abs/* or arxiv.org/pdf/*, the clipper sends an auto_capture event over the bridge with the URL; the librarian dispatches to the arxiv adapter (RFC-001 §3b) and ingests.

Default: off. Auto-capturing every arxiv abstract is opinionated; users opt in.

UI behaviour:

  • Toggle ON → small toast “Auto-capture armed for arxiv.org”
  • On a successful capture → unobtrusive notification dot on the Clipper toolbar icon for ~5s
  • On a duplicate (already captured) → silent (don’t spam)
  • On error → clipper toolbar icon shows red dot, click for details

Implementation note: the toggle state lives in chrome.storage.local for fast content-script reads; the librarian’s config.toml mirrors it as the authoritative copy. Sidepanel reads truth from amem_list_skills on every render; toggle writes go through amem_invoke_skill("arxiv- autocapture", { enabled: true }).

3b. LinkedIn inbox glance (id: "linkedin-inbox", kind: "run")

User clicks Run → librarian uses chrome_navigate to open linkedin.com/messaging/, then chrome_extract to read the visible inbox previews, then renders a digest in the sidepanel: who messaged, unread count, first line of each thread.

This skill exists for two reasons:

  1. Demo value. It demonstrates that the agent can operate the user’s already-logged-in Chrome — the v0.1 differentiator. Without a visible flagship for that capability, the CWS reviewer has nothing concrete to evaluate.
  2. Honest utility. The operator actually uses this. It is not a throwaway skill written for the demo.

UI behaviour:

  • Click Run → button disables, spinner, “Reading inbox…”
  • Success → digest renders inline below the card; expandable
  • Error (e.g. logged out) → “LinkedIn requires sign-in. Please open linkedin.com and sign in, then retry.”

3c. CWS demo recording (id: "cws-demo", kind: "run")

User (or the agent) clicks Run → librarian invokes record_demo() against the built-in CWS demo script (RFC-003). The sidepanel shows a recording state: live elapsed time, current step, “Stop” button.

This is the meta-skill: amem records its own CWS demo by driving its own extension. The recording captures the entire Chrome window including the sidepanel UI itself, which is why we need window-level macOS screencapture rather than chrome.tabCapture (full reasoning in RFC-003 §5).

UI behaviour:

  • Click Run → card expands with live status (current step, elapsed time)
  • “Stop” cancels gracefully, finalises whatever was captured
  • On finish → card shows amem://recording/<uuid> link + path to mp4
  • The recorded mp4 lands in ~/.amem/recordings/<uuid>.mp4

4. Auto-fire mechanism for URL-matched skills

The auto-capture flow lives in two places:

  1. Content script (clipper) — observes URL changes via chrome.webNavigation.onCommitted. Pattern matches against the active set of kind:"auto" skills with enabled:true. Patterns live in chrome.storage.local, synced from the librarian via the bridge on connect.
  2. Background service worker (clipper) — receives match events, forwards to bridge as auto_capture.

Why store patterns in chrome.storage.local rather than asking the bridge per navigation:

  • Latency. URL change → bridge round-trip → match would add 50–200ms; unacceptable for ambient capture.
  • Resilience. If the bridge briefly disconnects, the toggle behaviour shouldn’t change; the next reconnection re-syncs state.

Sync protocol on bridge connect:

clipper → librarian: { action: "skills_subscribe" }
librarian → clipper: { action: "skills_state",
                       params: { skills: [SkillCard, ...] } }
(thereafter, librarian pushes "skills_state" on any change)

Drift detection: every amem_list_skills MCP call also pushes the current state to the clipper, so any out-of-band CLI toggle propagates.

5. Bridge command flow for Run-skills

User clicks Run on a kind:"run" skill:

sidepanel UI                 background.js               librarian
     │  user clicks Run            │                          │
     ├─ "invoke_skill",───────────▶│                          │
     │   id: "linkedin-inbox" }    │                          │
     │                             ├─ WS send ───────────────▶│
     │                             │                          │
     │                             │           amem_invoke_skill("linkedin-inbox")
     │                             │                          │
     │                             │                          │ ┌─────────────┐
     │                             │                          │ │ runs the    │
     │                             │                          │ │ skill →     │
     │                             │                          │ │ chrome_*    │
     │                             │                          │ │ over bridge │
     │                             │                          │ └─────────────┘
     │                             │           ◀─ chrome_navigate /
     │                             │              chrome_extract pulses
     │                             │                          │
     │                             │◀─ "skills_state",────────│
     │                             │   { busy: false,         │
     │                             │     last_run: ... }      │
     │  "skills_state" forwarded   │                          │
     │◀────────────────────────────┤                          │
     │  re-render card             │                          │

The skill execution itself is just an MCP tool call (amem_invoke_skill). The librarian is the orchestrator; the clipper is the executor for DOM side-effects. The sidepanel only kicks the kickoff and re-renders state.

6. CWS listing positioning

Listing copy (proposed for store page):

amem Clipper The first agent-callable Chrome skills catalog.

amem Clipper turns your browser into a runtime for Chrome skills your AI agent can list and invoke over MCP. Three skills ship today:

  • Auto-capture arxiv — every paper you open lands in your local wiki
  • LinkedIn inbox glance — your agent can read your inbox without you opening the tab
  • CWS demo recording — amem records its own demos by driving its own extension

Custom skills coming v0.2.

Pairs with amem-librarian, the local Rust binary that runs on your machine. Your data stays on your disk.

Three claims worth defending:

  1. “First agent-callable Chrome skills catalog.” True if “catalog” means “a list of skills the agent can enumerate and invoke.” We have amem_list_skills and amem_invoke_skill over MCP; that is the catalog interface. v0.2 adds the user-installs-their-own dimension. The claim is honest with the v0.2 disclosure intact.
  2. “Three skills ship today.” True; all three are real and tested.
  3. “Your data stays on your disk.” True; storage layout in ~/.amem/ is unchanged; bridge is loopback only.

7. Honest disclosure: “Custom skills coming v0.2”

Below the catalog, a fixed footer:

┌─────────────────────────────────────────────────────┐
│  ✨  Custom skills coming v0.2                       │
│      Bring your own scripts. ==AmemSkill== headers   │
│      will install via drag-and-drop or URL paste.    │
│      Read the v0.2 design intent →                   │
└─────────────────────────────────────────────────────┘

The “Read the v0.2 design intent” link goes to docs.amem.sh/skills/v0.2, which renders the next section (§8) for transparency.

8. v0.2 design intent — Tampermonkey-style headers

When v0.1 hits a rule-of-three trigger (per RFC-001 §6) we ship a real skill engine. Sketch of the v0.2 user-facing format:

// ==AmemSkill==
// @id           youtube-autocapture
// @name         Auto-capture YouTube
// @description  Saves every YouTube video you watch to your wiki
// @kind         auto
// @match        https://www.youtube.com/watch*
// @capability   bridge:chrome_extract
// @capability   librarian:capture
// @version      1
// ==/AmemSkill==

export async function onMatch({ url, ctx }) {
  const title = await ctx.chrome.extract("h1.ytd-watch-metadata", "text");
  await ctx.librarian.capture(url, { title });
}

Key design choices, locked-in for forward-compat:

  • Header format borrowed from Tampermonkey/Greasemonkey. Familiar to anyone who’s written userscripts; no new format to learn.
  • Capabilities are explicit. Every skill declares which librarian and bridge verbs it touches. Unknown capabilities = install-time rejection.
  • Two execution targets. A skill is either DOM-side (runs in clipper) or librarian-side (runs in Rust via JS sandbox). The header decides.
  • Install paths. Drag-and-drop a .amemskill.js onto the sidepanel, paste a URL pointing to one, or import from a community registry once one exists.
  • Same MCP surface. amem_list_skills and amem_invoke_skill keep their v0.1 signatures. Custom skills appear in the same catalog, flagged with badges: ["custom"].

This section is design intent, not commitment. Numbers can change. What is locked-in is the v0.1 MCP surface; v0.2 will not re-shape it.

Privacy

  • The catalog UI itself sends no telemetry. Skill card render data comes exclusively from the loopback bridge.
  • Auto-fire patterns (URL globs) live in chrome.storage.local per-user, per-profile. They are not synced with Chrome Sync (we explicitly opt out by not declaring storage.sync permission).
  • Run-skill invocations log to the librarian only (~/.amem/skills.log, one line per invoke with timestamp, skill id, outcome). Off by default; enable with [skills] log = true in config.toml.
  • The “LinkedIn inbox glance” skill reads DOM content from the user’s own logged-in tab. It does not exfiltrate anywhere except the librarian’s local storage. The inbox digest is not auto-captured to the wiki — it renders inline only.

Failure modes

ModeCauseMitigation
Bridge disconnects mid-renderLibrarian crashed or restartedCards render in kind:"error" state with “Reconnect” affordance; on reconnect, skills_state re-syncs and cards refresh
Auto-skill toggle driftUser toggles in CLI (amem skills enable …) while sidepanel is openSidepanel listens for skills_state push; re-renders
Run-skill hangsSkill awaits a selector that never appearsRun buttons have a 60s soft timeout; user sees “Taking longer than usual…” + “Stop”
Pattern match fires on wrong URLGlob too loosev0.1 patterns are tightly scoped (e.g. arxiv.org/abs/* not *arxiv*); we err narrow
Sidepanel renders emptyamem_list_skills returned an empty arrayShow “Skills service unavailable. Is amem-librarian running?” with install link
LinkedIn UI changesLinkedIn redesigns inboxSkill returns kind:"error" with selector-not-found; we patch the selector and ship a librarian update — no extension update needed (selectors are server-side)
Recording skill UI conflictUser clicks Run on cws-demo while another chrome_* call in flightRecording acquires librarian-wide lock; concurrent invocations get RECORDING_IN_PROGRESS (RFC-001 §failure modes)

Concrete work

In rough order:

  1. (amem-clipper) Sidepanel tab strip + Captures / Skills shell — ~0.5d
  2. (amem-clipper) Skill card component (auto + run + error variants) — ~1d
  3. (amem-clipper) Bridge subscribe / skills_state handler + drift reconciliation — ~0.5d
  4. (amem-librarian) Hardcoded SKILLS table + amem_list_skills / amem_invoke_skill MCP tools (lives in RFC-001 §3d but the wiring to the catalog UI is here) — ~0.5d
  5. (amem-clipper) Auto-fire content-script — URL pattern matcher, chrome.storage.local cache, auto_capture emitter — ~1d
  6. (amem-clipper) “Custom skills coming v0.2” footer + link — ~0.25d
  7. (docs.amem.sh) skills/v0.2.md page documenting the design intent header — ~0.5d
  8. (amem-clipper) CWS listing assets: screenshots of the catalog, listing copy per §6, demo gif (recorded by the cws-demo skill itself) — ~0.5d

Total: ~4.25d (overlaps with RFC-001 §concrete-work step 6, which budgets 1d for the same UI work — net new is ~3.25d.)

Rejected alternatives

  • Ship without a “Skills” tab; make capture-only the v0.1 surface. Loses the CWS positioning. We are submitting at the same time as a hundred other capture extensions; we need the catalog story to stand out.
  • Render skills as a flat list of MCP tools. Technically accurate, user-hostile. “MCP tool” is jargon; “skill” maps to mental models from Tampermonkey, App Store, Raycast extensions, etc.
  • Make custom skills installable in v0.1 by accepting hardcoded patches. Doable but every patch is a librarian release; no actual install flow; misleading. Better to be honest and ship the disclosure.
  • Hide the “v0.2 coming soon” disclosure. Considered for marketing cleanliness; rejected for trust. Every user opening the sidepanel will immediately wonder “can I write my own?”; pretending otherwise breeds cynicism.
  • Use manifest_version / install registry now to keep “skills” honest. Equivalent to building a skill engine; rejected per RFC-001 §5.

Open questions

  • Card density. Three skills look fine; do we need pagination / search at 10+? Soft preference: defer to v0.2 when skill count is user-driven.
  • Skill icons. Use emoji per skill (📄, 💼, 🎬) or commission custom SVGs? Soft preference: emoji for v0.1 (zero design cost), SVGs if a CWS reviewer flags emoji as low-effort.
  • Sidepanel width on narrow screens. The skill cards assume ~360px; Chrome sidepanel can be narrower. Verify on 1280-wide laptop.
  • Should amem_list_skills filter by current tab URL? I.e. only show “LinkedIn inbox” when the user is on linkedin.com? Considered; rejected for v0.1 — the catalog is supposed to feel like a marketplace shelf, not a context menu.
  • Run-skill audit. Do kind:"run" invocations need a confirmation dialog (“Run LinkedIn inbox glance?”) or fire immediately? Soft preference: immediate for v0.1, confirmation if user feedback says surprising.

Roll-out

  • v0.1 (this round): three hardcoded skills, sidepanel UI shipped, CWS submission. The catalog is real but not extensible.
  • v0.1.x patches: site adapters and skill cards iterate based on actual usage. Each adapter or card is a librarian release; the extension changes only when the card schema changes (rare).
  • v0.2: skill engine ships per the design-intent sketch (§8). Custom skills install via drag-and-drop. Catalog UI gains an “Install” affordance; existing cards keep working unchanged.
  • v0.3+: community skills registry, skill versioning, signed skills. Out of scope here; track in a future RFC.

RFC-003 — Recording skill orchestration (v0.1)

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-05-09
  • Related: RFC-001 (function-based v0.1), RFC-002 (Clipper skills catalog), SPEC.md § “amem is its own best demo,” archived guide/self-recording.md (superseded by this RFC)

TL;DR

Replace the existing chrome.tabCapture self-recording skeleton with a scripted, librarian-driven recording pipeline. A YAML script declares a sequence of Chrome operations; the librarian drives Chrome through the bridge (which executes them in amem Clipper) while simultaneously running macOS window-level screencapture -v -l<windowID> against the Chrome window. Output: an mp4 at ~/.amem/recordings/<uuid>.mp4 plus an amem://recording/<uuid> URI. Optional handoff to video-use for post-processing (transcribe, cut filler, captions).

The whole thing surfaces as one MCP tool — amem_record_demo — and one skill card (cws-demo, RFC-002 §3c). v0.1 ships one built-in script (the CWS demo). Two more (feature-update demo, tutorial template) follow when the third script is requested.

Why librarian-side and not extension-side: chrome.tabCapture cannot record the sidepanel UI itself (sidepanels are out-of-tab surfaces), which is exactly the surface our CWS demo needs to show. Window-level macOS capture fixes that and only that.

Motivation

amem inherits crossmem’s principle — “a product that is its own best demo” (SPEC.md § README, also docs/src/guide/self-recording.md). We record our own marketing video by driving our own extension. This is real load-bearing infrastructure, not a marketing gimmick: every CWS re-submission, every feature announcement, every tutorial benefits.

The Day 1 skeleton (docs/src/guide/self-recording.md) used chrome.tabCapture from an offscreen document. Three problems made us abandon it for v0.1:

  1. chrome.tabCapture cannot capture sidepanels. It captures the tab’s rendered area; the sidepanel is outside that surface. The single most important thing our demo must show — the skills catalog sidepanel — is invisible to tabCapture. We could render the catalog inside a tab as a workaround, but then we are demoing a fake.
  2. No driving model. The skeleton assumed “an orchestrator” sends start_recording and “drives the extension UI.” There is no orchestrator design — just a hand-wave. v0.1 must ship a real one.
  3. MV3 service worker lifecycle is hostile to long recordings. Service workers can be evicted under memory pressure; offscreen documents help but add coordination complexity. Doing this in the librarian (a long-running Rust process) is simpler and more reliable.

Moving the recorder to the librarian means it can use macOS-native screencapture (window-level, captures everything in the Chrome window including all of Chrome’s chrome) and orchestrate the demo via the bridge. Same component split as RFC-001: librarian = brain, clipper = sensor.

Proposal

1. YAML script schema

Scripts live in two places:

  • crates/amem-librarian/builtin_scripts/*.yaml — built-in templates (CWS demo, etc.), compiled into the binary
  • ~/.amem/recordings/scripts/*.yaml — user scripts (v0.1 supports these but does not have a UI for managing them; CLI only)

Schema:

# ~/.amem/recordings/scripts/cws-demo.yaml
name: "CWS demo (v0.1)"
duration_target: 45s        # advisory; total run time including waits
window:
  app: "Google Chrome"
  match: "active"           # or { title_regex: "..." } for multi-window
record_cursor: true         # passed to screencapture
output:
  format: mp4
  resolution: source        # or 1080p, 720p (downscale via ffmpeg)
  filename: "cws-demo-{ts}.mp4"

steps:
  - id: open-arxiv
    type: navigate
    url: "https://arxiv.org/abs/1706.03762"
    wait_for: "h1.title"

  - id: capture-arxiv
    type: caption
    text: "amem captures any arxiv paper you open"
    duration: 3s

  - id: open-skills
    type: click_selector
    selector: "[data-amem-tab='skills']"
    wait_after: 500ms

  - id: show-skills
    type: caption
    text: "Three skills shipped — and your agent can call any of them"
    duration: 4s

  - id: click-linkedin
    type: click_selector
    selector: "[data-amem-skill='linkedin-inbox'] button.run"
    wait_for: "[data-amem-skill='linkedin-inbox'] .digest"

  - id: settle
    type: wait
    duration: 2s

  - id: terminal-finish
    type: terminal
    cmd: "amem recall --json 'attention is all you need'"
    show: "stdout"
    duration: 4s

2. Step types

Each step is one variant; steps execute sequentially (no v0.1 parallelism).

TypePurposeDriver
navigateLoad URL in active tabchrome_navigate over bridge
click_selectorClick an elementchrome_click
waitPause for N ms / N slibrarian sleep
wait_for (also a field on other steps)Block until selector / URL appearschrome_wait
extract_domRead DOM, save to script vars (for later assertions/captions)chrome_extract
captionRender an overlay caption for N secondslibrarian draws into a borderless overlay window
terminalRun a binary command, optionally render its stdout in an overlaylibrarian process + overlay

The terminal step is interesting: it lets a recording show CLI usage alongside the browser. The librarian opens a small floating window (via its own UI process, not the Chrome window) that renders the command and its stdout in monospace; screencapture picks up the overlay because we target the Chrome window but composite the overlay on top of it before each frame. Implementation detail: macOS lets us position a borderless NSWindow above the target Chrome window and screencapture -l<chromeWinId> will include it (Quartz compositing, same as visible-tab capture).

If the overlay approach proves unreliable across macOS versions, fallback v0.1: split the recording — capture Chrome window for browser steps, capture full screen for terminal steps, stitch in post via video-use. Decision deferred until we hit a Sonoma/Sequoia regression in testing.

caption steps render a similar borderless overlay at the bottom-center of the Chrome window with a translucent black bar.

3. macOS implementation — window-level screencapture

Core command:

screencapture -v -l <chromeWindowId> -V <duration_seconds> output.mov
  • -v — start recording immediately (no UI)
  • -l <windowId> — target a specific window. We get this from CGWindowListCopyWindowInfo(.optionOnScreenOnly) filtered to bundle identifier com.google.Chrome (or Brave / Arc / Edge variants — we hardcode the major Chromium IDs).
  • -V <seconds> — fixed duration. v0.1 sets this to duration_target + 10s buffer; we kill the process early if all steps finish before the timer.
  • Output .mov; converted to .mp4 via ffmpeg -i in.mov -c:v libx264 -crf 20 out.mp4 post-recording.

Why window-level not full-screen:

  • Privacy: full-screen captures the user’s desktop background, dock, notifications, other apps. Window-level captures only the Chrome window’s pixel rect.
  • Aesthetics: window-level avoids us cropping the recording in post.
  • Privacy posture matches the rest of amem (“data stays local; we don’t capture what we don’t need”).

Permission flow: macOS requires Screen Recording permission. On first invoke, the librarian prompts the user via TCC. If denied, we return { code: "SCREEN_RECORDING_DENIED", how_to_fix } so the agent can surface the system-settings link.

Linux / Windows: out of scope for v0.1. The recording skill is gated to macOS in v0.1 (#[cfg(target_os = "macos")]); on other OSes the skill renders as kind:"error" with “Recording requires macOS in v0.1.” Linux pipeline (probably wf-recorder for Wayland + scrot/ffmpeg-x11grab for X11) is a follow-up RFC.

4. video-use integration (optional post-processor)

The raw mp4 from §3 is usable as-is — but for marketing-grade output we want transcription, filler-word cutting, captions, and consistent encoding. video-use is the operator’s existing tool for that; integrating is a one-liner:

#![allow(unused)]
fn main() {
// crates/amem-librarian/src/skills/record_demo.rs
async fn record_demo(script: &Script) -> Result<RecordingOutput> {
    let raw_mov = run_screencapture(&script.window, script.duration_target).await?;
    let raw_mp4 = ffmpeg_convert(&raw_mov).await?;

    let final_mp4 = if script.post_process.unwrap_or(false) {
        video_use::process(&raw_mp4, &script.post_options).await?
    } else {
        raw_mp4
    };

    Ok(RecordingOutput {
        amem_uri: format!("amem://recording/{}", uuid),
        mp4_path: final_mp4,
    })
}
}

post_process: false is the v0.1 default (raw recording). Setting post_process: true in the YAML opts in. video-use is treated as an optional dependency: if it isn’t installed, the field is ignored with a warning.

The script’s post_options mirror video-use’s CLI flags one-to-one (so the integration stays a thin shim, not a redesign).

5. amem_record_demo MCP tool

Exposed indirectly through amem_invoke_skill("cws-demo") per RFC-001 §3d, but also as a direct tool for ad-hoc agent use:

amem_record_demo(
  script_path?:   string,        // path to YAML
  script_inline?: string,        // YAML literal (mutually exclusive with script_path)
  output_dir?:    string,        // defaults to ~/.amem/recordings/
)
  -> {
    amem_uri:    "amem://recording/<uuid>",
    mp4_path:    "/Users/.../<uuid>.mp4",
    duration_s:  number,
    steps_run:   number,
  }

Errors:

CodeMeaning
SCREEN_RECORDING_DENIEDmacOS TCC denied screen recording
BRIDGE_UNAVAILABLECannot reach amem Clipper to drive Chrome
WINDOW_NOT_FOUNDNo Chrome window matched script.window
STEP_FAILEDSome step failed; partial recording saved with partial: true
RECORDING_IN_PROGRESSAnother recording is active; refuse to start a second
UNSUPPORTED_OSNot macOS in v0.1

The tool is blocking: it returns when the recording finishes (typically 30–90s for v0.1 scripts). MCP clients already render “long-running tool call” affordances.

6. Use cases

The same pipeline serves three concrete needs, each justifying ship in v0.1:

6a. CWS promo video (the SPEC’s “own best demo” principle)

The CWS listing needs a 30–60 second demo. We script it (§1), the librarian records it, video-use polishes it, we upload. When v0.1.x patches change the UI, we re-run the same script. The demo never goes stale.

6b. Feature update demos (post-v0.1)

When v0.2 ships custom-skill installation, we want a demo of that flow. Same pipeline, new YAML script. v0.2 itself is a script.

6c. Tutorials

docs.amem.sh user guide pages can embed inline mp4s recorded from canonical YAML scripts in crates/amem-librarian/builtin_scripts/. When the UI changes, regenerate.

7. Why the librarian records, not the clipper

The cleanest framing of this RFC’s central decision:

OptionRecords sidepanel UI?Privacy?Lifecycle?Cross-OS path?
A. chrome.tabCapture (extension)❌ No (sidepanel is out-of-tab)✅ tab-only⚠️ MV3 service worker eviction risk✅ identical everywhere
B. getDisplayMedia (extension)✅ User picks the window✅ user opt-in per recording⚠️ same MV3 risks✅ uniform
C. macOS window screencapture (librarian)✅ entire Chrome window✅ window-only, no desktop leak✅ Rust process is long-lived❌ macOS-only v0.1

Option B is the second-best choice; we considered it seriously. We rejected it for v0.1 because:

  1. getDisplayMedia requires a user picker dialog every recording — incompatible with a “click Run, get a recording” agent-driven flow.
  2. The MV3 service worker would still need to coordinate with the librarian for the script driver, doubling the moving parts.
  3. We want the librarian to own the timeline anyway — it already owns the script, the bridge, the storage, the post-process step. Adding the recording itself keeps responsibility in one place.

Option C costs us OS portability in v0.1 — accepted tradeoff. Linux / Windows recording lands when there is a non-macOS user.

Privacy

  • Window-level screencapture: only the Chrome window’s pixels are captured. Desktop, dock, menu bar, other apps, notifications — all outside the frame.
  • The recorded mp4 lives in ~/.amem/recordings/<uuid>.mp4. Never uploaded by the librarian. The user can amem upload (future) or manually drag-drop to a CWS listing.
  • macOS Screen Recording permission is requested on first invoke; the user can revoke at any time in System Settings.
  • Caption / terminal overlays render the script’s literal text; no enrichment, no LLM rewriting (v0.1).
  • video-use (when used) is a local binary; transcription runs locally via whisper. No cloud calls in the default config.
  • The script itself is recorded into the mp4 metadata’s comment field (so anyone with the mp4 can see what was scripted). The user can disable with metadata: false in the YAML.

Failure modes

ModeCauseMitigation
Screen Recording permission deniedFirst-run user hasn’t granted TCCTool returns SCREEN_RECORDING_DENIED with open System Settings → Privacy → Screen Recording instruction; sidepanel card shows the same
Chrome window not foundUser has no Chrome window open, or app is Brave/ArcYAML window.app allow-list extended to common Chromium variants; tool errors with WINDOW_NOT_FOUND listing detected windows
Step times outSelector never appearsStep times out at wait_for budget (default 30s); recording stops; partial: true flag in result so user knows
Bridge disconnects mid-recordingLibrarian-clipper WS dropsRecording continues (it’s screencapture, not bridge-driven pixels), but subsequent steps can’t drive Chrome; we abort gracefully and save what we have
screencapture produces 0-byte fileKnown macOS bug in some Sonoma builds when target window is fully occludedPre-check: bring Chrome to front + verify visibility before starting; document the Apple bug in docs/troubleshooting.md
Multiple recordings requested concurrentlyTwo agents both call amem_record_demoLibrarian-wide recording lock; second call gets RECORDING_IN_PROGRESS
Output mp4 hugeHigh-resolution display, long demoDefault resolution: source for v0.1; users can opt to 1080p / 720p in YAML; ffmpeg downscale handled in §4
Terminal overlay flickers / drops framesmacOS compositor under loadDocumented; user can switch to “split capture + post-stitch” mode
Wrong Chrome window picked (multi-window)User has 3 Chrome windows; we pick wrongYAML window.match accepts title_regex; default is “active window” which uses the focused one at recording start

Concrete work

In rough order:

  1. (amem-librarian) YAML script parser + schema validator — ~0.5d
  2. (amem-librarian) macOS window-id resolver via Quartz CGWindowListCopyWindowInfo (FFI through core-graphics crate) — ~0.5d
  3. (amem-librarian) screencapture driver: spawn, monitor, kill-early on completion — ~0.5d
  4. (amem-librarian) Step executor: bridge-driven navigate / click_selector / wait_for / extract_dom (mostly reuses RFC-001 bridge wrappers) — ~1d
  5. (amem-librarian) Caption / terminal overlay window (NSWindow with borderless Cocoa view) — ~1.5d (riskiest step; if compositing misbehaves, fall back to post-stitch mode and trim to ~0.5d)
  6. (amem-librarian) ffmpeg mp4 conversion + optional video-use handoff — ~0.5d
  7. (amem-librarian) amem_record_demo MCP tool surface + structured errors — ~0.5d
  8. (amem-librarian) Built-in cws-demo.yaml script + smoke test on the actual amem Clipper UI — ~1d (this is the dogfood — script will reveal UI bugs)
  9. (docs.amem.sh) Replace guide/self-recording.md content with pointers to this RFC + how-to for users — ~0.25d

Total: ~6.25d (5d if overlay step falls back to post-stitch).

Rejected alternatives

  • Keep chrome.tabCapture from offscreen documents. Cannot capture sidepanel; demo would have to fake the catalog UI. Disqualifying.
  • Use getDisplayMedia from the extension. User picker dialog every time; service worker lifecycle headaches; doubles the moving parts. See §7 Option B.
  • Run a generic OBS / ffmpeg pipe and let the user start/stop manually. Loses the “agent-driven scripted demo” capability that makes self- recording load-bearing. We’d be back to manual demo production.
  • Embed a video-use clone in the librarian. video-use is its own product with its own scope; reimplementing transcription / cut-filler in amem is out of scope. Optional handoff is the right boundary.
  • Ship Linux / Windows recording in v0.1. Tripled scope for an audience we don’t currently have. macOS-only is acceptable v0.1 posture.
  • Record from a JS-only stack (puppeteer + ffmpeg-screen). Loses the user’s Chrome profile and login state (per RFC-001 §3a / browser automation rules). Non-starter.
  • Skip captions / terminal overlays. Dropping captions makes the resulting mp4 unsuitable for CWS listings without manual editing, which defeats the whole “own best demo” pipeline. Worth the implementation cost.

Open questions

  • Overlay window rendering technology. Native NSWindow + Cocoa view (proposed) or a small SwiftUI helper app the librarian shells out to? Soft preference: Cocoa from Rust via the cocoa / objc2 crates; SwiftUI helper if FFI gets miserable.
  • Caption style. Plain translucent black bar with white text v0.1; do we offer themes / fonts? Soft preference: defer to v0.2; one hardcoded style in v0.1.
  • Whether to embed the YAML script in the mp4 metadata by default. Useful for reproducibility, mildly leaky for privacy if the user shares the mp4. Soft preference: opt-in (metadata: true in YAML).
  • video-use handoff API stability. v0.1 calls it as a binary subprocess; if video-use grows a Rust crate API later, switch over. Not blocking.
  • Recording lock granularity. Per-librarian (current proposal) or per-script-id? Soft preference: per-librarian (simpler; scripted recording is sequential by nature).
  • Should the sidepanel show a live preview of the recording? Tempting but adds complexity (need to read the in-progress mp4 or echo screencapture frames). Defer to v0.2.

Roll-out

  • Day 4 of v0.1 build: YAML parser + window resolver + screencapture driver land. Smoke test: record a 5-second blank capture.
  • Day 5: step executor + bridge wiring. Smoke test: 10-second scripted recording opens arxiv, clicks something, finishes.
  • Day 6: caption / terminal overlay. Risk day; fallback path ready.
  • Day 7 (CWS submission day): full cws-demo.yaml runs, video-use polishes, we have the listing video. Submit.
  • Post-v0.1: track real script usage. When the third user-script exists (per RFC-001 §6 rule-of-three trigger), re-evaluate whether YAML schema needs versioning, whether scripts need a UI, whether Linux support belongs in v0.2.
  • Future: cloud-render mode (a remote machine runs the script and returns the mp4) for users without macOS. Out of scope here; tracked in a future RFC.

RFC-004 — Reference self wiki

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-05-14
  • Supersedes: archive/003-claim-grounding.md (broader fact-check pipeline; cut)
  • Numbering note: originally drafted as RFC-003 (unarchive of the claim-grounding RFC) but renumbered to 004 because 003-recording-orchestration.md landed on main first.

TL;DR

When the model is talking with the user, it should pull cites from ~/.amem/wiki/ on its own. Today the amem_recall and amem_cite MCP tools exist but the model only calls them when explicitly told. This RFC closes that gap with one prompt resource and one convenience tool.

Scope is small on purpose: just self-reference. No web dial-out, no LLM verifier, no clipper overlay, no iOS keyboard. Those were in the archived 003 and got cut because none of them needed to ship before the basic loop works.

Motivation — where the gap actually is

Anthropic already does grounding for things they can see:

WhatCitation source
Web search toolWeb pages
Citations APIDocuments you pass in context
claude.ai Projects + FilesFiles you uploaded to their cloud

What none of those touch: markdown files you wrote on your own disk. Anthropic can’t see ~/.amem/wiki/. MCP is the seam they left for it — and amem already exposes amem_recall + amem_cite over MCP.

The remaining gap is behavioural, not technical:

Today:  user says "what did I read on transformers"
        → model speculates from training data
        → only calls amem_recall if user types "@amem" or asks explicitly

Wanted: model auto-calls amem_recall when the topic is something the user
        might have captured, attaches amem:// cite when it hits, says
        nothing extra when it misses.

Proposal

1. MCP system_prompt resource

amem mcp serve exposes one resource:

URI:   amem://system/wiki-grounding
Mime:  text/plain
Body:
  When the user asks about a paper, dataset, technical spec, talk, or
  any source-able fact they might have captured, call `amem_recall`
  with the topic's key terms BEFORE answering from training data.

  - If a hit is returned: phrase the answer in terms of the wiki entry
    and append a cite "[<cite_key>](amem://<cite_key>)" so the user can
    click through. If they want BibTeX/APA/MLA, call `amem_cite`.

  - If no hit: answer normally from training knowledge. Do NOT invent
    a cite_key. Optionally suggest: "I don't see this in your wiki —
    want me to amem_capture <url>?"

  Skip recall for: code questions, logistics, jokes, opinions, the
  user's own preferences, very-well-known facts (e.g. "Python is dynamically typed").

MCP-aware clients (Claude Code, Cursor, Cline, Zed) merge resource content into the session system prompt automatically. No client patches.

2. New tool: amem_ground(query)

Single round-trip alternative to recall→cite chaining:

amem_ground(query: string, limit?: int) -> {
  hits: [{
    cite_key:   "vaswani2017attention",
    title:      "Attention Is All You Need",
    amem_uri:   "amem://vaswani2017attention",
    excerpt:    "...",
    bibtex:     "@article{vaswani2017attention, ...}"
  }],
  inline_md: "[Vaswani et al. 2017](amem://vaswani2017attention)"
}

Useful when the model knows it’ll cite (e.g. user asked a “what did the paper say” question). Saves one MCP round-trip vs. recall+cite separately.

3. amem:// URI scheme

Stable, filesystem-independent reference:

amem://<cite_key>               # whole wiki entry
amem://<cite_key>#chunk=<n>     # specific chunk (post-v0.1)

Plus a CLI handler so links in chat are clickable from terminal:

amem open amem://vaswani2017attention
  → opens ~/.amem/wiki/1776567380_vaswani2017attention.md in $EDITOR

What’s intentionally NOT in this RFC

These were in archived 003 and get pushed out:

CutWhy
Trust list + arxiv/wiki dial-out fetchamem capture <url> already exists; user-triggered capture is enough until v0.2
LLM-verify step (claim ↔ evidence)Over-engineering before the basic recall loop is proven
amem-clipper typing observer / contradiction toastAdds 3rd UI surface; ship reference-self alone first
amem audit chat-historyAfter-the-fact, not in-conversation
iOS keyboard extensionApple keyboard sandbox is brutal; far future
default_action = block modesPaternalistic; warn-only is fine for v0.1

If reference-self works and users want more, those come back in 004/005/… in their own RFCs, scoped tightly.

Verification — how we know it works

Two metrics from dogfooding (target: 1 week, 50+ conversations):

  1. Call rate: of conversations that mention a topic the user has in wiki, what % auto-trigger amem_recall? Target: ≥60%. Below 30% = prompt too weak; strengthen. Above 90% in conversations without relevant wiki content = noisy; soften.
  2. Hit rate: of amem_recall calls, what % return ≥1 hit? Target: ≥40%. Lower means model is searching too vaguely; refine the prompt’s “key terms” guidance.

Both numbers come from logging in amem-librarian MCP server (already logs recall calls; just need a daily summary).

Open questions

  • Does MCP system_prompt resource actually flow into Claude Code’s prompt? Need to confirm against current Claude Code MCP behaviour. If resources don’t auto-merge, fallback is to bake the rule into each tool’s description field (every tool tells the model when to call it).
  • Empty-result nudge: should amem_recall return {hits: [], suggest_capture: "<url>"} when it detects a URL-shaped query? Soft yes — turns misses into capture opportunities.
  • Stale wiki entries: if you captured arxiv v1 in 2024 and the paper updated to v3, the cite is technically wrong. Out of scope for v0.1; flag for later.

Roll-out

DayWorkOwner
1amem mcp serve exposes amem://system/wiki-grounding resourceamem-librarian
1Add amem_ground tool wrapping recall+citeamem-librarian
2amem open amem://... CLI resolveramem-librarian
3Dogfood — count call/hit rates over 1 week—
7+If metrics look good: ship; if not: tune prompt and repeat—

No new infra. No client patches. Doesn’t block CWS submission of amem-clipper.

Rejected alternatives

  • Auto-recall on every model turn — too noisy; would fire on “thanks” and “no”
  • Background daemon scanning Claude transcripts — privacy + complexity for marginal gain
  • Docs telling users to type @amem — that’s the current state; doesn’t work because nobody remembers
  • Force model to cite something even on misses — turns into hallucinated cite_keys; worse than silence
  • Ship as part of amem-clipper sidepanel UI — the gap is in Claude Code / Cursor / Cline, not the browser; clipper sidepanel doesn’t help

RFC-006 — Agent-driven file upload in logged-in Chrome

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-05-22
  • Related: RFC-001 (function-based v0.1), RFC-002 (Clipper skills catalog), amem-librarian#3 (window-only capture)

TL;DR

Today no MCP-driven path lets a Claude/Cursor/Cline agent upload a file in the user’s real, logged-in Chrome. Three reasons:

  1. JS-initiated <input type="file"> click → Chrome blocks the native picker (or opens it but JS can’t see/select)
  2. crossmem bridge has no execute_script / eval action (verified 2026-05-22), so we can’t even inject the DataTransfer workaround through it
  3. Playwright / CDP routes are banned by CLAUDE.md (debug-port issues)

This RFC adds file upload as a first-class amem capability by giving amem-clipper its own bridge to amem-librarian (separate from crossmem) and landing one new MCP tool: chrome_upload_file(selector, path).

Differentiation framing: when this ships, amem-clipper is the first non-debug-port stack that lets an agent complete a file upload on any logged-in site (CWS dashboard, GitHub PR attachments, Notion image upload, Slack file drops). Closing this hole is one of the largest practical agent gaps in 2026; TARS / Connector / chrome-devtools MCP all hit the same wall.

What “file upload” actually needs

Standard web upload flow:

1. user clicks <button> / <label for=fileinput>
2. browser opens native file picker
3. user selects file(s)
4. <input type=file>.files is set
5. page reads .files, kicks off XHR/fetch

Steps 1, 5 are normal DOM operations. Step 2–4 happen inside a sandbox JS can’t reach. The workaround the web has used for ~a decade:

// in page context
const dt = new DataTransfer();
dt.items.add(file);               // file is a File object
input.files = dt.files;
input.dispatchEvent(new Event('change', { bubbles: true }));

This BYPASSES the picker entirely. The page’s own change-handler runs as if the user picked the file. Works on >95% of normal <input type=file> sites. Sites with custom drag-drop-only zones may need a drop event variant.

Why this can’t run through crossmem today

crossmem bridge actions (verified 2026-05-22 via curl /command):

✅  navigate, click, type, wait, extract, screenshot, summarize,
    tab_info, ping
❌  execute_script, eval, inject_script, capture_page,
    chrome_runtime_send, fetch_resource

Without an execute_script verb, the agent can’t push the DataTransfer snippet into the page. To stay within CLAUDE.md rules (no Playwright, no debug port), the only options are:

A. Fork crossmem to add execute_script — out of scope (third party) B. Grow amem-clipper into its own bridge endpoint — chosen C. Use Native Messaging — feasible but adds OS-specific plist/registry plumbing; left as v0.2 hardening (see §Hardening)

Architecture (chosen path: B)

Add a second bridge daemon, embedded inside amem mcp serve, scoped to amem-specific commands. crossmem stays in charge of general agent computer-use; amem owns this new lane.

                 ┌────────────────────────────────────────────┐
agent (claude)──►│ amem-librarian                             │
                 │   stdio MCP server                         │
                 │   ┌─────────────────────────────────────┐  │
                 │   │ clipper_bridge (NEW)                │  │
                 │   │   tokio HTTP server on              │  │
                 │   │   127.0.0.1:7601                    │  │
                 │   │   ├─ POST /command (agent ↔ daemon) │  │
                 │   │   └─ GET  /poll    (ext ↔ daemon)   │  │
                 │   └────────────────┬────────────────────┘  │
                 └────────────────────┼───────────────────────┘
                                      │ long-poll JSON
                 ┌────────────────────▼───────────────────────┐
                 │ amem-clipper (Chrome MV3 extension)        │
                 │   background.js — poll loop                │
                 │   └─ on "upload_file":                     │
                 │      chrome.scripting.executeScript({      │
                 │        target:{tabId:active},              │
                 │        func: dataTransferInject,           │
                 │        args:[selector, base64, mime, name] │
                 │      })                                    │
                 └────────────────────────────────────────────┘

Why a separate port from crossmem (7600 → 7601):

  • amem can be installed without crossmem and still work
  • crossmem can be uninstalled without breaking amem
  • Daemons stay single-purpose: cross-extension fan-out vs amem-specific verbs
  • Avoid editing third-party code we don’t own

Wire protocol

POST /command (agent → daemon):

{
  "id":     "<uuid>",
  "action": "upload_file",
  "params": { "selector": "input[type=file]",
              "path":     "/Users/me/file.mp4" }
}

Daemon reads the file, base64-encodes, enqueues:

{
  "id":          "<uuid>",
  "action":      "upload_file",
  "params": {
    "selector":  "input[type=file]",
    "fileName":  "file.mp4",
    "mimeType":  "video/mp4",
    "base64":    "AAAAFGZ0eXBpc..."
  }
}

GET /poll?since=<lastId> (extension → daemon, long-poll up to 30s): returns next pending command or 204 on timeout.

POST /ack (extension → daemon):

{ "id":"<uuid>", "success":true, "error":null, "data":{...} }

Daemon returns the ack back to the original POST /command caller.

DataTransfer injection snippet (runs in page context)

(selector, base64, mime, name) => {
  const bin = atob(base64);
  const buf = new Uint8Array(bin.length);
  for (let i = 0; i < bin.length; i++) buf[i] = bin.charCodeAt(i);
  const file = new File([buf], name, { type: mime });
  const input = document.querySelector(selector);
  if (!input) return { ok: false, error: 'selector not found' };
  const dt = new DataTransfer();
  dt.items.add(file);
  input.files = dt.files;
  input.dispatchEvent(new Event('input',  { bubbles: true }));
  input.dispatchEvent(new Event('change', { bubbles: true }));
  return { ok: true };
};

Failure modes & mitigations

ModeCauseMitigation
Selector matches <button> not <input>Common — many sites hide the real inputResolver tries selector → if not input[type=file], walks up to <label for> / <form> / aria-controls to find the real input. Documented in tool description.
Site uses drag-drop only zoneNo <input type=file> to setv0.1 fails fast with “no compatible input”. v0.2 may dispatch a synthetic drop event with the file in DataTransfer.
Site validates with a custom event listenerMost use change; some only inputSnippet dispatches both.
File >100MBbase64-over-localhost is slow/memory hungryCap at 50MB in v0.1; bigger files return {error: "file too large; use v0.2 chunked path"}.
Multiple file inputs on pageWrong one selectedUser must provide a specific selector. Document :nth-of-type patterns.
File path outside $HOMESurprisingResolve path; refuse if outside $HOME unless --allow-system-paths is set.
Extension not connectedamem-clipper not installed / not pollingDaemon returns 502 after 10s with “amem-clipper not connected — load the extension”.

What’s NOT in this RFC

  • Native Messaging variant — cleaner long-term, postponed to v0.2 once cross-platform installer (mac/win/linux) is built. See §Hardening.
  • Drag-drop only sites — niche, defer to v0.2
  • Chunked uploads — same; v0.1 caps at 50MB total
  • Folder uploads — webkitdirectory inputs — defer
  • Cross-frame uploads — iframes that own the input — defer

MCP tool surface

One new tool registered with the existing amem mcp serve:

chrome_upload_file(selector: string, path: string) -> string
  description:
    Upload a local file to a logged-in Chrome page via amem-clipper.
    The selector should target a standard <input type="file"> or a
    parent <label for=...>/<button> that maps to one. Path must be
    inside $HOME unless --allow-system-paths is set in amem config.

Returns:

  • success: <fileName> uploaded into <selector> on ok
  • Error: <reason> on failure (selector miss, file too large, extension not connected, etc.)

Hardening (post-v0.1, separate RFCs)

  1. Native Messaging variant — replace the loopback HTTP poll with a NM port. Faster, no port choice/conflict, survives reboots. Requires per-platform manifest install.
  2. Drop-zone fallback — synthesize drop event with DataTransfer for sites that don’t expose an <input type=file>.
  3. CWS-submission skill — built on top of chrome_upload_file, automates the 4 file-pickers in the CWS dashboard flow (issue #amem-hq/11).
  4. File chunking — for >50MB uploads, slice base64 across multiple poll messages reassembled in the extension before injection.

Roll-out

DayWorkOwner
1This RFC mergedamem-hq
1amem-librarian clipper_bridge module + MCP tool stubamem-librarian
2amem-clipper background poll + DataTransfer injectionamem-clipper
2End-to-end smoke test against a public <input type=file> pageboth
3Document in docs/guide/file-upload.md + bake into CWS-submission skillamem-hq

Nothing in this RFC blocks the current CWS amem-clipper submission (that’s manual on the 4 file pickers); but shipping this RFC means the next extension submission could be one MCP call end-to-end.

Rejected alternatives

  • Add execute_script to crossmem — out of scope; crossmem is a third-party project we shouldn’t fork unilaterally
  • Use chrome-devtools MCP (CDP) — banned by CLAUDE.md, debug-port issues
  • Browser-use / TARS visual route — they hit the same native-picker wall
  • Server-side preview + manual user step — not agent automation, defeats the point
  • Ask the user to drag-drop into a sidepanel — friction; sidepanel-only inputs don’t help when the upload form is on the target site

Open questions

  • Should the daemon also handle download mirrors (amem_download(url, path))? Likely yes — symmetric verb, same protocol shape. Out of scope for this RFC.
  • Long-term: does amem-clipper replace crossmem for our users, given the new bridge channel? Initial answer: no, parallel — crossmem keeps its generic bridge role.
  • Multi-window Chrome: if user has two Chrome windows, which gets the upload? v0.1 picks the active tab of the focused window. Documented.

RFC-007 — Skill distillation: download → wiki → Claude skill

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-07-12
  • Related: SPEC.md “Pipeline” section, RFC-004 (reference self wiki), amem-clipper#3 (Paste Board)

TL;DR

amem Clipper’s positioning is a three-step pipeline: download the resource the user is reading/watching (web page, video) → compile it into an Obsidian vault (personal wiki, knowledge graph) → distill recurring procedural knowledge into Claude skills.

Steps 1–2 ship today. This RFC defines step 3: when a cluster of wiki nodes qualifies for distillation, how the compile works, and the guardrails that keep the skill list from drowning in junk.

Corrected 2026-08-11 (competitor scan). The original claim here — “every clipper competitor stops at storage; nobody closes the loop into agent capability” — is now half wrong, and the surviving half is narrower:

  • Closing the loop to agent access is a red ocean. LLM Wiki ships a bundled MCP server plus a published llm_wiki_skill installable with npx skills add; SiYuan, Karakeep, and basic-memory all serve MCP too. Claiming nobody reaches agent capability is false.
  • Compiling captured knowledge into new skills is still unclaimed. Every one of those, verified by reading their source, consumes hand-written SKILL.md files — none generates one from what the user captured.

So the moat is not “step 3”; it is the automatic generation half of step 3. Write it that way externally, because the wider claim does not survive contact with LLM Wiki.

Three-tier semantics

TierKind of knowledgeAgent usage
wiki nodedeclarative — “what is X”storage substrate
MCP recallreference — search + cite with provenanceon-demand lookup
skillprocedural — “how to do X”auto-triggered reflex

Distillation is selective. Of ~100 captures, maybe 3 clusters represent a repeatable workflow worth compiling; the other 97 stay reference material served via amem_recall. Compiling everything would pollute skill discovery and destroy trigger accuracy.

Qualification criteria (what makes a cluster skill-worthy)

A node cluster qualifies for distillation when ALL of:

  1. Procedural — the nodes describe steps/commands/decision rules, not just facts. Heuristic: imperative verbs, numbered steps, code blocks with commands (not just definitions).
  2. Recurring — signal that this workflow repeats:
    • ≥3 captures in the same topic cluster, OR
    • the same nodes surfaced in ≥3 distinct amem_recall queries, OR
    • the user explicitly says so.
  3. Self-contained — the distilled skill can execute from its own text + linked wiki nodes, without the original page being live.

v1 trigger is explicit only: amem compile-skill <cluster> (MCP tool + CLI). Auto-suggestion (“these 4 nodes look like a workflow — distill?”) is a later phase; auto-compilation without a human in the loop is a non-goal.

Compile output

~/.claude/skills/<slug>/SKILL.md     # or project .claude/skills/ when scoped
  • Frontmatter name + description follow skill-creator conventions (description states WHEN to trigger, with 中英 keywords the user actually says).
  • Body: distilled procedure, NOT a paste of the source nodes.
  • Provenance footer is mandatory: wikilinks back to the source ~/.amem/wiki/<node_id>.md nodes + original URLs. A skill whose sources died should be auditable and re-verifiable (amem_factcheck tie-in).

Guardrails

  1. Skill budget — warn when distilled skills exceed ~20; force review of the least-triggered before adding more.
  2. No secrets — compile refuses content matching credential patterns; vault references (vault get KEY) instead of literals.
  3. Eval before install — run the eval/skill-creator grader on the generated SKILL.md; below-threshold output lands as a draft in the wiki, not in ~/.claude/skills/.
  4. Idempotent recompile — re-running on the same cluster updates the existing skill (matched by provenance), never duplicates.

Video capture (step 1 scope note)

“Download” for video = transcript + key frames, not the media file: storage, copyright, and yt-dlp maintenance all argue against full downloads, and the agent consumes the text layer anyway. Full-media archival is a non-goal.

Acceptance

  • amem compile-skill <cluster> MCP tool + CLI verb
  • Qualification check (procedural + recurring + self-contained) with human-readable rejection reasons
  • SKILL.md output with provenance footer, eval-gated install
  • Idempotent recompile on provenance match
  • E2E: 3 captures on one workflow → compile → new skill triggers in a fresh Claude Code session

Non-goals

  • Auto-compiling every capture into a skill (junk-pollution failure mode)
  • Full video/media archival
  • Distilling from sources the user hasn’t captured (that’s the frontier model’s own knowledge, not amem’s)

RFC-001 — Bridge-first architecture + settings/feature-flags

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-04-21
  • Supersedes: SPEC.md § Roadmap Day 3 item “standalone mode”
  • Related: RFC-002 (RSS), amem-sh youtube pipeline

TL;DR

Reposition the Chrome extension as amem Clipper — a sensor for the agent in the browser. The native amem binary is the brain; sensors don’t function without it. Drop standalone mode from the roadmap: the binary is a hard prerequisite. When the bridge is unreachable or a feature flag is off, the relevant UI is grayed out in place with a single-click enable/install affordance, not hidden. Heavy capture features (YouTube, RSS, Drive) are feature-flagged in ~/.amem/config.toml and lazy-loaded on demand.

Positioning and naming

amem is the brand for agent memory. The CLI (amem-sh) is the brain — sync, compile, recall, fact-check. The Chrome extension is a sensor that feeds the brain when the user is in the browser. They are not peers; the sensor depends on the brain entirely.

Name: amem Clipper. Inherits the Evernote Web Clipper lineage (users know what “clipper” means), avoids collision with the internal “bridge” process name (amem-bridge server on WS 7600), and keeps the brand “amem” attached to the core (the brain) rather than to a single sensor.

Product model:

LayerNameRole
Brainamem-sh (CLI + library + MCP)The product. Storage, compile, recall, fact-check, agent API.
Sensoramem Clipper (Chrome extension)Browser sensor — surfaces capture moments to the brain. Cannot function without the brain.
Sensoramem Pockist (iOS native, shipped 2026-04-23)Mobile sensor — share-sheet, OCR, place extraction. Same dependency on the brain.
Infraamem-bridge (WS 7600 process)Loopback IPC between sensors ↔ brain.

This positioning reshapes every downstream decision:

  • “Does the sensor need feature X?” → Only if feature X is a capture moment. Storage, recall, fact-check always live in the brain.
  • “What happens with no binary?” → The sensor is dark. A sensor without a brain is a window with no eye behind it. Frontload install; don’t half-ship.
  • “Where do settings live?” → In ~/.amem/config.toml, owned by the brain. Sensors render them via a bridge RPC; they never hold authoritative state.

Motivation

The original SPEC envisioned two extension modes:

ModeDayRequirement
Bridge2Native amem binary + WS 7600
Standalone3No native binary; cloud sync via Drive

After shipping the YouTube pipeline we hit three things that make standalone look worse than we expected:

  1. Standalone is structurally crippled. Compile requires Ollama, transcription requires whisper-rs — neither runs in a Chrome MV3 extension. A standalone capture can store URLs but cannot compile, so the wiki never builds. Agent-side MCP is also dark because MCP is native-only.
  2. Feature cost is real. YouTube alone needs yt-dlp (~20 MB), ffmpeg (~60 MB), whisper model (75 MB–3 GB depending on size). Bundling all of this into default install breaks “offline-first with zero cloud dependencies by default” by shifting the pain from network to disk.
  3. Dual code paths are a maintenance tax. Standalone + bridge would mean two storage backends, two capture pipelines, two sets of bugs. The SPEC’s principle 4 (“complement aide, don’t duplicate”) applies to our own internals too.

Meanwhile, bridge mode is already the richer experience. A 30-second curl amem.sh/install | sh is less friction than a crippled standalone fork.

Proposal

1. Bridge is the only mode — UI degrades by graying out

Clipper on cold-start pings ws://127.0.0.1:7600/status. Based on the response, individual UI regions render as enabled, gray-disabled with a one-click enable, or gray-disabled with install CTA:

StateUI for capture-webUI for YouTubeUI for RSSGlobal banner
Bridge unreachablegray, tooltip “amem not running”graygray“Install amem → curl amem.sh/install | sh” with copy button
Bridge OK, all features offenabledgray, inline “Enable YouTube (~95 MB)” buttongray, inline “Enable RSS” button—
Bridge OK, YT enabledenabledenabledgray + enable—
Bridge OK, all onenabledenabledenabled—

Rationale: grayed controls are discoverable (user sees the feature exists, understands why it’s off) and honest (no hidden states). All-or-nothing install cards punish curious first-run users; per-feature gray-out is the “sensor goes dark when disconnected from the brain” pattern the positioning promises.

2. Feature flags live in ~/.amem/config.toml

Clipper renders a Settings page that maps to keys in this file via a new bridge RPC (settings_get / settings_set). The file is the single source of truth — both CLI and Clipper read/write the same keys. Clipper holds no authoritative config of its own, consistent with the peripheral positioning.

# ~/.amem/config.toml
version = 1

[features]
youtube = false         # enables YT capture + compile (lazy-downloads yt-dlp + whisper model)
rss     = false         # enables RSS subscription ingestion (see RFC-002)
drive   = false         # enables Google Drive backup (Day 3)

[youtube]
whisper_model = "tiny.en"  # tiny.en | base.en | small.en | medium.en

[bridge]
host = "127.0.0.1"      # MUST be loopback (see Security)
port = 7600
token_file = "~/.amem/bridge.token"

3. Lazy-load on feature enable

Enabling a flag from Clipper or CLI triggers a setup routine:

amem youtube setup      # CLI: downloads yt-dlp + whisper model + checks ffmpeg
amem rss setup          # (RFC-002)
amem drive setup        # (Day 3)

Clipper setup button → bridge RPC feature_setup({name}) → server runs the corresponding amem <name> setup, streams progress back over WS so Clipper can show a progress bar inline next to the (still grayed) control.

Graceful degradation in core flows. If a user runs amem capture <youtube-url> when features.youtube = false, the CLI prints:

YouTube capture is not enabled. To turn it on:
  amem youtube setup
This will download yt-dlp (~20 MB) and the tiny.en whisper model (~75 MB).

The MCP tool amem_capture returns an analogous structured error, so agents can surface it to their user.

4. Bridge auto-start

On first install, amem install (the curl|sh script) registers a per-user background service:

  • macOS: launchctl user agent (~/Library/LaunchAgents/sh.amem.bridge.plist)
  • Linux: systemctl --user unit (~/.config/systemd/user/amem-bridge.service)
  • Windows: Task Scheduler at-logon task (deferred; Windows support is a follow-up)

Goal: after first-run setup, the bridge is as available as Ollama is today — it just runs.

Security

Bridge-always means a persistent localhost WebSocket. Three defences, all MUST land before the “always on” posture ships:

DefenceMechanism
Loopback bindingServer binds 127.0.0.1, never 0.0.0.0. Reject --bind CLI flags that widen this.
Origin allow-listWS handshake rejects connections whose Origin: header is not chrome-extension://<amem-clipper-prod-id> (production ID) or chrome-extension://<amem-clipper-dev-id> (dev build).
Token authOn bridge start we mint a 32-byte random token to ~/.amem/bridge.token (mode 0600). Extension retrieves it via native-messaging handshake at install time. All WS messages must carry {"token": "..."} in their envelope. Tokens rotate on bridge restart.

Threat model

ThreatImpactMitigation
Another local program connects to WSCould trigger amem_capture → write files to ~/.amem/raw/Token auth + origin check kill 99% of this
Disk-fill DoS via repeated captureFill user’s diskRate-limit captures per minute; refuse when ~/.amem/ exceeds configurable quota
Malicious browser extension connects as usImpersonates our extension_idChrome refuses to forge Origin: for a different extension_id
RCE via yt-dlp / ffmpeg CVEArbitrary code executionUse pinned versions, track security advisories; same posture as Ollama
Prompt injection in captured contentPoisons MCP amem_recall outputSame risk as today’s arxiv/PDF pipeline; not new from bridge-always

Net risk: slightly higher than CLI-only (persistent WS endpoint exists), lower than an HTTP server accepting remote connections. Comparable to VS Code’s language-server loopback.

Migration

  • SPEC.md §Roadmap: strike Day 3 “standalone mode”; add “bridge security hardening + install polish” and “Drive backup” (Drive stays).
  • amem-clipper (repo renamed from amem-extension 2026-04-21): delete any standalone-only code paths (none should exist yet — Day 2 skeleton is bridge-only; this is a no-op today). Rename the product surface to “amem Clipper” in README, store listing, manifest name, and UI chrome.
  • docs: guide/extension.md renamed to guide/clipper.md; its “Standalone vs Bridge” section is being rewritten as “How install works”.
  • README (amem-hq): clarify install-first story on all public pages, label the repo as “amem Clipper (Chrome MV3 extension)”.

Rejected alternatives

  • “Pure cloud standalone” — extension + Drive only, no native binary. Breaks offline-first and agent-MCP. Also introduces OAuth complexity earlier than Day 3.
  • Bundle everything in the default install. Ships ~3 GB of whisper models most users never use. Opposite of the lazy-load principle.
  • Run whisper.wasm in the browser. Early 2026 performance is still 3–10× slower than native for base.en; model download in the extension also hits MV3 storage limits.

Concrete work

See GitHub issues linked from this RFC.

  1. amem youtube setup subcommand + graceful-degradation prompt in capture (amem-sh)
  2. Bridge: loopback binding + Origin: check + token auth (amem-sh)
  3. Bridge: auto-start service installers (macOS launchd, Linux systemd) (amem-sh)
  4. Clipper: bridge status probe, per-region gray-out UI, Settings page backed by config.toml via bridge RPC (amem-clipper)
  5. Clipper: rename product surface to “amem Clipper” (manifest, store listing copy, in-UI strings) — repo slug already renamed to amem-clipper 2026-04-21 (amem-clipper)
  6. SPEC.md + docs: drop standalone, document bridge-first (amem-hq)

Open questions

  • Do we treat Ollama as a similarly lazy-loaded feature? Arguably yes — PDF compile also blocks without it. Worth a follow-up RFC if so.
  • Drive backup (Day 3): should it require Pro/auth once shipped, or stay free? Product decision, not in scope here.

RFC-002 — RSS / Atom subscription management

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-04-21
  • Related: RFC-001 (feature flags, amem rss setup)

TL;DR

Add a first-class subscription layer on top of the existing capture pipeline. Users amem sub add <feed> to follow a source (arxiv category, blog, YouTube channel); a polling loop fetches the feed, dedups by GUID, and routes each new item through the existing capture + (optional) compile flow. MCP exposes amem_subscribe / amem_sub_list so agents can manage the user’s reading queue.

Motivation

amem’s capture flow is reactive: it only runs when a human (or agent) hands it a URL. That makes it useless for tracking ongoing sources:

  • Following arxiv cs.CL as new papers drop
  • Karpathy / Willison / lesswrong blog posts
  • A YouTube channel’s new uploads (YouTube publishes per-channel RSS natively)

Every knowledge worker we’ve talked to does some version of this manually today — Feedly/NetNewsWire for reading, then copy-paste URLs into whatever capture tool they use. amem can collapse both steps.

This also inverts the self-recording story: amem is designed for you to produce content into. RSS lets other people’s content flow in on the same rails, so the wiki grows continuously rather than only after active capture.

Proposal

1. Subscription storage

# ~/.amem/subscriptions.toml
version = 1

[[subscription]]
id            = "arxiv-cs-cl"
url           = "http://export.arxiv.org/rss/cs.CL"
title         = "arXiv cs.CL (Computation and Language)"
auto_compile  = false         # capture-only by default; compile is opt-in
poll_minutes  = 60
enabled       = true
added_at      = "2026-04-21T00:00:00Z"
last_polled   = "2026-04-21T00:30:00Z"

[[subscription]]
id            = "3b1b-channel"
url           = "https://www.youtube.com/feeds/videos.xml?channel_id=UCYO_jab_esuFRV4b17AJtAw"
title         = "3Blue1Brown"
auto_compile  = true          # small channel, OK to auto-transcribe
poll_minutes  = 240
enabled       = true

2. Dedup ledger

~/.amem/subscriptions/
  ledger.jsonl                 # append-only, one JSON object per seen item
  state/{sub_id}/last_etag     # HTTP caching

Each ledger line:

{"sub_id":"3b1b-channel","guid":"yt:video:aircAruvnKk","captured_at":"2026-04-21T00:30:00Z","cite_key":"3blue1brown2017neural"}

Dedup is GUID-based. If a feed republishes an item (edit, repost), the existing capture wins; we don’t re-download.

3. CLI

amem sub add <url> [--auto-compile] [--poll-minutes N] [--title "..."]
amem sub list [--json]
amem sub remove <id>
amem sub enable|disable <id>
amem sub fetch [<id>]          # one-shot poll, honours etag
amem sub daemon                # long-running poller (used by service unit)
amem rss setup                 # install daemon (macOS launchd / Linux systemd user)

4. Poll algorithm

For each enabled sub whose now - last_polled >= poll_minutes:

  1. GET feed with If-None-Match: {last_etag} and If-Modified-Since: {last_polled}
  2. 304 → update last_polled, skip
  3. 200 → parse via feed-rs, iterate items
  4. For each item not in ledger:
    • Route to existing cite::cmd_capture(item.link) (auto-picks arxiv / PDF / YouTube based on URL)
    • If auto_compile = true → also call the appropriate cmd_compile
    • Write ledger line
  5. Update ledger + state

Failures per-item don’t block the rest of the feed. Aggregate failures re-queue with exponential backoff (15 min → 2 h cap).

5. MCP surface

amem_sub_add(url, auto_compile?) -> sub_id
amem_sub_list() -> [{ id, title, last_polled, enabled, ... }]
amem_sub_remove(id)
amem_sub_fetch(id?) -> {fetched: N, captured: M, errors: [...] }

This lets an agent maintain its own research feed without a human in the loop: “follow every arxiv paper that cites Vaswani 2017” becomes a single MCP call.

6. Gated behind features.rss

Disabled by default. amem rss setup enables it, installs the daemon, and writes features.rss = true to ~/.amem/config.toml (per RFC-001).

Non-goals

  • Rich reader UI. amem is not Feedly. Reading lives in the wiki + amem recall. If people want visual unread counts, that belongs in an extension page, not the core.
  • OPML import on day 1. Easy add later; skip for MVP to keep surface small.
  • Arbitrary scheduling cron. poll_minutes is enough; cron-syntax scheduling is out of scope.
  • Podcast audio-only feeds. These would need whisper anyway — treat them as RFC-002b when YouTube pipeline is stable on more models.

Risks

RiskMitigation
A popular arxiv category fills disk (dozens of papers/day)poll_minutes default 120 + per-feed disk quota + user confirmation on first-time auto_compile = true
Feed publisher rate-limits usHonour Retry-After, respect 429; back off to 6 h for repeat offenders
Duplicate captures when arxiv updates a paper’s versionKeep first ingest; subsequent versions append a note to the existing wiki entry rather than creating a new cite_key
RSS spec is loose — malformed feeds break parserfeed-rs handles common variants; log + skip malformed entries, do not abort the poll

Concrete work

  1. Rust crate additions: feed-rs = "2", toml_edit = "0.22" (config writes preserve comments) (amem-sh)
  2. amem sub subcommand family (amem-sh)
  3. amem rss setup installer and amem sub daemon long-runner (amem-sh)
  4. MCP tools (amem-sh)
  5. SPEC.md: add subscription to the storage layout section (amem-hq)
  6. Docs: new guide/subscriptions.md page (amem-hq)

Extension UI for managing subscriptions is deferred — CLI first.

RFC-003 — Claim-grounding & fact-check pipeline

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-04-28
  • Related: RFC-001 (bridge-first), RFC-002 (RSS), RFC-004 (planned: arxiv/wiki capture extension)

TL;DR

Fact-check is the core of amem, not a feature added on top. This RFC specifies the loop that makes amem distinct from every other “personal RAG” or “agent memory” product: every factual claim a user (or an agent speaking for them) makes is grounded against (1) the user’s own captured wiki and (2) a user-curated trust list of authoritative sources, and the result is written back so next time it’s a local hit.

Two automatic behaviours:

  1. Cite from your own wiki — amem_recall is invoked on the claim’s key terms; if there’s a hit, the claim is annotated with amem://<id> and the evidence chunk.
  2. Dial out to fetch + verify — when there’s no local hit, amem_factcheck fetches from the user’s trust list (arxiv, wikipedia, specific blogs they trust), runs a verification step (LLM compares claim ↔ retrieved evidence), and captures the result back into the wiki.

Net effect: the things you assert in conversation become traceable to either your own captured knowledge or to sources you have explicitly chosen to trust. Whatever can’t be traced is surfaced as opinion.

Motivation — what amem actually is

amem is the brand for agent memory. amem-sh (the CLI) is the brain; amem Clipper and amem Pockist are sensors that feed the brain through the browser and the phone. Frontier models access the brain through MCP.

A “memory” that just stores and recalls is half a product. The other half — and the one that nobody else is doing well — is grounding: when the agent speaks for the user, every factual claim must trace to either captured content or a trusted source. Without that, the memory is just a fancier ChatGPT context window.

Two unsolved problems amem-grounding addresses:

  1. Agents speaking for the user invent plausible facts unless rigorously prompted. Even good models hallucinate at ~5–30% on long-tail factual claims.
  2. Humans speaking for themselves repeat half-remembered facts without knowing they’ve contradicted something they read last month.

amem already owns the storage layer (sensors capture; brain compiles). This RFC adds the outbound retrieval loop that closes the system.

Existing tools don’t fill this gap:

ToolWhy it doesn’t ground claims
Glasp / ReadwiseCapture + highlight only; no retrieval at speak-time
Notion AI / Mem.aiCites within their own product; doesn’t help when typing in Slack
Perplexity / Claude Web SearchWeb-wide ranking, not the user’s trust list, no persistence
ChatGPT memory / Claude memoryBlack-box server-side, no source surfaced, no user-curated audit
Letta / Mem0Agent memory with no grounding obligation; cites are not enforced

amem is uniquely positioned because it already owns the user’s local corpus, already runs a sensor stack feeding it, and already exposes everything via MCP.

Proposal

1. New MCP tool: amem_factcheck

amem_factcheck(claim: string, context?: string) -> {
  verdict:   SUPPORTED | CONTRADICTED | NO_EVIDENCE,
  evidence:  [{ source_uri, snippet, captured_at, similarity }],
  amem_uri:  "amem://<id>"
}

Pipeline:

claim
  ├─ amem_recall(claim) ──────────────────► local hit?
  │                                              │
  │                                       ┌──────┴──────┐
  │                                       ▼             ▼
  │                                      HIT          MISS
  │                                       │             │
  │                                       │             ▼
  │                                       │      walk trust list
  │                                       │      (priority order)
  │                                       │      └─► fetch via adapter
  │                                       │            └─► amem_capture
  │                                       │                  (becomes local
  │                                       │                   hit next time)
  │                                       │
  │                                       └─► LLM verify:
  │                                           "Does <evidence> support
  │                                            <claim>?"
  │                                           → SUPPORTED / CONTRADICTED
  │                                             / NO_EVIDENCE
  │
  └────────────────────────────────────────► return verdict + evidence + uri

amem_recall (already shipped Day 1) is the local search; amem_factcheck adds the dial-out + verify steps.

2. amem:// URI scheme

Every cite uses one canonical format:

amem://<item_id>           # whole captured item
amem://<item_id>#chunk=<n> # specific chunk

Why a custom scheme rather than https://...:

  • Stable across ~/.amem/ filesystem reorganisations
  • Works offline (the URI resolves to local content first, with https:// fallback for external clients)
  • Identifies content uniquely whether the user clipped it or it was pulled via factcheck

A small CLI handler resolves amem://... to the corresponding local file or chunk for human-readable preview (amem open amem://abc#chunk=3).

3. Trust list in ~/.amem/config.toml

[grounding]
enabled = true
default_action = "warn"          # warn | block | log

[[grounding.sources]]
name = "arxiv"
url_pattern = "arxiv.org/abs/*"
priority = 100                   # higher = checked first

[[grounding.sources]]
name = "wikipedia-en"
url_pattern = "en.wikipedia.org/wiki/*"
priority = 80

[[grounding.sources]]
name = "fed-h15"
url_pattern = "federalreserve.gov/datadownload/Output.aspx?rel=H15"
priority = 90

[[grounding.sources]]
name = "user-rss"                # see RFC-002
priority = 50

Design stance: amem does not ship an opinionated default trust list. Users curate their own. This is a deliberate epistemic position — the tool helps you ground claims in sources you chose, not in sources the tool’s author chose for you. (amem grounding init --opinionated can offer a starter list for first-runs who don’t want to think about it.)

4. Surfaces (phased)

PhaseSurfaceEffortEffect
AAgent system prompt rule via MCP system_prompt resource1 dayEvery MCP-aware client (Claude Code, Cursor, …) auto-cites when speaking for the user
Bamem_factcheck tool implementation (trust list + dial-out + verify)3 daysReal grounding capability available to any caller
Camem clipper input overlay on claude.ai / chatgpt.com / slack.com / gmail.com1 weekCatches human-typed claims, not just agent-generated ones
DWeekly amem audit chat-history + iOS keyboard extensionTBDPost-hoc + cross-device coverage

Phase A in detail (ship-this-week)

amem mcp serve registers a system_prompt MCP resource:

Resource URI:  amem://system/grounding-rule
Content:
  When generating any factual claim about the world (not opinions, not
  questions, not logistics), you MUST first call `amem_recall(claim)`.
  - If a result is returned, append `[src: amem://<id>]` to the sentence.
  - If no result is returned, call `amem_factcheck(claim)`.
  - If factcheck returns NO_EVIDENCE, prefix the sentence with
    "Without source: ".
  - If factcheck returns CONTRADICTED, stop and surface the contradiction
    to the user before proceeding.
  Logistics, opinions, jokes, and the user's own preferences do not need
  citations.

MCP-aware clients (Claude Code, Cursor, Claude.ai mobile) read this resource and merge it into their system prompt automatically. Zero new infra; zero client work.

Phase B — amem_factcheck implementation

Per-source adapter strategy is intentionally minimal in v1:

#![allow(unused)]
fn main() {
// Hardcode the two common adapters. Generalise to a trait only after
// the third source ships (rule of three).
async fn factcheck(claim: &str) -> Verdict {
    if let Some(hit) = amem_recall(claim).await? {
        return verify_with_llm(claim, &hit);
    }
    for source in trust_list_sorted_by_priority() {
        match source.name.as_str() {
            "arxiv"        => arxiv_factcheck(claim).await,
            "wikipedia-en" => wiki_factcheck(claim).await,
            _              => generic_http_factcheck(&source, claim).await,
        }
    }
    Verdict::NoEvidence
}
}

Both adapters reuse code from RFC-004 (arxiv/wiki capture extension), which ships alongside.

Phase C — clipper overlay

The amem Clipper extension (already has content scripts on claude.ai / chatgpt.com / gemini.google.com for autosave, per current manifest) acquires a new responsibility: debounced typing observer.

UI sketch:

[user typing in claude.ai compose box]
  └─ debounce 800ms
     └─ extract latest sentence ending in . ! ?
        └─ background.js → bridge POST /factcheck
           └─ if CONTRADICTED → side-panel toast:
              ┌────────────────────────────────────┐
              │ ⚠ "GPT-4 hits 50% on ARC-AGI"     │
              │ Your wiki [Chollet 2024]: 30%      │
              │ [insert correct] [I have a new src]│
              │ [different metric] [skip]          │
              └────────────────────────────────────┘

The overlay is non-blocking; the user can ignore it. It’s a hint, not a gate.

Phase D — audit + iOS

After-the-fact: amem audit chat-history --since=last-week ingests Slack / iMessage / Gmail exports (where the user has authorised), runs each message through amem_factcheck, and produces a markdown report grouped by:

  • ✅ supported (cited from your wiki or trust list)
  • ⚠ contradicted (you said X, your wiki says Y)
  • ❓ unverified (no evidence in your trust list)

iOS keyboard extension: same idea as Phase C, but in a system keyboard. Painful to build (Apple keyboard sandbox is nasty), so deferred until A/B/C have shown value.

Privacy

  • Trust list ≠ “the web”: amem_factcheck only contacts hosts the user has explicitly listed. There is no fallback to a generic web search.
  • Outbound traffic is logged: every dial-out is recorded in ~/.amem/grounding.log with timestamp, host, and claim hash.
  • Claims never leave the loopback bridge unless they’re being sent to a trust-list source. The local LLM verify step uses the locally configured Claude API key (or a local model if configured) — same posture as existing amem capture enrichment.
  • No federation by default: the captured factcheck results stay in ~/.amem/. They are not shared with any community cache.

Failure modes

ModeCauseMitigation
Over-citationCasual chat triggers factcheck on every “hi”Phase A prompt rule explicitly excludes “logistics, opinions, jokes”. Phase C overlay only fires on full-sentence claims.
Latencyfactcheck takes 1–3s; agent feels slow(1) cache verified claims by content hash; (2) run async, let agent draft sentence first then annotate; (3) skip factcheck on claims < 5 words
Source biasUser’s trust list is itself wrongThis is on the user. amem does not adjudicate truth, it just enforces traceability. (Documented as a feature, not a bug.)
Contradiction noiseOld captured content disagrees with current consensusSurface both with timestamps; let user mark old item as superseded_by
Evidence ≠ supportLLM verify says SUPPORTED but it’s a misreadShow the evidence snippet next to the claim so user can sanity-check the reasoning step
Privacy regressiondial-out reveals reading habits to sourceDocumented in trust list config; user can disable per-source or globally

Concrete work

In rough order:

  1. (amem-sh) MCP system_prompt resource with the grounding rule (Phase A)
  2. (amem-sh) amem:// URI scheme + resolver CLI
  3. (amem-sh) amem_factcheck tool with hardcoded arxiv + wiki adapters (Phase B)
  4. (amem-sh) ~/.amem/config.toml [grounding] section + parser
  5. (amem-clipper) input observer + side-panel toast (Phase C)
  6. (amem-sh) amem audit chat-history (Phase D, slack/imessage/gmail importers)
  7. (amem-pockist) iOS keyboard ext (Phase D, much later)

Rejected alternatives

  • “Build a generic SourceAdapter trait + trust-graph DSL” — premature abstraction. Hardcode arxiv + wiki; revisit on the third source.
  • “Use Perplexity / Claude Web Search as the dial-out” — abandons the user-curated trust list, which is the whole point. The web-search products are a different product category (entropy reduction over global web), not what amem is for (entropy reduction over your chosen sources).
  • “Push factcheck verdicts to a community cache” — privacy mess + trust mess + premature for a pre-CWS product. Federation is RFC-005-or-later.
  • “Block sending when CONTRADICTED” — too paternalistic. default_action = "warn". Block is opt-in for users who want stricter discipline.

Open questions

  • Should Phase A’s prompt rule include language for how to phrase uncertainty (e.g., “according to my wiki, …”) vs. leaving that to the client? Soft preference: prescribe a template, since otherwise every client prints citations differently.
  • Should amem_factcheck block on contradiction in the verdict, or always return both verdict + evidence and let the caller decide UX? Soft preference: always return both; UX gating is a client concern.
  • Trust list versioning: when the user changes their trust list, do previously stored factcheck results need re-evaluation? Open.
  • Should agent-generated claims that get NO_EVIDENCE be auto-filed somewhere (~/.amem/unverified/) for the user to either capture-the-source or reject-as-opinion later? Probably yes, but UX needs design.

Roll-out

  • Day 8: ship Phase A (1 day of work). amem mcp serve exposes claim-grounding-rule resource. amem-using devs immediately notice their agent starts citing.
  • Day 9–11: ship Phase B. amem_factcheck lands.
  • Day 12+: Phase C in amem-clipper, scheduled after CWS approval + sidepanel UI work.
  • Phase D: open-ended, after first 5 external testers report grounding is the feature they actually use.

RFC-005 — Keyboard-driven region capture (multi-source)

  • Status: Draft
  • Authors: @yiidtw
  • Created: 2026-05-06
  • Related: RFC-001 (bridge), RFC-003 (grounding — this RFC supplies the citable units), RFC-007 (planned: macOS Accessibility hint mode)

TL;DR

Make every visual region a user can point at — a figure on a webpage, a table on a PDF, a frame in a design tool — addressable by a stable amem:// URI. Reference once, reference forever, even after the source artifact is renamed, mutated, or deleted by an agent.

Two sources ship together in v0:

  1. Chrome hint mode — keyboard hotkey on amem Clipper draws vimium-style hints over <figure>, <img>, <table>, headings, and selectable blocks; user types the hint; region is captured + assigned amem://<uuid>.
  2. Sioyek annotation watcher — amem-sh tails Sioyek’s local SQLite db; every keyboard rectangle-mark (r mode) the user makes auto-becomes an amem://<uuid> capture without changing Sioyek’s UX.

Unified output: every captured region surfaces in the user’s session with a short alias like @fig-A that paste-able into agent prompts. The agent reads the region image + metadata via the existing amem_recall MCP tool.

Motivation — the fig 4 → fig 2 problem

A scenario the operator hits weekly:

User:  "Claude, change fig 4's color to blue."
Claude: <does the modification, but renames it to fig 2 in the process>
User:  "fig 4 is now where? was it deleted?"
Claude: <confused about identifier history>
User:  Spends 5 minutes verbally re-anchoring "the figure showing X-axis Y"
       …or gives up and re-screenshots.

The bug isn’t the agent. The bug is that fig 4 is a name binding, not an identity reference. As soon as the agent mutates the underlying artifact, the binding breaks but the user’s mental model still points at “that thing.”

The solved-problem analogues live elsewhere in software:

  • Git’s commit hash (content-addressable; renaming the branch doesn’t invalidate the commit)
  • Notion’s per-block immutable IDs
  • Roam’s ((block-ref))
  • amem’s existing amem://<uuid> (already content-addressable for whole-document captures)

This RFC extends amem:// from document-grade addressing to region-grade addressing — and gives the user a keyboard-only path to mint such IDs from anywhere they read.

The killer secondary effect: every amem:// minted this way is automatically a citable unit for RFC-003 fact-check. Grounding gets concrete targets (“amem://b3f29... shows X”) rather than vague text recalls.

Proposal

1. amem:// URI extended for regions

Existing scheme:

amem://<item-id>                  # whole captured document
amem://<item-id>#chunk=<n>        # specific text chunk

Add region addressing:

amem://<region-id>                # always a 16-char base32 of UUID
                                  # mints a NEW item, not a sub-ref of an
                                  # existing one — content-addressable, so a
                                  # region you saved last week resolves to
                                  # the same id even if the source is gone

A region is a first-class amem item with these required fields:

# example: ~/.amem/raw/regions/b3f29x7q4t8m6w2n.toml
uri        = "amem://b3f29x7q4t8m6w2n"
short      = "fig-A"             # session-scoped alias, may collide across sessions
source_app = "amem-clipper"      # or "sioyek", "macos-ax", ...
captured_at = 2026-05-06T11:23:08Z

# What was captured
image_path  = "raw/regions/b3f29x7q4t8m6w2n.png"
ocr_text    = "Figure 4: Self-attention scores across heads…"

# Where it came from (best-effort, may be partial for some sources)
source_url  = "https://arxiv.org/pdf/1706.03762"
source_doc  = "Attention Is All You Need (Vaswani 2017)"
source_page = 7
source_bbox = [120, 340, 480, 620]    # x, y, w, h in source coordinates
source_dom  = "main > article > figure:nth-of-type(4)"  # if web

# Optional, for grounding
parent_item = "amem://efb1...."  # if region was extracted from an existing
                                  # whole-doc capture

The point is: the URI alone is enough to recover the image + context even after the source is offline / mutated / removed. amem caches the image bytes locally; the source metadata is provenance, not the source of truth.

2. New MCP tool: amem_capture_region

amem_capture_region(target?: TargetSpec, mode?: "hint" | "wait")
  -> { uri: "amem://...", short: "fig-A", image_path, source_meta }

TargetSpec :=
  | { app: "chrome", tab: "active" | <tabId> }
  | { app: "sioyek", current_doc: true }
  | (omitted)            // amem decides based on which source is "live"

Behaviour by mode:

  • mode: "hint" (default for Chrome) — agent calls the tool; bridge sends enter_hint_mode to amem Clipper; user types a hint; the call blocks until the user picks (or 60s timeout aborts).
  • mode: "wait" (default for Sioyek) — agent calls the tool; amem waits for the next new entry in Sioyek’s db (or 60s timeout); returns the region that landed.

Both are blocking from MCP’s perspective — clients (Claude Code, Cursor, etc.) already render “tool running” UI for long-running calls. The user isn’t surprised; they expect the tool to wait for their selection.

3. Source 1 — Chrome hint mode (amem-clipper)

Keyboard flow:

User in Chrome: presses ⌥G  (or invoked via amem_capture_region MCP call)
       ↓
Content script enumerates candidate elements:
  • <figure>, <img>, <table>, <video>, <pre>, <code>
  • <h1>–<h4>
  • Block elements with computed area > 8000 px²
  • PDF.js text-layer divs (when PDF rendered in Chrome)
       ↓
Each element gets a 2-letter hint label: aa, ab, ac, …
       ↓
Overlay rendered at element's top-left corner with high z-index
       ↓
User types: "ac"
       ↓
Content script captures the region:
  • Screenshot via chrome.tabs.captureVisibleTab + crop to bbox
  • Plus: outerHTML of the element (DOM range archive)
  • Plus: element's computed CSS rect, page URL, page title
       ↓
Posts to bridge: { action: "region_captured", payload: <region-record> }
       ↓
amem-sh receives, mints amem://<uuid>, stores PNG + metadata, returns short alias

Hint allocation: depth-first DOM traversal, skipping invisible elements (zero size or display:none). Two-letter hints support up to 676 elements per page; in practice ~50–100 visible candidates is plenty.

Activation: extension’s existing capture button gets a long-press / shift modifier for “hint mode,” or a new dedicated icon. The hotkey ⌥G is a reserved keyboard shortcut declared in manifest.json.

4. Source 2 — Sioyek annotation watcher (amem-sh)

Sioyek already provides keyboard-driven rectangle selection via the r command — and it stores the result in a stable SQLite database. This RFC does not require changes to Sioyek; we just listen.

~/.config/sioyek/local.db       ← Sioyek writes (highlights, bookmarks, marks)
       │
       ▼ fsnotify
amem-sh sioyek-watcher loop:
  every change event ↓
       ▼
  SELECT * FROM highlights
  WHERE creation_time > last_seen_creation_time
       ▼
  for each new row:
    • read (begin_x, begin_y, end_x, end_y, page, document_path, type, creation_time)
    • render that bbox of <document_path>:<page> via mupdf/pdfium  → PNG
    • OCR the rendered region via Vision/tesseract → text
    • mint amem://<uuid>, store under raw/regions/
    • update last_seen_creation_time
  end
       ▼
  emit MCP notification (if any client subscribed) so the agent can refresh

Trade-offs:

  • ✅ Zero Sioyek modification, zero plugin maintenance
  • ✅ Keyboard-only (Sioyek’s r mode already is)
  • ✅ Cross-platform (SQLite is portable; mupdf works on Mac/Linux)
  • ⚠️ Couples to Sioyek’s db schema. We pin ~/.config/sioyek/local.db at SQLite version we tested with; on Sioyek upgrade we re-validate.
  • ⚠️ User has to be using Sioyek for this source to fire — non-Sioyek PDF users go via Chrome PDF viewer (see RFC-005 §3 / “PDF in Chrome” path)

5. Short alias system (@fig-A)

The full URI amem://b3f29x7q4t8m6w2n is unwieldy in chat. amem maintains a per-session alias table:

~/.amem/state/aliases.json
{
  "session_id": "2026-05-06-am",
  "started_at": "...",
  "aliases": {
    "fig-A": "amem://b3f29x7q4t8m6w2n",
    "fig-B": "amem://4ek38u2t1y9q5w0v",
    "tbl-A": "amem://nz0p7qm6r3s4d2f8"
  }
}

Alias scheme:

  • Prefix indicates type: fig- (figure / image), tbl- (table), txt- (text block), pg- (whole page screenshot), dom- (DOM range)
  • Suffix is alphabetic, in capture order this session
  • Resets at session start (user can amem alias persist to keep them)

CLI surface:

amem aliases                    # list current session aliases
amem alias persist               # promote session aliases to permanent
amem alias resolve @fig-A        # → amem://b3f29x7q4t8m6w2n
amem alias forget @fig-A         # remove from current session

In agent conversations, the amem_recall tool already resolves amem:// URIs. Aliases get resolved client-side by amem-sh before the MCP call: the user pastes @fig-A, the agent’s amem_recall sees the full URI.

6. Normalized capture record (cross-source schema)

All sources produce records of the same shape, regardless of provenance:

#![allow(unused)]
fn main() {
struct RegionCapture {
    uri:           String,         // "amem://<uuid>"
    short_alias:   Option<String>, // "fig-A"
    captured_at:   DateTime,

    // Always present
    image_path:    PathBuf,        // PNG, full-resolution
    image_bytes:   u64,            // for budget tracking

    // Optional but encouraged
    ocr_text:      Option<String>,
    image_caption: Option<String>, // alt text, figure caption, etc.

    // Source provenance (one of these blocks present)
    source: SourceProvenance,
}

enum SourceProvenance {
    Web {
        url:        String,
        title:      String,
        dom_path:   String,
        bbox_css:   [f32; 4],
        outer_html: Option<String>,  // archived for posterity
    },
    Sioyek {
        document_path: String,
        document_hash: String,       // for tracking renames / moves
        page:          u32,
        bbox_pdf:      [f32; 4],     // PDF coordinates
        document_title: Option<String>,
    },
    MacOSAX {                         // RFC-007, placeholder
        bundle_id:     String,
        window_title:  String,
        bbox_screen:   [f32; 4],
    },
}
}

Every consumer (amem_recall, the side panel, amem cite) handles a RegionCapture regardless of source. New sources just add new SourceProvenance variants.

Privacy

  • All captures land only in ~/.amem/raw/regions/ on the user’s machine. Never uploaded.
  • Source provenance metadata is verbose by design (DOM path, PDF coordinates) so the user can audit. Verbosity stays local.
  • The Sioyek watcher reads ~/.config/sioyek/local.db only; it never writes to it. Sioyek’s own data integrity is unaffected.
  • amem Clipper’s hint mode uses the same chrome.tabs.captureVisibleTab permission already declared. No new permissions.
  • Region OCR runs locally (Vision on macOS, tesseract on Linux). No cloud OCR.

Failure modes

ModeCauseMitigation
Hint mode times out (60s)User got distracted, never pickedMCP tool returns a structured timeout error; agent can retry or apologise
User picks two hints simultaneously (race)Multiple keypresses queuedLast-arrived wins; emit a debug log
Sioyek db schema changes after upgradeApple/Sioyek release breaks SQLWatcher pins schema version, gracefully disables itself with a warning if the schema doesn’t match; user gets amem doctor sioyek to update
Same source region captured twiceUser picks the same hint or makes the same Sioyek annotationContent-hash dedupe; second capture gets the same amem://uri as the first
Region image too large (e.g., a full-screen 4K screenshot)High-resolution monitor + lazy hintCap raw image at 4MB; downscale rest; preserve original under raw/regions/full/ if the user wants it
OCR misses non-Latin scriptsDefault tesseract langBoth Vision (macOS, multi-language) and tesseract are configured for zh-Hant, zh-Hans, en, ja to match amem-pockist

Concrete work

In rough order of dependency:

  1. (amem-sh) Extend amem:// URI parser to accept region IDs (compat with existing item IDs — same UUID space)
  2. (amem-sh) Add RegionCapture data model + storage layout under raw/regions/
  3. (amem-sh) Add amem_capture_region MCP tool (blocking, mode parameterised)
  4. (amem-sh) Add amem alias CLI subcommands + session alias state
  5. (amem-bridge) Add enter_hint_mode + region_captured verbs to bridge protocol
  6. (amem-clipper) Add hint overlay content script (~3 days; see §3)
  7. (amem-clipper) Wire ⌥G hotkey + side-panel “hint mode” toggle
  8. (amem-sh) Add Sioyek watcher (~/.config/sioyek/local.db poller + PDF region renderer via mupdf)
  9. (amem-sh) OCR pipeline (Vision + tesseract fallback) for region images
  10. (docs.amem.sh) User guide page: “Capturing regions for AI references”

Estimated total: two weeks of focused work, parallelisable across amem-sh / amem-bridge / amem-clipper. Sioyek source ships in week 1 (simpler — db reads only); Chrome hint mode ships in week 2.

Rejected alternatives

  • “Write a Sioyek Lua plugin to bind a hotkey to amem capture” — duplicates Sioyek’s built-in r mode, requires per-Sioyek-version maintenance, and breaks if Sioyek loses Lua support. The fs-watch approach decouples completely.
  • “Mouse-driven rectangle selector” — operator’s stated preference is hand-stays-on-keyboard. Mouse mode could be a v2 nice-to-have but isn’t the design center.
  • “Just screenshot and let the agent OCR/describe” — loses stable identity. Two screenshots of the same figure get different IDs; agent can’t track “the same thing” across calls.
  • “Push everything through Chrome only (drop Sioyek)” — operator uses Sioyek for daily PDF reading. Forcing them into Chrome is a UX regression for the work Sioyek is good at.
  • “Build native macOS overlay with Accessibility API now” — would cover all native apps generically, but is multi-week macOS work and is unjustified before Chrome + Sioyek prove the model. Deferred to RFC-007.

Open questions

  • Alias naming policy — should aliases be fig-A style (semantic prefix + letter), or pure @a1, @a2, …? Soft preference: semantic prefix; users immediately know @fig- ≠ @tbl-. But it requires classification at capture time (heuristic on element tag / PDF caption detection).
  • Cross-session alias persistence — should aliases auto-persist if the user uses them in a chat that the agent also persists into amem (closing the loop)? Probably yes; soft preference: aliases referenced in any captured conversation get auto-persisted.
  • Multi-monitor / multi-window — when the user has two Chrome windows on different monitors, which one gets the hint overlay? Soft preference: only the focused window. If the agent calls amem_capture_region(target={app: chrome, tab: ...}) with no tab specified, route to the active window’s active tab.
  • PDF.js inside Chrome — should the hint mode treat PDF text-layer divs as targets (currently yes per §3), or should it route to the Sioyek source instead? The user’s choice of viewer is the answer; if the PDF is in Chrome, hint mode handles it.
  • Should amem_factcheck (RFC-003) accept a region URI as the claim context? Soft preference: yes — amem_factcheck(claim, context_uri: "amem://<region>") makes grounding richer.

Roll-out

  • Week 1 — RFC-005 implementation kickoff. Sioyek watcher first (cleanest path, no extension changes). Ships behind feature flag features.sioyek_capture = true.
  • Week 2 — amem Clipper hint mode lands. Side panel shows live capture log. Ships behind features.region_capture = true until stable.
  • Week 3 — wire amem_factcheck to accept region URIs (RFC-003 Phase B integration). Now grounding loop is end-to-end.
  • Post-CWS — promote both flags from beta to default. Update docs.amem.sh/clipper and add a new docs.amem.sh/sioyek page.
  • Future — RFC-007 macOS Accessibility hint mode, when first user asks “I want this for Sketch / Figma desktop / Adobe.”

Privacy Policy

The policy for amem Clipper and the amem CLI lives at:

https://yiidtw.dev/projects/amem-clipper/policy

That is the URL registered with the Chrome Web Store, and it is the only copy.

Why it is not here

A privacy policy is a URL a store holds on file, and it has to outlive the product’s branding, its domain, and whether a given repo is public this month. This one was served from docs.amem.sh, built out of a private repo — and when that repo went private on 2026-08-22 the page’s own header links, its contact channel, and four links in its body all became 404s at once. The document that tells a user how to reach you should not depend on the thing it is documenting.

Keeping a second copy here would recreate the failure this project keeps running into: CWS_LISTING.md sat at v0.2 for months while the live store text moved on, and was read as authoritative anyway. A pointer cannot go stale in that way.