amem.sh
Shared knowledge for agents and humans.
amem is a local-first knowledge engine. It captures references (arXiv papers, PDFs, web pages), compiles them into structured wiki notes with SHA-256 provenance, and serves them to AI agents over MCP. Everything runs on your Mac — no cloud.
Why amem
Frontier models write better papers when they can cite real sources. Humans maintain better notes when capture is friction-free. amem serves both audiences from one knowledge base:
- For agents: MCP tools
amem_capture,amem_compile,amem_cite,amem_recall— grounded citations with verifiable provenance - For humans: a CLI and a Chrome extension that drop into your existing workflow, storing everything as readable markdown under
~/.amem/
Relationship to aide.sh
aide orchestrates agents; amem gives them memory. The two are complementary:
| Layer | Tool | Concern |
|---|---|---|
| Orchestration | aide.sh | Dispatch, budgets, teams |
| Knowledge | amem.sh | Capture, compile, recall |
aide can use amem as its memory sync method ([sync.memory] method = "amem" in aide.toml).
Quick links
Installation
Prerequisites
- macOS (primary target) or Linux
- Rust toolchain (
cargo) - Ollama running locally — required for the
compilestep pdftotext(frompoppler) —brew install poppler
Install the CLI
The CLI is in private beta and is not publicly distributed yet. If you have access to the source:
cargo install --path .
The Chrome extension works without it — captures stop at page title and URL, which is the part that needs no daemon.
Verify
amem --version
amem --help
Register with Claude Code
cargo install does not auto-register the MCP server. Run this once:
amem mcp install
This registers amem under your user scope (~/.claude.json). After restarting Claude Code, /mcp will show amem · ✔ connected, giving any agent access to amem_capture, amem_compile, amem_cite, amem_recall.
The one-liner shell installer (install.sh) runs this automatically if claude is on PATH.
To undo: amem mcp uninstall.
Pull an Ollama model
ollama pull llama3.1
Override the default model with AMEM_OLLAMA_MODEL=<name> if needed.
Next
- Quick Start — capture your first paper
- Concepts — how amem thinks
Quick Start
Capture a paper, compile it into a wiki note, and recall it.
1. Capture
amem capture https://arxiv.org/abs/1706.03762
Downloads the PDF to ~/.amem/raw/ and prints a cite key (e.g., vaswani2017attention).
2. Compile
amem compile vaswani2017attention
Parses the PDF, chunks it with SHA-256 provenance, runs Ollama paraphrase passes, and writes ~/.amem/wiki/{ts}_vaswani2017attention.md.
3. Recall
amem recall "attention mechanism"
Grep-searches your wiki and returns excerpts with cite keys.
4. Cite
amem cite vaswani2017attention --format bibtex
Prints a formatted citation.
5. Hook it up to Claude
claude mcp add amem -- amem mcp serve
Then ask Claude: “Cite vaswani2017attention in APA format using the amem MCP server.”
Next
- MCP Tools — the four tools agents use
- Storage Layout — what’s in
~/.amem/
Concepts
Knowledge, not chat logs
Most “second brain” tools store a stream of captures — articles, notes, clippings — and hope you can find them later. amem does something different: it compiles captures into wiki notes with verbatim quotes and verifiable provenance. The wiki is the product; the raw captures are just source material.
This follows Andrej Karpathy’s recommendation of maintaining a personal wiki as the substrate for long-term thinking.
Dual audience
amem serves both agents (over MCP) and humans (via CLI + extension) from the same store. An agent citing a paper sees the same markdown a human sees. There’s no hidden “agent memory” that drifts from what’s on disk.
Provenance by construction
Every chunk carries a SHA-256 hash of its source. If the original file changes, amem verify detects drift. Citations stay grounded.
Offline-first
Zero cloud dependencies by default. Ollama runs locally. The only network calls are to fetch papers from their public URLs (arXiv, DOI resolvers). You can operate amem entirely air-gapped after initial capture.
Three interfaces
- CLI —
amem capture,amem compile,amem recall,amem cite - MCP server —
amem_capture,amem_compile,amem_cite,amem_recallfor agents - Chrome extension — one-click web page capture + self-recorded demos
MCP Tools
amem exposes four MCP tools over stdio. Start the server with amem mcp serve.
amem_capture
Download a paper and generate a cite key.
Input: url (string) — arXiv URL, DOI, PDF URL, or local file path
Output: cite_key (string), raw_path (string)
amem_compile
Parse + chunk + paraphrase a captured source into a wiki note.
Input: cite_key (string)
Output: wiki_path (string), chunk_count (number)
amem_cite
Format a citation in a supported style.
Input: cite_key (string), format (string, one of bibtex / apa / mla / chicago / ieee)
Output: citation (string)
amem_recall
Search the wiki for matching chunks.
Input: query (string), limit (number, optional, default 10)
Output: list of {cite_key, excerpt, sha256, score}
Register with Claude
claude mcp add amem -- amem mcp serve
Register with other MCP clients
Any MCP client that supports stdio transport works. Point it at the amem mcp serve command.
Storage Layout
Everything lives under ~/.amem/.
~/.amem/
├── raw/ # original captures
│ ├── vaswani2017attention.pdf
│ └── ...
├── wiki/ # compiled notes
│ ├── 20260418_vaswani2017attention.md
│ └── ...
└── index.md # auto-maintained TOC
raw/
Original files exactly as downloaded. amem verify re-hashes against these.
wiki/
Compiled markdown notes. One file per cite key. Filename prefix is the compile timestamp so recompiles preserve history (you can git init ~/.amem/wiki && git commit if you want versioning).
index.md
Auto-regenerated on every amem compile — a flat list of all cite keys with paths. Never edit by hand.
Custom location
Override with AMEM_HOME=<path> (planned — not yet wired up).
Citation Formats
amem cite <cite_key> --format <fmt> supports:
| Format | Flag | Use |
|---|---|---|
| BibTeX | bibtex | LaTeX papers |
| APA | apa | Social sciences |
| MLA | mla | Humanities |
| Chicago | chicago | History, arts |
| IEEE | ieee | Engineering |
Default: bibtex.
Example
$ amem cite vaswani2017attention --format apa
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N.,
Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need.
Metadata sources
- arXiv API (for arXiv papers)
- CrossRef (for DOIs)
- Extracted from PDF metadata as fallback
amem Clipper (Chrome Extension)
amem Clipper is the browser-side companion to the native amem CLI — a Chrome MV3 side-panel extension for capturing web pages into your amem knowledge base. Positioning and install flow are defined by RFC-001.
Status
Bridge mode (Day 2 — in progress): extension talks to amem locally over WebSocket on port 7600. Extension captures → bridge forwards → amem stores.
Standalone mode (Day 3 — planned): when amem isn’t running, the extension falls back to Google Drive backup. Requires a new OAuth client (not shipped yet).
Install
The extension will be published on the Chrome Web Store as a new listing under the same developer account as crossmem (the v0 prototype). crossmem remains published and unchanged.
Published URL will appear here after approval.
Self-recording workflow
The extension ships with a self-recording skeleton: chrome.runtime.sendMessage({cmd: 'start_recording'}) triggers a tabCapture session via the offscreen document, producing amem-recording-<iso>.webm in Downloads.
This is how amem produces its own demo videos — the extension records itself being used. See Self-Recording Workflow.
Self-Recording Workflow
One of amem’s foundational features: amem is its own best demo. The extension records itself being used, producing the marketing video and Chrome Web Store listing assets automatically.
How it works
- An orchestrator (human or agent) sends
{cmd: 'start_recording'}to the extension’s background service worker. - The background worker opens an
offscreen.htmldocument with aMediaRecorderconsumingtabCapture. - The agent then drives the extension UI through the bridge — clicking the side panel, capturing pages, showing the wiki.
{cmd: 'stop_recording'}finalizes the.webmand saves to Downloads asamem-recording-<iso>.webm.
Why this matters
- Demos stay in sync with reality — the recording shows the current UI, not a stale screenshot
- Chrome Web Store listings can be refreshed in minutes
- Agents learn to operate amem by watching their own recordings
Status
The recording skeleton is shipped. The agent-driven recording scenarios are Day 2 work. The complete workflow ships alongside the Chrome Web Store submission.
amem capture
Download a paper or PDF into ~/.amem/raw/ and generate a cite key.
Usage
amem capture <url> [--doi <doi>] [--cite-key <key>]
Arguments
<url>— arXiv URL, DOI, PDF URL, or local file path--doi <doi>— override DOI lookup (useful when a PDF URL doesn’t resolve)--cite-key <key>— override the auto-generated cite key
Examples
amem capture https://arxiv.org/abs/1706.03762
amem capture https://example.com/paper.pdf
amem capture 10.1038/nature14539 --doi 10.1038/nature14539
amem capture ./local-paper.pdf --cite-key smith2024local
Output
Prints the cite key on success. Exits non-zero on failure.
See also
- amem compile — next step after capture
amem compile
Parse a captured source, chunk with SHA-256 provenance, run Ollama paraphrase, and emit a wiki note.
Usage
amem compile <cite_key>
Environment
AMEM_OLLAMA_MODEL— override the paraphrase model (default:llama3.1)
Output
Writes ~/.amem/wiki/{timestamp}_{cite_key}.md and updates ~/.amem/index.md.
Example
$ amem compile vaswani2017attention
Chunked into 47 chunks. Ollama paraphrase 47/47 ok.
Wrote ~/.amem/wiki/20260418_vaswani2017attention.md
See also
- amem recall — search compiled notes
amem recall
Grep-search compiled wiki notes and return excerpts.
Usage
amem recall <query> [--limit N]
Arguments
<query>— free-text search term--limit N— max results (default 10)
Example
$ amem recall "attention mechanism" --limit 3
vaswani2017attention (sha256:a1b2c3…): "…the Transformer, based solely on attention mechanisms…"
bahdanau2014neural (sha256:d4e5f6…): "…a neural machine translation model learns to pay attention to…"
See also
- MCP Tools — same backend exposed to agents
amem cite
Print a formatted citation for a captured source.
Usage
amem cite <cite_key> [--format bibtex|apa|mla|chicago|ieee]
Default format: bibtex.
See also
- Citation Formats — full format reference
amem mcp
Subcommands for the MCP stdio server and its Claude Code registration.
amem mcp serve
Start the MCP stdio server. Exposes amem_capture, amem_compile, amem_cite, amem_recall to any MCP client.
amem mcp serve
amem mcp install
Register the current amem binary with Claude Code at user scope. Idempotent — safe to re-run; it removes any stale registration first. Requires claude on PATH.
amem mcp install
install.sh calls this automatically if Claude Code is installed.
amem mcp uninstall
Remove the registration from Claude Code.
amem mcp uninstall
See also
- MCP Tools — tool schemas
- MCP Protocol — implementation details
Philosophy
Wiki as substrate
Karpathy argues that a personal wiki — verbatim quotes, paraphrase, cross-links — is the durable substrate of long-term thinking. Apps come and go; markdown with provenance stays readable for decades.
amem treats the wiki as the product, not a byproduct. Capture is just staging; compile is the commitment.
Offline-first, always
Cloud dependencies are a liability. amem runs entirely on a user’s Mac: Ollama for LLM, pdftotext for PDF parsing, local disk for storage. The only network calls fetch papers from their canonical URLs.
This matters for three reasons:
- Sovereignty — your knowledge base doesn’t evaporate when a vendor pivots
- Speed — local I/O beats round-trips
- Reproducibility — SHA-256 provenance only means something if the source sits next to the hash
Dual audience
MCP for agents, CLI + extension for humans — same store, one source of truth. If an agent cites a chunk, a human can read that exact chunk by grep-ing ~/.amem/wiki/.
Complement, don’t compete
aide orchestrates agents. amem gives them memory. Neither duplicates the other. [sync.memory] method = "amem" in an aide.toml is the contract.
No legacy
amem is a clean v1 rewrite of crossmem. No code was ported verbatim; every module was reconsidered against these principles. crossmem remains published as v0 for historical continuity.
MCP Protocol
amem implements the Model Context Protocol over stdio using the rmcp Rust crate.
Transport
- Stdio only (Day 1)
- HTTP/WebSocket bridge for the Chrome extension is Day 2 work
Server identification
server_name:amemserver_version: matches the crate version
Tools
See MCP Tools for input/output schemas.
Error conventions
- Malformed inputs →
InvalidParams - Missing cite_key →
ResourceNotFound - Ollama unreachable →
Internalwith hintstart ollama - File IO errors →
Internalwith the OS error
Session model
Each amem mcp serve process is one session. No state shared between sessions beyond what’s on disk under ~/.amem/.
Relationship to aide.sh
amem and aide are complementary layers of the same system.
aide — orchestration
aide.sh dispatches work between agents. An Aidefile defines each agent’s persona, budget, triggers, and skills. aide dispatch <agent> <task> spawns an isolated Claude Code process and returns a bounded summary.
amem — knowledge
amem captures, compiles, and serves references. An agent running under aide can call amem_recall over MCP to ground its reasoning in real sources.
Integration point
In aide.toml:
[sync.memory]
method = "amem"
conflict = "causal"
When this is set, aide uses amem as its cross-agent memory substrate. Agents share a wiki; a fact learned by one agent is available to all.
Division of labor
| Concern | aide | amem |
|---|---|---|
| Which agent runs? | ✓ | — |
| How much budget? | ✓ | — |
| What does the agent read? | — | ✓ |
| Where are citations stored? | — | ✓ |
| How are secrets gated? | ✓ | — |
| What’s the SHA-256 of chunk 7? | — | ✓ |
Using one without the other
- amem alone: CLI + extension + MCP. Useful for human-led research and single-agent setups.
- aide alone: works fine, just without shared knowledge. Agents reason from their own seed context.
- Both together: the full stack.
formace-00 — Zotero reference corpus (ops)
Set up 2026-08-16. Backend for issue #24 (Zotero data dir as the librarian’s reference corpus, fine-grained ref check).
Where things live
| What | Path | Disk |
|---|---|---|
| Zotero 7 program (140.12.0esr) | /mnt/storage1/users/ydwu/opt/zotero | big disk (19T) |
Zotero data dir (zotero.sqlite, storage/*.pdf) | /mnt/storage1/users/ydwu/zotero-data | big disk |
| Zotero profile (prefs) | ~/.zotero/zotero/5sun4qjy.default | system disk |
| systemd user unit | ~/.config/systemd/user/zotero.service | system disk |
| amem corpus (source of the import) | ~/.amem/raw (157 files, 78 with metadata) | system disk |
Nothing of size sits on / — it runs ~82% full (319G free of 1.8T).
Service
systemctl --user status zotero # active/enabled
systemctl --user restart zotero
journalctl --user -u zotero -n 50
loginctl enable-linger ydwu is set, so the unit starts at boot without a
login session. Zotero is a GUI app, so the unit runs it under xvfb-run.
Library contents
78 papers + 78 PDF attachments, imported from ~/.amem/raw/*.meta.json by
generating BibTeX with file = {…} fields and POSTing to
http://127.0.0.1:23119/connector/import. Cite keys match amem’s
(belrose2023eliciting etc.), so Zotero items and amem wiki nodes line up.
Read the library:
curl -s "http://127.0.0.1:23119/api/users/0/items?limit=200"
curl -s "http://127.0.0.1:23119/api/users/0/collections"
The local read API is off by default; it is enabled via
extensions.zotero.httpServer.localAPI.enabled in the profile’s prefs.js.
Security — this is a shared machine
formace-00 has 25+ accounts (a class of students). Three exposures were found and closed on setup; keep them closed:
- Zotero’s HTTP server on
127.0.0.1:23119has no authentication. That is fine on a personal desktop and a data leak here — on a multi-user box localhost is not a security boundary, so any local account could read the whole library. The unit gates the port by uid via an iptables owner match, added inExecStartPreand removed inExecStopPost(so it needs noiptables-persistentand cannot outlive the service). zotero-datawas created world-readable (drwxrwxr-x) → nowdrwx------.~/.amemwas also world-readable (drwxrwsr-x) → nowdrwx------. The 157 raw PDFs and every wiki node were readable by every account.
New directories under ~ inherit a group-writable umask on this host — check
permissions after creating anything that holds corpus or credentials.
Gotchas hit during setup
pkill -f "opt/zotero/zotero"kills your own SSH command, because the pattern matches the command line running it. The script dies silently mid-way. Usepkill -x zotero-bininstead.- Zotero’s Linux download is served as
.tar.bz2but is actually XZ;tar -xfauto-detects, an explicit-jfails. - First launch must be given a display (
xvfb-run) to createzotero.sqlite; it never exits on its own, so run it undertimeoutwhen you only want initialisation. - Do not read
zotero.sqlitewhile Zotero holds it. Prefer a better-bibtex.bibauto-export (plain text, no lock) as the metadata source, or copy the sqlite first.
Chunked reference layer (built 2026-08-16)
src/refindex.rs in amem-librarian. Paragraph-level chunks, each hashed, with
BM25 retrieval — no model, no GPU, so the pipeline is verifiable end to end
before embeddings go on top.
export PATH="$HOME/.cargo/bin:$PATH" # cargo is NOT on the default ssh PATH
AM=/mnt/storage1/users/ydwu/cargo-target/release/amem
$AM refindex build --bib /mnt/storage1/users/ydwu/zotero-data/amem-corpus.bib
$AM refcheck "<a sentence from your draft>" [--cite-key rafailov2023direct] [--limit 5]
State on formace-00: 78 refs → 12,751 chunks, 13 s, index at
~/.amem/refindex/chunks.jsonl.
What it is for: a hit carries cite_key#p<page>c<n>, the page, a SHA-256 and
the query terms that actually matched, so “this paper does not support that
sentence” becomes visible rather than a hunch. Measured on the Mac corpus, a
DPO claim scores 32.2 against rafailov2023direct and 5.9 when misattributed
to vaswani2017attention.
Content addressing immediately paid for itself — identical hashes across different keys exposed duplicate papers in the corpus:
| identical chunks | same paper filed twice as |
|---|---|
| 197 | wei2022chain = wei2022cot |
| 156 | dpo_2023 = rafailov2023direct |
| 97 | merrill2023expressive = merrill2024expressive |
Worth deduping before the corpus grows; 12,751 chunks carry only 12,129 unique hashes.
Cross-machine ref check (built 2026-08-16)
amem refserve — a separate service from clipper_bridge on purpose. That
one drives the user’s real logged-in Chrome, so exposing it off-host would hand
a remote caller the browser. This one is read-only over the chunk corpus, so it
is the only piece with any business on a non-loopback address.
| service | amem-refserve.service (user unit, linger already on) |
| bind | 100.86.146.58:7602 — the tailnet address, not 0.0.0.0, so the lab LAN cannot see it |
| auth | bearer token, constant-time compare; token in vault as AMEM_REF_TOKEN, on the box at ~/.config/amem/refserve.env (mode 600) |
| guardrail | the server refuses to start on a non-loopback bind with no token, or a token under 16 chars |
From the Mac:
export AMEM_REF_TOKEN=$(bash ~/.claude/skills/vault/run.sh get AMEM_REF_TOKEN)
amem refcheck "<sentence from your draft>" --remote http://100.86.146.58:7602 \
[--cite-key rafailov2023direct] [--limit 5]
Verified from the Mac: no token → 401, wrong token → 401, valid token →
12,751 chunks searched. --cite-key vaswani2017attention on a compositionality
claim correctly returns NO SUPPORT, which is the wrong-citation signal working
across the network.
GET /status is deliberately unauthenticated and deliberately boring
({ok, chunks, refs}) so a health probe reveals nothing about the corpus.
Bibliography chunks are excluded
Reference lists mention every topic in the field while asserting none of them,
and they score well — a claim phrased like a paper’s title matches that title
in every bibliography citing it. Observed live: querying the CoT paper’s own
title returned three other papers’ reference lists above any real text. Chunks
are now flagged by signal-counting (arxiv:/doi:/bare URLs weighted double,
et al., In Proceedings, and 2022./2022b. year stamps) and skipped at
search time. 16.5% of the formace corpus flags; after the fix the same query
returns the actual paper’s title page first.
Dense retrieval (built 2026-08-16)
BM25 alone only finds passages sharing words with the claim. Measured: “transformers
struggle to compose multiple reasoning steps” never surfaced dziri2023faith, the
paper that argues exactly that, because it says “compositionality” and “error
propagation”. With vectors, all three top hits are that paper — including its
“Error Propagations: The Theoretical Limits” section.
ollama pull nomic-embed-text # 768-d, ~274 MB
$AM refindex embed # 12,994 vectors in 4m11s on the RTX 6000
Vectors live at ~/.amem/refindex/vectors.bin (raw LE f32, fixed stride, ~38 MB)
with vectors.meta.json carrying a fingerprint of the corpus they were built
from. A mismatch is refused, not warned about — vectors out of step with the
chunks would attribute quotes to the wrong papers. Re-run refindex embed after
any refindex build.
Fusion is reciprocal-rank, not a weighted score sum: BM25 is unbounded and corpus-dependent while cosine is in [-1,1], so weighting needs calibration that drifts as the corpus grows.
Read the components, not the fused rank. RRF orders well but says nothing about whether a passage is support, so hits carry both:
| bm25 | cosine | |
|---|---|---|
genuine support (dziri2023faith#p9c5) | 16.6 | 0.82 |
unrelated paper (liu2024spinquant#p5c4) | 1.8 | 0.52 |
Metadata verification — CrossRef and DataCite (built 2026-08-16)
amem refverify answers the question refcheck cannot: does this citation
exist, and is the metadata right? (refcheck answers what it says.)
Which registry matters. arXiv registers preprints through DataCite
(10.48550/arXiv.*); CrossRef 404s them. Verified by hand: CrossRef returns 404
for 10.48550/arXiv.1706.03762 while DataCite returns “Attention Is All You
Need”, 2017, publisher arXiv. Routing everything to CrossRef made all 7 local
preprints report ABSENT — indistinguishable from fabricated references. With
prefix routing: 7/7, and 74/75 on formace.
An arXiv id is turned into an exact DOI, so most lookups are direct resolution rather than fuzzy search. Short titles additionally need the first author’s surname to corroborate: “Attention Is All You Need” is five tokens and matched a 2025 book chapter at 71% Jaccard, reported as a year error when it was the wrong paper entirely.
Result on the 78-ref corpus — every duplicate content hashing suspected, confirmed by DOI:
10.48550/arxiv.2201.11903 ← wei2022chain = wei2022cot
10.48550/arxiv.2305.18290 ← dpo_2023 = rafailov2023direct
10.48550/arxiv.2310.07923 ← merrill2023expressive = merrill2024expressive
Note on the last one: both records carry year 2023 and DataCite agrees, so only
the cite key string merrill2024expressive is wrong — the metadata never was.
Not done yet
- better-bibtex plugin — not installed; stable cite keys currently come
from amem’s own metadata, not from Zotero. The
.bibcurrently indexed is generated from~/.amem/raw/*.meta.json, not exported by Zotero. - Dedupe the three duplicate cite-key groups, now confirmed by DOI.
- Re-embedding is manual —
refindex buildinvalidates the vectors by fingerprint (correctly), but nothing re-runsrefindex embedfor you. - The clipper bridge is still loopback-only, deliberately — see the
cross-machine section. Only
refserveis exposed. - GUI curation — no VNC on the box. Adding/annotating items by hand needs either a VNC server or doing curation on the Mac and syncing.
OpenWiki / llm-wiki-agent interop research for amem-librarian
Research date: 2026-08-10. Read-only study. Both repos cloned to
/private/tmp/claude-501/-Users-ydwu-claude-projects-amem-hq/09fb9b1e-7475-4775-aad3-0eeb124d78aa/scratchpad/.
0. Headline finding
OpenWiki does not have a format of its own. It emits Google’s Open Knowledge Format (OKF). The interop question is therefore not “should amem talk to OpenWiki” but “should amem emit OKF”. Those are different decisions with different risk profiles, and the second one is the one worth making.
Second finding: OpenWiki shipped a personal mode with Gmail / X / Slack /
Notion / Hacker News / web-search connectors writing to ~/.openwiki/wiki.
That is the same product shape as amem — sensors feeding a local agent wiki.
OpenWiki is now a competitor, not just an interop target. Details in §6.
1. Repos found
| OpenWiki | llm-wiki-agent | |
|---|---|---|
| URL | https://github.com/langchain-ai/openwiki | https://github.com/SamurAIGPT/llm-wiki-agent |
| Author | Brace Sproul (Head of Applied AI, LangChain) + LangChain | Anil Chandra Naidu Matcha (SamurAIGPT) |
| Version | npm openwiki 0.3.1 | no release; pyproject says 0.1.0 |
| Language | TypeScript (Node ≥22), React/Ink CLI | Python 3.10–3.13 |
| License | MIT | MIT |
| First commit | 2026-06-22 | 2023-04-21 (repurposed later) |
| Last commit | 2026-08-07 | 2026-08-03 |
| Commits total | 255 | 102 |
| Commits last 30d | 164 | 11 |
| Contributors | 60+; top 3 = Colin Francis 139, bracesproul 63, Brace Sproul 51 | 10; top 3 = Anil 32(+9+3), watsonk1998 23, Tony Lin 17 |
| Bus factor | ~2–3 (LangChain-backed, corporate) | ~2 (single maintainer + one heavy contributor) |
The ai.miraheze.org background page’s claim of >10k stars is consistent with
the repo’s activity, though star count is not verifiable from a clone.
Engine
OpenWiki is a DeepAgents documentation agent (deepagents +
@langchain/*). Model providers are pluggable and numerous — OpenAI (default
gpt-5.6-terra), Anthropic, Gemini/Vertex, Bedrock, Copilot, OpenRouter,
Baseten, Fireworks, Nebius, NVIDIA, Ollama, LM Studio, any OpenAI-compatible
gateway. Keys live in ~/.openwiki/.env. BYO key; no hosted LangChain service.
llm-wiki-agent has two execution paths: standalone Python scripts calling
litellm (LLM_MODEL env, default claude-3-5-sonnet-latest), or — the
headline mode — no API key at all, just Claude Code / Codex / Gemini CLI
reading CLAUDE.md / AGENTS.md / GEMINI.md and driving the workflow with
its own file tools.
2. Storage layout
amem (current, on disk)
~/.amem/
raw/ -> symlink to /Volumes/YDExtSSD/amem/raw
recordings/ -> symlink to external SSD
wiki/ 34 flat .md files, no subdirectories
Flat namespace, node id encoded in filename with a colon:
url:d5ffad083d4a.md, arxiv:1706.03762.md, x:AnthropicAI.md.
Namespace counts: url: 30, arxiv: 1, x: 1, plus 2 legacy files.
~/.amem/index.md does not exist, although SPEC.md §“Storage layout”
lists it as “auto-maintained”. No log.md. No config.toml either.
amem has two coexisting node schemas in one directory, which matters for any mapping work:
Schema A — legacy PDF/arxiv compile pipeline (1776567380_vaswani2017attention.md):
cite_key: vaswani2017attention
title: "Attention Is All You Need"
authors: [ ... ]
year: 2017
arxiv_id: "1706.03762"
raw: "/Users/ydwu/.amem/raw/1776567380_vaswani2017attention.pdf"
pdf_sha256: "bdfaa68d…"
chunks: 15
Body: # Title → ## Citations (APA/MLA/Chicago/IEEE/BibTeX) → ## Chunks
with per-chunk ### p1c1 (p.1) + `sha256:3b3e5b…`.
Schema B — clipper node (the format the task description gives, 32 of 34 files):
id: url:d5ffad083d4a
type: "page"
title: "ForMACE Lab"
url: https://…
host: "formace-lab.zulipchat.com"
first_seen_at: 2026-07-28T09:48:19.062Z
captured_at: [ …, … ] # list, grows on re-capture
tldr: "…"
tags: [amem-clipper]
Body: # Title → > tldr → **Source**: <url> → ## Content (verbatim
extracted text) → ## Links → ## Connections (empty in v0.2; the file
literally carries <!-- amem-clipper v0.2 leaves this empty; v0.3 populates with [[<node_id>]] wikilinks. -->).
Only Schema A carries SHA-256. The clipper nodes — 32 of 34 — have no hash at all, so the “provenance / SHA-256 excerpts” differentiator currently exists only on the legacy arxiv/PDF path.
OpenWiki
Two modes, two roots (src/config/openwiki-home.ts,
src/okf/index-sync.ts:66):
code mode: <repo>/openwiki/ (virtual root /openwiki)
personal mode: ~/.openwiki/wiki/ (virtual root /)
~/.openwiki/ mode 0700, ACL-restricted on Windows
wiki/ the knowledge bundle
connectors/<id>/
config.json never contains raw secrets, only env var names
state.json {version, lastRunAt, latestIds, runs[]}
raw/ raw dumps + manifests
logs/
conversation_history/
skills/ bundled SKILL.md dirs, synced on each run
.env provider keys
install-id telemetry install id
Nested directories are first-class: /sources/gmail.md, /topics/ai-research.md.
Every directory gets its own generated index.md.
llm-wiki-agent
Repo-rooted, not $HOME-rooted (tools/_utils.py):
raw/ immutable source documents, never modified
wiki/
index.md catalog, updated every ingest
log.md append-only chronological record
overview.md living synthesis across all sources
sources/ one page per source document (kebab-case.md)
entities/ people, companies, projects (TitleCase.md)
concepts/ ideas, frameworks, methods (TitleCase.md)
syntheses/ saved query answers
graph/ graph.json, graph.html, .cache.json, .refresh_cache.json
tools/ ingest/query/heal/health/lint/build_graph/refresh/…
One file per role, not per source: a single ingest fans out into one source page plus N entity pages plus M concept pages.
3. File format
OKF — the actual standard
Spec: GoogleCloudPlatform/knowledge-catalog/okf/SPEC.md. The published
spec is now v0.2. OpenWiki still emits and hardcodes v0.1
(src/okf/index-sync.ts renderIndex() writes okf_version: "0.1";
src/agent/prompts/personal.ts:97 instructs “MUST follow the Google Knowledge
Catalog OKF v0.1 schema”).
OKF v0.2 frontmatter:
| Field | Status | Shape |
|---|---|---|
type | REQUIRED | short string, uncontrolled vocabulary (“BigQuery Table”, “Playbook”, “Reference”) |
title | recommended | string |
description | recommended | one-sentence summary, retrieval-optimized |
resource | recommended | canonical URI of underlying asset |
tags | recommended | list of strings |
sources | optional | list of {resource(req), id, title, author, usage_count, last_modified} |
usage_window | optional | {from, to} |
generated | optional | {by: <actor>, at: <ISO8601>} |
verified | optional | list of {by: <actor>, at: <ISO8601>} |
status | optional | draft | stable | deprecated (default stable) |
stale_after | optional | YYYY-MM-DD |
| extensions | allowed | any additional keys |
Actor convention (§7): <producer>/<version> for tools, human:<id> for
people, process:<id> for automation. Consumers classify trust by the
human: prefix.
Structural rules:
- Reserved files
index.mdandlog.mdare not concepts; everything else.mdis a concept document. okf_versionmay appear only in the bundle-rootindex.md— the sole place frontmatter is allowed in an index.- Cross-links are standard markdown links, not wikilinks:
[label](/tables/customers.md)bundle-absolute, or[label](./other.md)relative. - Conformance (§11): parseable YAML + non-empty
type+ reserved-file rules. Consumers MUST NOT reject a bundle for unknown keys, unknowntypevalues, broken links, or missingindex.md. - Extensions (§10): “Consumers SHOULD preserve unknown keys when round-tripping and MUST NOT reject documents with unrecognized fields.”
v0.1 → v0.2 breaking changes (§13.1): timestamp → generated.at; body
# Citations section → frontmatter sources.
OpenWiki’s validator (src/okf/frontmatter.ts) enforces exactly the v0.1
subset: type required; type|title|description|resource|timestamp must be
non-empty strings when present; tags must be a list of non-empty strings.
It knows two of its own extension fields, openwiki_generated and
openwiki_translation_pending.
llm-wiki-agent
Schema lives in CLAUDE.md (also AGENTS.md, GEMINI.md) — the agent reads
it as instructions. It is prose convention, not a validator; nothing in
tools/ enforces frontmatter shape.
title: "Page Title"
type: source | entity | concept | synthesis # closed vocabulary
tags: []
sources: [] # list of source slugs (NOT OKF's structured sources)
last_updated: YYYY-MM-DD
Source pages additionally use date: YYYY-MM-DD and source_file: raw/….
Body convention for a source page: ## Summary → ## Key Claims →
## Key Quotes → ## Connections → ## Contradictions. Domain templates
exist for diary/journal and meeting notes.
Links are [[WikiLinks]] — tools/_utils.py:extract_wikilinks() is
re.findall(r"\[\[([^\]]+)\]\]", content). Resolution is by filename stem,
case-insensitive, ignoring directories (tools/ingest.py:validate_ingest,
tools/lint.py:page_name_to_path). Obsidian-compatible by design, with a
documented vault-symlink pattern in the README.
wiki/index.md is a hand-structured catalog with fixed sections (Overview /
Sources / Entities / Concepts / Syntheses); new entries are string-spliced
under the right heading by tools/ingest.py:update_index().
wiki/log.md: append-only, ## [YYYY-MM-DD] <operation> | <title>,
deliberately grep-parseable (grep "^## \[" wiki/log.md | tail -10).
Operations: ingest, query, health, lint, graph.
4. Update loop, dedupe, provenance
OpenWiki
openwiki personal --init/--update, oropenwiki ingest <connector|all>.- Connector tools (deterministic, no model) fetch and write
~/.openwiki/connectors/<id>/raw/plus a manifest, and updatestate.jsonwithlastRunAt/latestIds/ aruns[]summary. Incrementality is per-connector watermarks —latestIds, commit SHAs, Notion object ids + last-edited timestamps + content hashes. migrateWikiToOkf()runs before the agent, normalizing every concept page’s frontmatter so the agent sees a conformant wiki.- The DeepAgents agent reads raw dumps and writes/edits wiki pages.
synchronizeWikiIndexes()regenerates every directory’sindex.mddeterministically.validateWikiInternalLinks()stamps broken links inline.- Mermaid fences validated; failures downgraded to
textfences with a reason comment. - Post-run snapshot; unchanged runs record no new metadata (“no-op runs are free”).
Dedupe is agent-mediated — the model decides whether to edit an existing concept page or create a new one. There is no content-hash dedupe at the wiki layer and no stable node-identity field.
Provenance: source URI in resource; connector raw dumps retained on disk;
state.json run history. No hashes in frontmatter. No verbatim archive of
the source in the wiki page — pages are agent-synthesized prose. OKF v0.2’s
sources[] / generated / verified families are not emitted, because
OpenWiki targets v0.1.
llm-wiki-agent
python tools/ingest.py <file> (or “ingest raw/x.md” to Claude Code):
- Auto-convert non-
.mdvia markitdown (20+ formats incl. pdf/docx/pptx/ xlsx/epub/ipynb and wav/mp3 audio transcription). sha256(source_content, truncate=16)computed and printed.- Build context =
index.md+overview.md+ 5 most-recently-modified source pages (that’s the contradiction-detection window). - One LLM call returns a strict JSON envelope:
title,slug,source_page,index_entry,overview_update,entity_pages[],concept_pages[],contradictions[],log_entry. - Write pages, splice index, append log.
- Post-ingest validation: broken
[[wikilinks]]+ pages missing fromindex.md, both printed.
Dedupe/staleness: tools/refresh.py re-hashes each source_file and compares
to graph/.refresh_cache.json; only changed sources are re-ingested
(--force overrides). tools/build_graph.py caches by SHA-256 so only
changed pages get re-processed by the semantic pass.
Provenance: source_file: points at the immutable raw/ document, date:,
sources: [] slug list, plus the hash cache. Hashes are truncated to 16
chars and stored in a cache file, not in the page frontmatter — so a page
alone is not verifiable.
5. Entry points and MCP posture
| OpenWiki | llm-wiki-agent | amem | |
|---|---|---|---|
| Install | npm i -g openwiki | git clone + pip install -r requirements.txt | Rust CLI + MCP server |
| Serves MCP | No | No (zero MCP references in the repo) | Yes — amem_capture/compile/cite/recall/… |
| Consumes MCP | Yes — src/connectors/mcp-client.ts, full JSON-RPC over stdio + HTTP, openwiki_list_mcp_tools / openwiki_call_mcp_tool, backend: "mcp-http" | "mcp-stdio" | No | n/a |
| Agent integration | writes AGENTS.md/CLAUDE.md blocks between <!-- OPENWIKI:START/END --> markers | is driven by Claude Code/Codex/Gemini via CLAUDE.md + .claude/commands/*.md slash commands | MCP tools |
| API key | required (any of ~15 providers) | optional — none needed in Claude Code mode | n/a |
| Skills | bundles skills/mermaid-diagrams, skills/write-connector; syncBundledSkills() installs into ~/.openwiki/skills | .claude/commands/wiki-{ingest,query,lint,graph}.md | RFC-007 distillation (not yet shipped) |
This asymmetry is the interop lever. OpenWiki is an MCP client with a
generic mcp-stdio / mcp-http connector backend and an McpReadOnlyOperation
config. amem is an MCP server. An OpenWiki user can therefore configure amem
as a connector today, with zero code in either project — no format conversion
involved. That is a distribution channel, not a schema problem.
6. Competitive read (unrequested but load-bearing)
OpenWiki personal mode overlaps amem’s core pitch: local markdown wiki,
private (~/.openwiki, mode 0700), BYO key, sensors pulling from the user’s
tools, agent-maintained, MIT. With LangChain’s distribution and 164
commits/month it will not stay behind.
Where amem still differs, on the evidence:
- Real browser sensor. OpenWiki’s web reach is Tavily search and public
APIs. It cannot see a logged-in page. amem Clipper runs in the user’s real
Chrome session — Zulip DMs, X timeline, paywalled reading.
url:d5ffad083d4a(a Zulip login-gated page) is exactly the capture OpenWiki structurally cannot make. - iOS share-sheet sensor. No equivalent.
- Verbatim archive. amem keeps
## Content— the actual extracted text. OpenWiki keeps only agent-written synthesis in the wiki (raw dumps live separately underconnectors/*/raw/and are not the wiki page). - Hash-level provenance + citation formatting. OpenWiki has neither; llm-wiki-agent has truncated hashes in a side cache.
- Fact-check. Neither project has any grounding/verification tool.
amem_factcheck(RFC-003) has no counterpart in either codebase. - Skill distillation. OpenWiki ships skills to its own agent; neither project compiles user knowledge into reusable agent skills. RFC-007 is still unique.
Where amem is behind: index generation, link inference, staleness policy, audit log, schema self-healing, multi-format ingest, graph visualization. Section 9 lists what to take.
7. Field-by-field schema mapping
amem clipper node (Schema B) as the source of truth.
| amem field | OKF v0.2 target | llm-wiki-agent target | Verdict |
|---|---|---|---|
id (url:d5ffad083d4a) | none — keep as extension amem_id; optionally sources[].id | filename stem | lossy. OKF has no document-identity field; identity is the path. Colons in filenames also need sanitizing. |
type: "page" | type (required) | type: source | lossless but degraded. amem’s type encodes media kind; OKF’s encodes concept kind; llm-wiki’s encodes role. "page" is legal OKF but carries almost no routing signal — remap to Web Page / Paper / Profile. |
title | title | title | lossless |
url | resource | source_file (local path) | lossless to OKF. lossy to llm-wiki, which expects a repo-relative raw path, not a URL. |
host | none — extension amem_host | none | missing both; trivially re-derivable from resource. |
first_seen_at | none — extension, or sources[].last_modified (date-only) | date (date-only) | lossy: both flatten ISO-8601 datetime to a date, and neither has “first seen” semantics. |
captured_at[] (list) | none | none | lossy — worst case. Re-capture history is a list; OKF’s nearest list-valued time field is verified[], which means something else. Encode as extension amem_captured_at. |
tldr | description | ## Summary body section | lossless to OKF (rename). Structural move for llm-wiki. |
tags | tags | tags | lossless |
## Content (verbatim) | no standard section | no standard section | lossy. Both target formats expect synthesized prose. Survives only as an unrecognized body section — which conformant consumers must tolerate, but no consumer will understand. |
## Links | — | — | lossy, low value |
## Connections [[node_id]] | [label](/node_id.md) markdown links | [[PageName]] | lossless to llm-wiki (same syntax, resolved by stem). Requires rewrite for OKF — wikilinks are not OKF edges. Currently empty at v0.2, so this is a forward-looking cost, not a migration cost. |
Legacy Schema A (arxiv/PDF nodes):
| amem field | OKF v0.2 | llm-wiki-agent | Verdict |
|---|---|---|---|
cite_key | sources[].id | source slug | lossless-ish |
authors[] | sources[].author (single actor) | — | lossy — OKF’s author is one actor per source entry, not a list. |
year, arxiv_id | extension fields | — | lossy |
raw (local path) | sources[].resource (bundle-relative or absolute) | source_file | lossless |
pdf_sha256 | none — extension | cache only | missing from both standards. amem-unique. |
per-chunk sha256: | none | none | missing from both. amem-unique; the substrate amem_cite depends on. |
## Citations (APA/BibTeX/…) | v0.1 had # Citations; v0.2 retired it in favour of sources | — | churn risk: the one body section OKF standardized was removed in the very next version. |
chunks: 15 | extension | — | lossy |
Fields amem should adopt from OKF, not just map to:
| OKF v0.2 field | Why amem wants it |
|---|---|
verified: [{by, at}] | The standard’s designated slot for verification events. amem_factcheck output belongs here: verified: [{by: "amem/factcheck@0.3", at: …}]. amem would be the first producer populating the trust family with real verification rather than self-attestation. |
generated: {by, at} | Distinguishes clipper-captured from agent-synthesized nodes; human: prefix convention gives free provenance classing. |
status, stale_after | amem has captured_at[] but no staleness policy; stale_after is a re-capture trigger. |
sources[] | Structured multi-source provenance — a node captured from 3 URLs is currently inexpressible in amem’s single url field. |
8. Interop options
(a) Point the other tool at ~/.amem/wiki/ — symlink or config
Effort: S. Recommendation: reject.
Mechanically near-free, but it is not a read: OpenWiki mutates the directory it is pointed at.
migrateWikiToOkf()(src/okf/index-sync.ts) walks every.mdand callsnormalizeConceptContent(), which rewrites any page lacking a usabletype— replacing its frontmatter with a minimal derived block and stampingopenwiki_generated: true. amem’s Schema A nodes have notypefield at all, so every legacy arxiv/PDF node would have itscite_key,authors,year,arxiv_id,raw,pdf_sha256, andchunksdeleted on first run. Onlyopenwiki_translation_pendingis on the preservation list.synchronizeWikiIndexes()writes anindex.mdinto every directory.validateWikiInternalLinks()rewrites files to stamp broken links.- The agent then edits page bodies as prose — destroying
## Contentverbatim text, which invalidates every chunksha256.
Also: personal mode’s wiki root is hardcoded to ~/.openwiki/wiki
(openWikiLocalWikiDir), so “pointing” means symlinking that path at
~/.amem/wiki — full write exposure, no read-only mode.
llm-wiki-agent is gentler (its WIKI_DIR is repo-relative and its tools are
opt-in), and its [[wikilink]] syntax already matches amem. But it expects
wiki/{sources,entities,concepts,syntheses}/ subdirectories and a
hand-structured index.md; a flat directory of url:*.md files yields
“unindexed page” warnings for all 34 nodes and an empty graph.
Secondary blocker for both: colons in filenames. url:d5ffad083d4a.md is
illegal on Windows, ambiguous in some markdown link resolvers, and gets
percent-encoded to url%3Ad5ffad083d4a by OpenWiki’s index generator
(encodeURIComponent).
(b) One-way exporter amem export --okf — recommended
Effort: M (~2–4 days for a conformant v0.2 bundle).
Target OKF v0.2 the specification, and note in the docs that OpenWiki
currently reads v0.1. Write to a fresh directory (~/.amem/export/okf/), never
in place.
Transform:
tldr→description;url→resource;title,tagspass through.type: "page"→ a real concept type derived from the id namespace (url:→Web Page,arxiv:→Paper,x:→Social Profile).- Emit
generated: {by: "amem-clipper/0.2", at: <latest captured_at>}. - Emit
sources: [{resource: <url>, id: <node_id>, last_modified: <first_seen_at date>}]; for Schema A addauthorand therawpath, and carrypdf_sha256as an extension. - Preserve everything unmappable under
amem_*extension keys (amem_id,amem_host,amem_captured_at,amem_pdf_sha256,amem_chunks) — §10 guarantees consumers must not reject them. - Sanitize filenames:
url:abc.md→web-page/abc.md, i.e. use the namespace as a directory. Kills the colon and gives OKF the nested structure it expects. - Rewrite
## Connections[[id]]→[title](/web-page/abc.md). - Generate root
index.mdwithokf_version: "0.2"and per-directory indexes — reuse the algorithm in §9.1. - Emit
log.mdfromcaptured_athistory.
What breaks / what you accept: the export is a projection, not the node.
Chunk-level sha256 blocks and ## Content have no OKF home and survive only
as extension body — legal, but no OKF consumer will act on them. Re-importing
an OKF bundle edited elsewhere is explicitly out of scope (see (c)).
Cost is bounded because it is pure output — nothing in amem’s write path changes, nothing in amem’s schema is held hostage to OKF’s version churn, and if OKF v0.3 breaks again you edit one exporter.
(c) Bidirectional sync
Effort: L. Recommendation: reject, and say so in the docs.
Five independent reasons, any one sufficient:
- No stable identity in OKF. Identity is the file path. amem’s
idis the join key and has no standard home, so a page renamed by the other tool is unmatchable on the way back. - The other writer is an LLM. OpenWiki’s agent rewrites body prose
non-deterministically. Round-tripping mutates
## Content, which invalidates every chunksha256— the exact substrateamem_citeandamem_factcheckstand on. Sync would corrupt the differentiator. - Lossy in both directions. §7 shows
captured_at[], chunk hashes, and verbatim content have no OKF representation; a round trip cannot restore what the projection dropped. - No merge model. Two writers, no vector clocks, no CRDT, no conflict surface. Last-writer-wins over a knowledge base is silent data loss.
- Version skew is already real. OpenWiki writes v0.1 while the spec is v0.2, so a bidirectional bridge must translate between two versions of a moving format on every hop.
(d) Ship amem as an OpenWiki MCP connector — the cheap win
Effort: S (documentation only).
OpenWiki already speaks MCP as a client (src/connectors/mcp-client.ts,
backend: "mcp-stdio" | "mcp-http", readOnlyOperations). amem already
serves MCP. So an ~/.openwiki/connectors/amem/config.json pointing at
amem mcp serve with readOnlyOperations: [{type: "tool", name: "amem_recall"}]
makes amem a source OpenWiki reads from — no format conversion, no
exporter, no schema commitment.
Strategically this is the right shape: amem is the sensor-and-provenance
layer; OpenWiki becomes one more consumer of amem_recall. It also inverts
the competitive dynamic — instead of amem exporting into LangChain’s format,
LangChain’s tool queries amem’s API.
9. Ideas worth stealing
9.1 Deterministic per-directory index generation —
openwiki/src/okf/index-sync.ts:synchronizeWikiIndexes() /
synchronizeDirectory() / renderIndex(). Zero LLM calls; reads title +
description from each page’s frontmatter; sorts by href; writes only when
the rendered content differs from what’s on disk, so scheduled runs don’t
churn git. amem’s SPEC promises ~/.amem/index.md and it does not exist —
this is a ~100-line port and amem already has the two fields it needs
(title, tldr).
9.2 Normalize-before-agent, never reject —
openwiki/src/okf/frontmatter.ts:normalizeConceptContent() +
migrateWikiToOkf(). A non-conformant page is repaired deterministically
(derive title from first H1, else filename) and stamped
openwiki_generated: true; the agent later sees that flag and enriches it.
PRESERVED_EXTENSION_FIELDS carries control markers across the rebuild.
Directly applicable to amem’s own Schema A/Schema B split and to the
v0.2→v0.3 ## Connections migration: repair silently, mark for enrichment,
never fail a run.
9.3 Broken-link stamping instead of failing —
openwiki/src/agent/wiki-link-validator.ts:validateWikiInternalLinks().
Broken links get an inline <!-- openwiki: broken internal link … -->
comment; existing stamps are cleared at the start of each pass so they
never accumulate, and a later run finds the comment and repairs the href.
A self-healing loop that degrades instead of erroring.
9.4 Degrade-and-repair for generated artifacts — same pattern for Mermaid
(src/mermaid/validate.ts, README §Diagrams): an invalid diagram becomes a
plain text fence with a reason comment rather than a broken render; the next
update finds it and fixes it. Quality recovers over successive runs instead of
requiring a correct one-shot.
9.5 The health/lint cost boundary — llm-wiki-agent/tools/health.py vs
tools/lint.py, boundary documented as a table in CLAUDE.md. health is
deterministic, zero LLM calls, free, run every session (empty/stub files,
index sync, log coverage). lint is semantic, costs tokens, run every 10–15
ingests (orphans, broken links, contradictions, gaps). “Run health first —
linting an empty file wastes tokens.” This is the correct shape for an
amem doctor, and the explicit run-order rule is the valuable half.
9.6 Two-pass graph with typed edges —
llm-wiki-agent/tools/build_graph.py. Pass 1 parses [[wikilinks]] →
EXTRACTED edges (deterministic). Pass 2 asks the model for implicit
relationships → INFERRED (with confidence) or AMBIGUOUS. Louvain
community detection clusters topics; SHA-256 cache means only changed pages
are re-inferred; output is a self-contained graph.html. amem’s
## Connections is empty at v0.2 — this is a ready-made v0.3 design, and the
EXTRACTED/INFERRED distinction keeps model guesses auditable rather than
laundering them into the same link namespace.
9.7 Contradiction flagging at ingest time —
llm-wiki-agent/tools/ingest.py, the contradictions[] field of the JSON
envelope, checked against index.md + overview.md + the 5 most recent
source pages. Cheap precursor to amem_factcheck: catch conflicts when
writing, not at query time. Their own README frames it as the RAG
differentiator (“Contradictions surface at query time (maybe)” vs “Flagged at
ingest time”).
9.8 Append-only, grep-parseable log — llm-wiki-agent/wiki/log.md,
## [YYYY-MM-DD] <operation> | <title>, designed so
grep "^## \[" wiki/log.md | tail -10 is the read API. log.md is also an
OKF reserved filename, so adopting it is free conformance. amem has no audit
trail today.
9.9 Refresh-by-hash staleness — llm-wiki-agent/tools/refresh.py
re-hashes each source_file, compares against graph/.refresh_cache.json,
re-ingests only what changed. amem already stores captured_at[] but has no
policy that consumes it; combine with OKF stale_after for a re-capture
trigger.
9.10 Multi-format ingest via markitdown —
llm-wiki-agent/tools/ingest.py:convert_to_md() gets pdf/docx/pptx/xlsx/
html/epub/ipynb and wav/mp3 transcription from one dependency. amem’s
Pockist share-sheet would inherit a large format surface for very little code.
9.11 Secrets by reference, never by value —
openwiki/src/connectors/: connector config.json stores env var names;
values live only in ~/.openwiki/.env; ~/.openwiki is mode 0700 with
Windows ACL restriction (src/platform/windows-acl.ts). Matches RFC-007’s
“no secrets, use vault references” guardrail and shows the config-file shape.
9.12 .openwikiignore as a read boundary — gitignore syntax; when active,
it filters filesystem discovery and restricts shell execute. The README is
careful about what it does not promise (“does not guarantee a topic is
never mentioned, since the agent may still infer an ignored area from other
allowed evidence”). Good model for an .amemignore on the clipper, and good
copy for honest scoping language.
10. Risks
Schema churn — high. OKF v0.1 → v0.2 shipped two breaking changes
(timestamp→generated.at, # Citations→sources) inside roughly a month,
and one of them retired the only body section the format had standardized.
OpenWiki has not caught up: it validates and emits v0.1 today. Anything amem
builds against “OpenWiki’s format” is targeting a stale snapshot of a moving
spec. Mitigating factors: type is the sole required field, and §11 forbids
consumers from rejecting unknown keys — so a minimal export (type + title +
description + resource + tags) is very unlikely to break, while the rich
provenance families are where churn will bite.
Velocity churn — high for OpenWiki. 164 commits in 30 days, 255 total
since 2026-06-22, with active refactors landing in the exact modules that
matter here (#611 reorganized the CLI, #513 restructured the repo into
domain directories, #371 added the link validator). Any code-level coupling
will rot fast. Format-level coupling will not.
Bus factor. OpenWiki ~2–3 with LangChain behind it — low abandonment risk, high direction risk (it is a company’s product and will follow the company’s roadmap). llm-wiki-agent ~2, single-maintainer, 11 commits in the last month and most of them cosmetic star-history chores — treat as a design reference, not a dependency.
License — no obstacle. Both MIT. amem’s repos are private/proprietary; MIT permits use and derivation with attribution, so reading their code for ideas is fine and vendoring a file is fine with the notice retained. Implementing OKF creates no license relationship at all: a data format is not copyrightable, and OKF is published by Google as an open spec. The clean rule: implement the format, do not vendor the code.
Coupling cost. Option (b) is one output-only module with no upstream
dependency — its failure mode is “the export is stale”, which is cheap.
Options (a) and (c) put a competitor’s LLM agent inside amem’s write path;
their failure mode is silent corruption of the provenance data that
amem_factcheck is supposed to stand on. That asymmetry is the whole decision.
Strategic risk. Emitting OKF is also a small act of standard adoption in
LangChain-and-Google’s direction. It is worth it — the format is genuinely
better specified than anything amem would invent, and verified/sources
give the fact-check story a standard vocabulary — but amem’s internal schema
should stay amem’s, with OKF as an export target only.
11. Recommendation
Do (b) + (d): a one-way amem export --okf targeting OKF v0.2, plus a
documented recipe for registering amem as an OpenWiki MCP connector.
Rationale:
- The real standard is OKF, not OpenWiki. Build against the spec; treat OpenWiki as one consumer that happens to lag at v0.1.
- One-way export is bounded, reversible, and touches nothing in amem’s write path. If OKF v0.3 breaks, one module changes.
- OKF’s
verified: [{by, at}]andsources[]giveamem_factchecka standard vocabulary to publish into. amem would be the first producer filling the trust family with actual verification rather than self-attestation — that is a positioning asset, not just plumbing. - The MCP connector path costs a documentation page and inverts the dependency: LangChain’s tool queries amem, rather than amem exporting into LangChain’s world.
Order of work: steal 9.1 (index generation) and 9.5 (health/lint split) first — both are pure wins independent of any interop decision, and 9.1 closes a gap between SPEC.md and reality. Then the exporter. Then the MCP connector doc.
What amem should NOT do:
- Do not let OpenWiki write to
~/.amem/wiki/.migrateWikiToOkf()strips frontmatter from any page without atype— that is every legacy arxiv/PDF node, includingpdf_sha256andchunks. - Do not build bidirectional sync. §8(c).
- Do not adopt OKF as amem’s internal schema. It has no field for
captured_at[], no field for chunk hashes, and no document identity — the three things amem’s provenance model is built on. Export to it; don’t live in it. - Do not target OpenWiki’s v0.1 dialect. Emit v0.2 and let OpenWiki catch up; v0.1 consumers tolerate the extra keys by §11 conformance anyway.
- Do not vendor their code. Reimplement the patterns; the value is in the design, and both codebases are moving too fast to track.
- Do not chase connector parity (Gmail/Slack/Notion/X). That is LangChain’s strength and a treadmill. amem’s edge is the logged-in browser session and the iOS share sheet — sensors OpenWiki structurally cannot build — plus fact-check and skill distillation, which neither project has attempted.
amem 競品掃描 — 最接近的對手是誰
研究日期:2026-08-11。唯讀研究,未修改任何 project repo。
方法:先讀 amem 自家 SPEC/RFC/design 與 ~/.amem/wiki/ 實際產出,再做廣掃,再對前三名做原始碼層級驗證。
llm_wiki 與 llm_wiki_skill 已 clone 至本 scratchpad 逐檔查核。
0. 結論先講
最接近的競品是 LLM Wiki(nashsu/llm_wiki)。
它不是「同類產品」,而是同一張架構圖的另一個實作:Chrome MV3 clipper 抓當前分頁 → 本機 Rust daemon 用 LLM 兩段式摘要 → 寫成 Obsidian 相容的 markdown wiki → 內建 MCP server 讓 frontier model 取用並附引用。amem 的 SPEC「Mental model」四層,它四層都有。
更關鍵的是兩者同源。amem SPEC 的 storage layout 寫「wiki/ # compiled wiki notes (Karpathy style, agent-friendly chunks)」;llm_wiki 的 README 首段就寫「This project is based on Karpathy’s LLM Wiki pattern」。兩個產品在讀同一份 gist。這不是巧合式撞車,是同一個公開設計模式的兩個實作,而對方已經做到 16,173 stars。
前一份 OpenWiki 研究漏掉它,是因為那份研究從「interop 對象」出發,掃的是 LangChain 生態。llm_wiki 不在那條線上,它在 Karpathy gist 那條線上——也就是 amem 自己站的那條線。
1. 為什麼是它:逐項對照 amem 的 Mental model
以下每一列都經我親自 clone 原始碼查核,非讀行銷文案。
| amem SPEC 的層 | amem 現況 | llm_wiki 現況 | 證據 |
|---|---|---|---|
| Clipper(Chrome MV3 web sensor) | 有 | 有,MV3,activeTab+scripting+Readability.js+Turndown.js,Alt+Shift+L | extension/manifest.json、extension/clipper-core.js |
| 本機 daemon(Rust) | amem-librarian(Rust) | 有,Tauri v2 Rust backend,本機 HTTP API 127.0.0.1:19827(clip)與 :19828(API) | src-tauri/、mcp-server/README.md |
| AI 摘要編譯成 wiki | 有 | 有,two-step chain-of-thought ingest(先分析再生成) | README「3. Two-Step Chain-of-Thought Ingest」 |
| Obsidian 相容 markdown | 有(扁平 ~/.amem/wiki/) | 有,而且更完整——自動產生 .obsidian/ 設定 | README:79「the wiki directory works as an Obsidian vault」、README:370 |
| MCP 供 frontier model 取用 | 有(amem_recall/amem_cite/amem_ground…) | 有,bundled MCP server,10 個 tool | mcp-server/src/index.ts |
| 附引用的 grounded recall | amem_ground 回傳 hits + inline_md | 有,[1] [2] 編號引用 + Cited references panel + 引用持久化 | README:229、README:242-243 |
| 來源可追溯 | frontmatter url | 有,每頁 frontmatter sources: [] 指回 raw 檔 | README:132 |
| Raw 原件保留 | ~/.amem/raw/ | 有,raw/sources/ 明訂 immutable | README 目錄結構 |
| index.md / log.md | SPEC 承諾但檔案不存在 | 有,wiki/index.md + wiki/log.md 皆已實作 | README 目錄結構 |
| 蒸餾成 Claude skill(RFC-007) | 未實作 | 部分——見 §3,這格要講精確 | llm_wiki_skill/SKILL.md |
| iOS sensor | amem-pockist(TestFlight) | 無 | 全 repo 無 mobile |
| fact-check | RFC-003 規劃中 | 無 | 全 repo 無 claim verification |
成熟度:16,173 stars、1,918 forks,建立於 2026-04-08(僅 4 個月),最新 v0.6.8(2026-08-08),最後 push 2026-08-10,GPL-3.0(LICENSE 檔為 GPL v3 全文,作者 Yong Su;GitHub API 因檔頭多一行版權宣告而回 NOASSERTION)。桌面三平台 macOS / Windows / Linux。作者 nash_su。
四個月 16k stars,代表這個方向的市場需求已被驗證,也代表時間視窗正在關閉。
2. Top 6 競品對照表
軸線取「真正決定重疊度」的九項。「登入頁」欄位特別標註推論與明文的差別。
| LLM Wiki | SiYuan | Obsidian Clipper(+社群 MCP) | Karakeep | basic-memory | SurfSense | |
|---|---|---|---|---|---|---|
| Capture 面 | Chrome MV3,抓當前分頁 DOM | 官方 Chrome/Edge extension | Chrome/Firefox/Safari(含 iOS/iPadOS) | Chrome/Firefox/Safari + iOS/Android app | 無 | Chrome extension |
| 登入頁 | 結構上可(讀 live DOM),官方未明文 | 同上,未明文 | 同上,未明文 | 同上,未明文 | — | 舊 README 宣稱可抓 authenticated,現行 docs 已無此頁(404),無法驗證 |
| 抓自己的 AI 對話 | 無專用支援 | 無 | 無官方 template | 無 | — | 無 |
| Mobile share sheet | 無 | iOS/Android/HarmonyOS app | Safari extension,非 share sheet | 有 | 無 | 無 |
| 儲存 | 本機 markdown 檔 | .sy JSON(非 plain md) | 本機 markdown 檔 | SQLite/Postgres + assets | 本機 markdown 檔 | server 端 DB |
| Obsidian 相容 | 是,自動產 .obsidian/ | 否(僅單向匯入/匯出) | 原生 | 否 | 是 | 否 |
| 誰消費 | agent(MCP)+ 桌面 UI | agent(MCP)+ UI | 人(vault)/agent 靠第三方 MCP | agent(MCP)+ UI | agent 為主 | agent(MCP)+ web UI |
| MCP | serve(10 tools,內建) | serve + consume(官方,30+ tools) | 官方無;社群 mcp-obsidian 4,287★ | serve(官方 @karakeep/mcp) | serve(原生,20+ tools) | serve |
| Provenance | sources: [] + raw immutable + [1] 引用 | clipper 寫入原始 URL + 時間戳 | frontmatter 有 source URL/author/date | monolith 全頁存檔(verbatim 最強) | 無 source 追蹤 | 引用式回答 |
| 內容 hash | 僅圖片去重用 sha2,非 provenance | 無 | 無 | 無 | 無 | 無 |
| Fact-check | 無 | 無(web_search 是查資料非驗證) | 無 | 無 | 無 | 無 |
| Skill 迴路 | 消費 skill(/skill)+ 官方 access skill | 消費 skill(SKILL.md 機制) | 無 | 無 | 無 | 無 |
| License | GPL-3.0 | AGPL-3.0 | MIT | AGPL-3.0 | AGPL-3.0 | 未在 README 明示 |
| 平台 | mac/Win/Linux 桌面 | 桌面+行動+Docker | 瀏覽器 | 自架+行動 | CLI/MCP | Docker 自架+雲 |
| 商業模式 | 免費 OSS | Free / $64 買斷 / $148 年 | App 個人免費;Sync $4/mo | 免費自架 | 本機免費;雲端 $15/mo | 自架免費;雲 PAYG |
| 成熟度 | 16,173★,4 個月,v0.6.8 (2026-08-08) | 45,733★,v3.7.3 (2026-07-21) | 4,993★,1.7.1 (2026-07-22) | 28,249★,v0.33.1 (2026-08-01) | 3,627★,v0.22.1 (2026-06-13) | 15,876★,v0.0.36 (2026-08-06),自陳未達 production |
star / release 數據以 GitHub REST API 於 2026-08-11 查核。
表格讀出來的三件事:
- MCP 不再是差異點。 六家有五家 serve MCP,SiYuan 甚至雙向。amem 把「MCP 存取層」當賣點已經失效。
- 「本機 markdown + agent 可讀」也不再是差異點。 llm_wiki、basic-memory、Obsidian 生態都做到了。更要命的是 Claude Code 自己的 auto memory 就寫在
~/.claude/projects/<project>/memory/的明文 markdown,是免費預設值。 - 內容 hash 與 fact-check 兩格全業界皆空。 這是 amem 唯一真正無人佔領的象限——但見 §4,amem 自己也還沒真的佔住。
3. 對 LLM Wiki 的攻防拆解
3.1 幾乎完全重疊的部分
- capture 機制一模一樣。 兩者都是 MV3 extension 在使用者真實 session 讀 live DOM,經 Readability 抽取後 POST 到 localhost 的本機 daemon。amem 用 :7601,llm_wiki 用 :19827。連「登入頁能不能抓」的答案都一樣:結構上可以,因為 content script 跑在使用者已登入的分頁裡。
- compile 目標一模一樣。 LLM 摘要 → Obsidian 相容 markdown wiki → wikilink 知識圖譜。
- agent 取用方式一模一樣。 本機 HTTP API 包一層 MCP server 給 Claude Code / Codex。
- 知識來源三層架構一模一樣。 raw immutable → wiki → schema,因為都照 Karpathy gist。
3.2 amem 真正領先的地方
(a)iOS sensor。 llm_wiki 完全沒有行動端。amem-pockist 已上 TestFlight。share sheet 是行動端唯一低摩擦的捕捉入口,這格對方短期補不上(要做原生 app + 一套同步)。
(b)登入頁與自有 AI 對話是「明講的產品主張」而非副作用。 兩邊技術上都能抓登入頁,但 llm_wiki 從未把它寫成賣點,也沒有 per-site recipe。amem 的 design memo(2026-07-31)已經把 per-site recipe 定為要建的東西,且 ~/.amem/wiki/ 裡真的躺著 Gmail 搜尋結果頁與 Zulip 登入牆後的 capture。把「別人抓不到的頁面」做成明確能力,是 llm_wiki 沒佔的位置。
(c)chunk 級 SHA-256 + 格式化 citation。 llm_wiki 的 sha2 依 src-tauri/Cargo.toml 註解明寫是「for the dedup cache (Phase 3) — same image hash」,只用於圖片去重;agent workspace 追蹤用的是非密碼學的 DefaultHasher。amem 的 cite.rs 有 text_sha256 per chunk、pdf_sha256,並輸出 APA/MLA/Chicago/IEEE/BibTeX。全業界只有 amem 做這件事——但只在 PDF/arxiv 路徑上,見 §4。
(d)fact-check 是規劃中的產品核心。 llm_wiki 只有 Read Sources Only(限縮回答來源)與人工 Review 佇列,沒有對外部 trust list 驗證 claim 的機制。全掃描 18+ 個標的皆無。
(e)分發成熟度(潛在)。 llm_wiki 的 extension 沒有上 Chrome Web Store,README 只教 chrome://extensions → Developer mode → Load unpacked。對非開發者是硬門檻。但這一格 amem 目前也還沒兌現——amem-clipper(id jgknnaaaobkdggmjhhbklagidilmadii)CWS 查詢顯示 crx 0.3.0、無 published version、in review。要贏這格得先真的上架。
3.3 LLM Wiki 領先的地方
(a)規模與動能。 16,173★ / 1,918 forks / 4 個月 / 每兩天一個 build。amem 是單人私有 repo。這是最大的落差,且會自我強化。
(b)多格式 ingest 遠勝。 PDF(含 MinerU)、DOCX、PPTX、XLSX、EPUB/MOBI、影音、圖片 vision caption。amem 只有 web clip 與 PDF。
(c)wiki 維護機制成熟。 index.md / log.md / overview.md / Lint / Review 佇列 / cascade 刪除(含 dead wikilink 清理)/ 4-signal 知識圖譜 + Louvain 分群。amem 的 ## Connections 至今是空的(節點裡還留著 <!-- amem-clipper v0.2 leaves this empty --> 註解),SPEC 承諾的 index.md 也不存在。
(d)Deep Research 迴路。 偵測知識缺口 → 自動 web search(Tavily/SerpApi/SearXNG)→ 合成新 wiki 頁 → 自動 ingest。amem 沒有主動補洞能力。
(e)purpose.md。 讓使用者宣告「這個 wiki 為什麼存在」,LLM 每次 ingest/query 都讀。這是把使用者意圖變成可執行 context 的簡單好設計,amem 沒有對應物。
(f)scenario templates。 Research / Reading / Personal Growth / Business 各自預設 purpose.md 與 schema.md,降低冷啟動成本。
3.4 一個必須修正的 amem 內部說法
RFC-007 寫:「Every clipper competitor (Obsidian Web Clipper, Readwise, Notion) stops at storage; nobody closes the loop into agent capability. Step 3 is the moat.」
這句話現在有一半是錯的,需要改。
llm_wiki 已經把迴路收到 agent capability:MCP server + 官方發布的 llm_wiki_skill,npx skills add 一鍵裝進 Claude Code / Codex。所以「沒人閉環到 agent capability」不成立。
但精確地看,amem 的主張仍然活著,只是範圍小得多。 我逐字讀了 llm_wiki_skill/SKILL.md:它是人手寫的、唯讀的存取型 skill,內容是「怎麼呼叫 127.0.0.1:19828 的 API」,frontmatter 的 description 甚至明訂「DO NOT trigger on generic ‘search my notes’」。它讓 agent 讀 wiki,不是把 wiki 內容編譯成新 skill。llm_wiki 的「Agent Skills」功能同樣是掃描並啟用既有 SKILL.md——消費 skill,不生產 skill。
所以 RFC-007 該改成這樣:閉環到 agent 存取(MCP + access skill)已是紅海;把重複出現的程序性 capture 叢集自動編譯成新的 SKILL.md,目前仍無人做。moat 不是「step 3」整段,而是 step 3 裡的「自動生成」那半段。這個修正很重要,因為前者站不住,後者站得住。
3.5 amem 該怎麼做
定位語言
- 停用「local-first markdown wiki for agents」這類描述。llm_wiki、basic-memory、Obsidian 生態、Claude Code 內建 memory 全都符合這句話。
- 改押兩件別人結構上做不到或沒做的事:(1)別人抓不到的頁面(登入牆後、自己的 AI 對話、行動端 share sheet);(2)記憶可被稽核(chunk hash + verbatim + 格式化引用 + fact-check)。
- 「sensor」這個字要繼續用,而且要講滿。競品的入口是「你手動匯入的文件」,amem 的入口是「你本來就在讀的東西」。這是敘事上的真差異。
table-stakes 但目前缺席,必須補
~/.amem/index.md與log.md。SPEC 已承諾、對手已實作、與任何 interop 決策無關。前一份 OpenWiki 研究的 §9.1 已經給了可直接移植的演算法。## Connections真的填上 wikilink。空的 Connections 讓 wiki 退化成一堆孤立檔案,知識圖譜是這個品類的基本盤。- Lint / health 這類 wiki 維護迴路。
- 把 amem-clipper 真的推上 CWS。這是對 llm_wiki 少數幾個可立即兌現的優勢,卡在 review 就等於沒有。
不要做的事
- 不要追多格式 ingest 與知識圖譜視覺化的 parity。那是 llm_wiki 有 60 個貢獻者級動能的地方,單人追不上,且不是 amem 贏的理由。
- 不要把 MCP 當賣點寫在 landing page 第一屏。
4. amem 最沒防守的一塊
「verbatim archive + SHA-256 provenance」這個差異化主張,在實際出貨的路徑上不存在。
這是本次掃描對 amem 最不利、也最該立刻處理的發現。三個獨立證據:
~/.amem/wiki/34 個檔案中,32 個沒有任何內容 hash。 只有兩個 legacy PDF 節點(1776567380_vaswani2017attention.md、1776825486_belrose2023eliciting.md)帶pdf_sha256與 per-chunksha256:。所有url:*clipper 節點都沒有。- clipper 路徑的 SHA-256 是拿來算 ID,不是算內容。
amem-librarian/src/clipper_bridge.rs的sha12()把正規化後的 URL 做 SHA-256 再截前 12 個 hex,產出url:d5ffad083d4a這種節點 id。內容本身從未被 hash。 amem_ground不是 fact-check。 讀src/mcp/mod.rs:172的 tool description,它是「search the local wiki for the topic and return JSON hits」——對自有 wiki 的檢索加引用。SPEC 定義的核心差異化amem_factcheck(recall + 撥出去打 trust list + 驗證)仍是 RFC-003 規劃中。
也就是說:對外講的兩個核心差異(可驗證的 provenance、fact-check),一個只存在於非主力的 PDF 路徑,一個還沒開始。而主力路徑正是 clipper——它產出 32/34 的節點,是整個產品的入口。
附帶的品質風險:clipper 節點的 ## Content 雖然是 verbatim,但實測是未經整理的 DOM 傾印。url:e55e42200ccb.md(Gmail)裡是大段錯亂的表格標記與 。llm_wiki 用 Readability + Turndown 做過一輪清洗。「verbatim archive」若拿不出乾淨可引用的文字,對 fact-check 也撐不起來。
優先修的順序很清楚:先讓 clipper 節點帶 chunk 級 SHA-256 與乾淨的 verbatim 文字,再談 fact-check。 沒有前者,amem_factcheck 沒有可驗證的底層資料可站。這是整個 amem 論述的地基,而地基目前只鋪在 2/34 的面積上。
5. 亞軍與其他值得記錄的
SiYuan(45,733★, AGPL-3.0) — 功能面最完整的對手:官方 clipper(寫入原始 URL + 時間戳)、官方 MCP 雙向、AI Agent、SKILL.md 機制、三大行動平台官方 app。輸在資料主權敘事:.sy 是 JSON 不是 plain markdown,且非 Obsidian 相容,要匯出才拿得回來。amem 在「檔案就是你的」這點上贏它。
Obsidian Web Clipper(4,993★, MIT)+ 社群 mcp-obsidian(4,287★) — 「組裝出來的 amem」。官方 clipper 已有依網站自動套用 template 的機制,也就是 per-site recipe 概念已經存在於出貨產品裡。缺的是中間層:沒有常駐 daemon、沒有 AI 摘要編譯、沒有 provenance 層,且 capture 與 agent 存取由互不相關的專案提供。amem 的機會正是那個中間層。
Karakeep(28,249★, AGPL-3.0) — provenance 的另一種解法,且在「保真」這軸上其實比 amem 強:用 monolith 做全頁存檔對抗 link rot,加 yt-dlp 影片存檔。有官方 MCP、瀏覽器 extension、iOS/Android app。輸在儲存是 DB 不是 markdown,且無 Obsidian、無 agent 導向的 wiki 編譯。若 amem 要講「verbatim archive」,Karakeep 是必須被比較的對象。
basic-memory(3,627★, AGPL-3.0) — MCP 原生 + 本機 markdown + Obsidian 相容,agent-first 程度最高。完全沒有 capture 面,知識靠對話產生。它證明了「MCP + markdown」這半邊已被佔滿。
SurfSense(15,876★) — 曾以「儲存 authenticated 頁面」為 extension 主打,但現行 README 與 docs 已無該頁(/docs/browser-extension 回 404),無法從一手來源證實現況。官方自陳「not yet production-ready」。
Reor(8,573★,2026-03-07 已封存) — 值得引用的死亡案例。它有 local-first + 本機 LLM + 本機 vector DB 的完整組合,死在沒有 capture surface,匯入要人手動拷貝 markdown 進資料夾。這反向支持 amem 押 sensor 的判斷。
平台風險(非競品但更緊迫) — Claude Code auto memory 預設開啟,寫在 ~/.claude/projects/<project>/memory/ 的明文 markdown,使用者可直接編輯。「本機 markdown、agent 自寫、可稽核」已是 Anthropic 免費內建。它唯一缺的是 provenance——只記 modified 時間戳,不記來源。這再次指向同一個結論:amem 只能靠 provenance 與 sensor 活,不能靠儲存格式活。
6. 建議的定位句
amem 記錄你真正讀過的東西——包括登入牆後面那些——並且每一句話都能追回原文。
拆解:「你真正讀過的」= sensor(對手要你手動匯入);「登入牆後面」= 結構性優勢,且 llm_wiki 沒佔;「追回原文」= provenance 與 fact-check,全業界空白。整句刻意不提 markdown、不提 local-first、不提 MCP——這三個詞現在都無法把 amem 和任何人區分開。
但這句話目前只有前半是真的。 後半要成立,得先讓 clipper 節點帶上內容 hash 與乾淨的 verbatim 文字。在那之前,這是一句願景,不是一句產品描述。
7. 未能驗證的項目
- 登入頁擷取能力:llm_wiki、SiYuan、Obsidian Clipper、Karakeep 官方文件皆未明文陳述。本報告全部標為「結構上可推論」,因為 MV3 content script 在使用者已登入的分頁執行。
- SurfSense extension 現況:
/docs/browser-extension404,surfsense_browser_extension/README.md只有 Plasmo 建置指令。「儲存 authenticated 頁面」的說法來自舊版 README 與 fork 快照,無法確認是否仍為現行能力。 - llm_wiki 的 GitHub license 欄位回
NOASSERTION,但 LICENSE 檔內容為完整 GPL v3 全文加一行作者版權宣告。以檔案內容為準。 - llm_wiki 是否計畫做 mobile 或 fact-check:repo 內無 roadmap 佐證,僅能陳述「目前沒有」。
amem Clipper + Librarian — 使用者痛點 vs 產品現況
研究日期:2026-08-17。唯讀研究,未修改任何 repo,未發任何 issue/PR。
方法:先親自讀 amem 原始碼與 ~/.amem/ 實際產出,再挖四類一手來源。
一手來源為 GitHub issue tracker(10 個 repo)、HN Algolia API、forum.obsidian.md
Discourse API、Firefox AMO ratings API。前一份競品掃描(2026-08-11)的結論視為已知,
不重推。
0. 結論先講
amem 對「登入頁擷取」這個品類最大痛點答得很好,對「存下來的東西是否可信」答得最差。
第二點是致命的,因為它正是 amem 對外的定位句。SPEC 說 fact-check 是核心, 定位句說「每一句話都能追回原文」。實際出貨的 clipper 路徑做不到: 14 個節點在 6000 字截斷、沒有任何內容 hash、沒有 raw 備份、重剪直接覆蓋舊檔。
第三點是 RFC-007 的 moat 已經消失。不是變弱,是消失。Anthropic 官方文件明寫 產生 SKILL.md 不需要工具,Claude Code 二進位檔本身就會自動寫 SKILL.md。
1. amem 現況查核(我親自讀碼,非讀 SPEC)
這張表是後面所有判斷的地基。每一列都有檔名行號。
| 項目 | SPEC/定位的說法 | 程式碼實際行為 | 證據 |
|---|---|---|---|
| 內容 hash | 「每一句話都能追回原文」 | clipper 節點完全沒有 hash 欄位 | amem-librarian/src/clipper_bridge.rs:570-597 render_wiki_md 的 frontmatter 只有 id/type/title/url/host/first_seen_at/captured_at/tldr/tags |
| hash 涵蓋率 | — | 2/35 個 wiki 節點有 hash,都是 legacy PDF | ~/.amem/wiki/ 實測;30 個 url:* 節點與 1 個 clip:* 節點皆無 |
| verbatim 全文 | 「raw immutable → wiki」三層 | body 截斷在 6000 字,尾巴接 …(truncated) | clipper_bridge.rs:588 |
| 截斷實際比例 | — | 14/30 個 clipper 節點已被截斷 | grep -l '…(truncated)' ~/.amem/wiki/*.md |
| raw 層 | ~/.amem/raw/ 保留原件 | clipper 路徑不寫 raw。raw/ 只有 14 個 arxiv PDF 與 meta | background.js 只 POST /save_wiki_node,無 raw 寫入路徑 |
| 重剪語意 | 去重、防 re-clip spam | 保留 first_seen_at 與 captured_at 清單,然後整檔覆蓋 | clipper_bridge.rs:359-397,std::fs::write 覆寫;captured_at 上限 20 筆 |
| recall 檢索 | amem_recall / amem_ground | token 計數 grep,非 BM25、非向量 | src/query/grep.rs:1 檔頭自述「MVP search… Token-count scoring」;src/mcp/mod.rs:167,175 走 helpers → grep |
| BM25 + 向量 | 已建 | 已建,但只服務 Zotero 參考文獻語料,不服務 wiki | src/refindex.rs(948 行)、src/embed.rs(519 行),入口是 refindex build / refcheck |
| Connections | v0.3 填 wikilink | 31/35 個節點仍是字面 placeholder 註解 | <!-- amem-clipper v0.2 leaves this empty --> |
| index.md / log.md | SPEC 承諾 | 不存在 | ~/.amem/ 與 ~/.amem/wiki/ 均無 |
| AI 對話自動存 | manifest 宣傳「Rich extraction on Claude / ChatGPT / Gemini」 | 三個 autosave content script 都是 8 行 stub,只有 console.debug | extension/content-scripts/{claude,chatgpt,gemini}-autosave.js 檔頭自述「Day 1-2: stub only」 |
| AI 對話手動擷取 | 同上 | 真的可用。三個 extractor 各 200-230 行,含 artifact 與 binary 判別 | extension/extractors/claude.js 等;background.js:77-83 依 host 派送 |
| 離線容錯 | — | 無佇列、無重試。daemon 掛掉則該次擷取遺失 | background.js 唯一 alarm 是 amem-bridge-reconnect,只服務 bridge 輪詢 |
| 權限面 | — | <all_urls> + 12 個 permission,含 tabCapture、downloads、scripting | extension/manifest.json |
| CWS 上架 | 前次掃描判定「卡在 review」 | 已公開上線。v0.3.0,更新 2026-08-13,3 users,0 則評論 | CWS 公開頁實測;chrome-extension skill 的 status 判讀為誤報 |
兩件要更正前次掃描的事。
第一,amem-clipper 已經上架了。前次掃描說「卡在 review 就等於沒有」,
那是 CWS API 的 uploadState: NOT_FOUND 造成的誤判。公開頁面查得到 v0.3.0。
第二,「verbatim archive」比前次掃描的描述更糟。前次說是「未整理的 DOM 傾印」, 問題是髒。實際上更嚴重:47% 的節點被截斷,而且沒有 raw 備份可以還原。 截斷後的 markdown 就是唯一一份副本。
2. Part A — 痛點排序
排序依據是「頻率 × amem 目前答得多差」。答得好的排在後面。
| # | 痛點 | 頻率 | amem 現況 | 落差 |
|---|---|---|---|---|
| 1 | 靜默的部分擷取與保真度崩壞 | 7 repo,~90 comment,~70 reaction | 失敗 | 最大 |
| 2 | 存了找不回來 | 4 repo + HN dominant | 失敗 | 大 |
| 3 | 權限過寬引發的供應鏈恐懼 | 安裝決策點,有卸載實例 | 失敗,且無謂 | 大 |
| 4 | AI memory 錯誤、過期、無法驗證 | 2026 最熱,HN 180 pts | 部分 | 中 |
| 5 | 離線時擷取不進去 | ~92 reaction | 忽略 | 中 |
| 6 | 讀了不回頭的墳場 | HN 十年 dominant,tracker 僅 41 reaction | 忽略 | 中(但被高估) |
| 7 | 登入頁與付費牆擷取 | 全調查最高 reaction mass | 結構性解決 | 無 |
| 8 | 重複擷取 | 6 repo,兩位維護者宣告做不到 | 解決,但用資料遺失換 | 反轉風險 |
| 9 | 本機 LLM 連線摩擦 | 5 repo,~80 comment | 部分 | 小 |
| 10 | 上架被下架的存亡風險 | 有死亡案例 | 已上架 | 無 |
痛點 1 — 靜默的部分擷取與保真度崩壞(amem 答得最差)
陳述:擷取失敗最常見的形態不是報錯,是安靜地存了一個不完整或錯誤的東西。
一手證據:
kepano/defuddle#352(0 comment,2026-07-28,OPEN)—getElementSelector()未 escape 冒號 id,React streaming SSR 觸發後 Defuddle 自己拋錯,靜默退回抓整個<body>。這是「靜默部分擷取」的機制本體。 https://github.com/kepano/defuddle/issues/352obsidianmd/obsidian-clipper#196(4 comment,2024-11-21)— jwhitley 回報存到了 頁面上從不可見的隱私聲明,可見內文一個字都沒存。他自己說「只發生在某些頁面, 我找不出規律」。https://github.com/obsidianmd/obsidian-clipper/issues/196obsidianmd/obsidian-clipper#37(19 comment,32 reaction,CLOSED)— 「Download pictures to local」是該 repo 史上最高 reaction 的 closed issue。 遠端圖片連結會隨原站一起腐爛。https://github.com/obsidianmd/obsidian-clipper/issues/37karakeep-app/karakeep#1652(16 reaction,OPEN)+#1522(11 reaction)+#594/#999/#1306— 同一個 archive.org 中介需求被獨立提出五次,合計 27+ reaction。https://github.com/karakeep-app/karakeep/issues/1652deathau/markdownload#366(7 comment,OPEN)— 「Only half of an article is extracted」,多位獨立回報者。一位指出根因是 Readability.js,另一位量化 「砍掉大約 16 段」。https://github.com/deathau/markdownload/issues/366gildas-lormeau/SingleFile#1744(2 comment,OPEN)— Perplexity 與 ChatGPT 這類 對話視窗擷取被截斷。維護者:「It’s not an easy problem to deal with in a generic way.」https://github.com/gildas-lormeau/SingleFile/issues/1744- HN 44095939(bayindirh,2025-05-26)— 把需求講到最準:「Because I want the version I have seen. Not the edited/updated one.」 https://news.ycombinator.com/item?id=44095939
AMO 逐字評論獨立佐證同一件事。SingleFile 1033 則評分中 52 則低星,前三主題是
capture 不完整(4)、掉圖片與互動元素(3)、慢(2)。MarkDownload 有一則 3 星
(2023-06-17, alex):「it only converts what is visable on the screen rather than
the whole page」。https://addons.mozilla.org/en-US/firefox/addon/single-file/reviews/
amem 現況:失敗,而且是最壞的一種失敗。
三個獨立缺陷疊在一起:
- 6000 字截斷,14/30 節點已中彈(
clipper_bridge.rs:588)。 - 沒有 raw 備份,截斷後無法還原。
- 重剪整檔覆蓋,舊內容直接消失,沒有 diff、沒有版本、沒有 hash 可以偵測變化。
第三點值得單獨看。amem 用 URL 正規化把重剪折疊回同一個檔案,這解決了痛點 8。 但實作方式是覆寫,所以它同時製造了痛點 1。使用者兩週後重剪一個被改過的頁面, 第一版就永久消失了,而且系統不會告訴他。
對照 basicmachines-co/basic-memory#124(6 comment,4 reaction,OPEN)裡那句話:
「5 minutes into using it I realized I HAVE to have a mechanism for checkpoints…
the LLM could easily mess up the memory in one prompt, and you want to be able to go
back.」https://github.com/basicmachines-co/basic-memory/issues/124
這條落差最傷,因為 amem 的定位句押的就是這一格。全業界沒人做 chunk 級 hash, 這是真空象限。但 amem 目前只在 2/35 的面積上鋪了地基,而那 2 個還是不走 clipper 的 legacy PDF 節點。
痛點 2 — 存了找不回來(amem 答得差)
陳述:使用者的抱怨分兩種,檢索失敗與忘記自己存過。第二種更常被講。
一手證據:
- HN 22105561(279 pts / 268 comment,2020-01-21)— hooande 一句話定義了它: 「My problem with bookmarks isn’t managing them, but remembering that they exist.」同 thread 的 JohnFen 補上為什麼「再 Google 一次」不算替代方案: 「Most of the things I bookmark are things that were really hard to find in the first place.」https://news.ycombinator.com/item?id=22105561
- HN 44927588(jtqq,2025-08-16)— 自建 Emacs + org-roam + Elfeed + Wallabag 全套, 結論仍是「Retrievability is the biggest pain point right now」。下一步計畫是加 向量資料庫。https://news.ycombinator.com/item?id=44927588
karakeep-app/karakeep#2489(0 comment,2026-02-16,OPEN)— 「Full Text search STILL unavaliable… this seems as the most obvious use case」。六個月零回覆。 https://github.com/karakeep-app/karakeep/issues/2489karakeep-app/karakeep#1955(4 comment,OPEN)— 部分關鍵字搜尋回傳不完整結果。 https://github.com/karakeep-app/karakeep/issues/1955basicmachines-co/basic-memory#951(5 comment,OPEN,維護者自撰)— 有硬數字。 LoCoMo 1,986 個 query,single-hop R@5 basic-memory 0.486 vs mem0 0.592。 281 個獨有 miss 裡 146 是真 retrieval miss,主因是跨對話實體混淆。 https://github.com/basicmachines-co/basic-memory/issues/951nashsu/llm_wiki#5(15 comment,OPEN)— 50GB+ 語料庫每次開啟 loading 極久, 且作者給不出建議上限。#397(3 reaction)— 6566 頁重建索引要 2 小時以上。 https://github.com/nashsu/llm_wiki/issues/397
amem 現況:失敗,而且失敗的方式很諷刺。
amem_recall 與 amem_ground 走 src/query/grep.rs,檔頭自述是「MVP search…
Token-count scoring」。它把每個 wiki 檔整份讀進記憶體再數 token 命中次數。
35 個檔案可以,3,500 個不行。
諷刺的地方在於:BM25 加密集向量檢索已經寫好了。src/refindex.rs 948 行、
src/embed.rs 519 行,含 RRF 融合、bibliography chunk 過濾、向量指紋校驗。
但它服務的是 Zotero 的 78 篇 PDF 語料,不是 wiki。
也就是說,好的檢索建在對的技術上、錯的語料上。而 wiki 才是 clipper 的產物, 是產品入口。
~/.amem/ 沒有 index.md 也沒有 log.md,## Connections 在 31/35 個節點裡是空的。
所以「忘記自己存過」這半邊的痛點,amem 連最便宜的解法都沒有。
痛點 3 — 權限過寬引發的供應鏈恐懼(amem 答得差,且是無謂的差)
陳述:反對的不是 telemetry,是 <all_urls> 加 scripting 這個後門面積。使用者
會因此卸載,並且會寫下來。
一手證據:
karakeep-app/karakeep#2782(9 comment,2026-05-10,CLOSED)— 全調查最好的 單一物證。使用者 psla:「This permission is a no-go from me.」接著:「As much as I trust you and the project, supply chain attacks are a thing, and I generally don’t allow extensions that request full control over the page. Have you considered using optional permissions?」維護者 24 小時內改成optional_permissions並發 1.2.11。https://github.com/karakeep-app/karakeep/issues/2782karakeep-app/karakeep#2790(2026-05-12,CLOSED)— 另一位獨立提出,附威脅模型: 「its a great backdoor if repo ever gets compromised」,並標all_urls🔴 HIGH、scripting🔴 HIGH。https://github.com/karakeep-app/karakeep/issues/2790gildas-lormeau/SingleFile#1030(7 comment,OPEN 四年)— 「if I install this extension on Firefox, it will inject a script into every page I load even if I don’t intend to save the web page」。維護者答:必須永遠注入,無法關閉。 https://github.com/gildas-lormeau/SingleFile/issues/1030obsidianmd/obsidian-clipper#165(2024-11-14)— brian-burton 因 0.9.6 新增下載權限 而卸載,並在 template 欄位寫「Sorry, I can’t provide a template because I’ve uninstalled Web Clipper」。https://github.com/obsidianmd/obsidian-clipper/issues/165- AMO 逐字,Instapaper 1 星(2025-08-30,
dev421):「The permissions required by this add-on are ridiculous. “Access your data for all websites”!? You only need the current tab URL!」https://addons.mozilla.org/en-US/firefox/addon/instapaper-official/reviews/
amem 現況:失敗,而且有一半的權限沒有換到任何功能。
manifest 有 <all_urls> 加 12 個 permission,含 tabCapture、downloads、
scripting、tabGroups。這正是讓 psla 拒裝 karakeep 的權限輪廓。
更糟的是四個 always-injected content script 裡有三個是 8 行 stub。
claude-autosave.js、chatgpt-autosave.js、gemini-autosave.js 都只做一件事:
設一個 flag 然後 console.debug。它們對 claude.ai、chatgpt.com、gemini.google.com
永遠注入,換到零功能。
這是純負債。使用者付出「這個 extension 在我的 AI 對話頁上跑程式」的信任成本,
產品沒有拿到任何東西。真正的 AI 對話擷取在 extractors/ 裡,是按需
executeScript 派送的,不需要宣告 content script。
補一個對 amem 有利的觀察。維護者本人有時是隱私鷹派,而使用者站在他那邊。
kepano 在 #112 以安全理由永久拒絕雙向整合:「I don’t think most Obsidian users
would trust giving their browser full access to their Obsidian vault.」同 thread 的
writtenfool 附議:「I would not want, at anytime, for a web app to have access to my
local app.」amem 的 daemon 是本機的,這是可以講的故事,但前提是權限面要乾淨。
痛點 4 — AI memory 錯誤、過期、無法驗證(amem 部分答)
陳述:2026 年最熱的一條。抱怨不是記不住,是自動記下了錯的東西再拿去誤導使用者。
一手證據:
- HN 48776232「Memorizing session transcripts isn’t useful」(180 pts / 159 comment,
2026-07-03)。Fabricio20 最完整:「I specifically disabled claude memory in a
project because it kept writing down thigns to memory that didn’t need to be in
memory, including severly wrong statements that then would confuse it later.」
他還指出自動記憶被自動重新啟用,要同時關
autoMemoryEnabled與autoDreamEnabled。https://news.ycombinator.com/item?id=48776232 - 同 thread,mastax 指出污染的正回饋:「If you allow any low value things into memory, Claude will notice that established pattern and start trying to add low value memories」。semiquaver:「Claude’s own memories severely mislead it」。
- HN 47531882(ravikirany22,2026-03-26)— 唯一有量化的一筆:「We’ve been auditing TypeScript repos and finding 10-84% of symbol references in AI config files are stale… it’s getting a confident lie.」 https://news.ycombinator.com/item?id=47531882
nashsu/llm_wiki#458(2 reaction,OPEN)— 「Concerns About LLM-Based Knowledge Bases: Hallucinations and the Need for Faithful Original Text Retrieval」: 「When I query specific regulations or raw tables, the output is often hallucinated… frequently diverges significantly from the original content.」 https://github.com/nashsu/llm_wiki/issues/458nashsu/llm_wiki#229(5 reaction,OPEN)— 要一個修正入口。使用者指出唯一的修復 路徑是手改 markdown,而這與「Wiki 頁面全部由 LLM 維護」的核心設計理念相悖。 https://github.com/nashsu/llm_wiki/issues/229
這裡有一個對 amem 極重要的相關性,是 GitHub 挖掘得到的最有價值推論之一。 幻覺抱怨只出現在 llm_wiki 一家。 其他八個 repo 的 LLM 只寫 tag 與 summary, 不寫知識本體,所以沒有幻覺抱怨。llm_wiki 是唯一讓 LLM 撰寫 wiki 的,也是唯一 被抱怨幻覺的。
amem 的 clipper 節點把 LLM 產物(tldr)與原文(Content)並置。這個結構比 llm_wiki 安全。但這個安全性依賴原文真的在旁邊,而 47% 的節點原文被截斷了。
amem 現況:部分答,且答案正在漏氣。
有的部分:frontmatter 帶 url、first_seen_at、captured_at 清單。這比 Claude
Code 內建 memory 多(後者只記 modified 時間戳,不記來源)。tldr 與 Content 分離
也對。
漏的部分:沒有內容 hash,所以無法偵測來源頁面變了。amem_ground 的 tool
description 自述是「search the local wiki for the topic and return JSON hits」
(src/mcp/mod.rs:172),這是對自有 wiki 的檢索加引用,不是驗證。
SPEC 定義的 amem_factcheck 仍在 RFC-003 規劃中。
refverify(CrossRef + DataCite)與 refcheck(BM25 + 向量)是真的驗證,
而且做得好。它們只服務 PDF 語料。
痛點 5 — 離線時擷取不進去(amem 忽略)
陳述:擷取的那一刻經常是離線的,沒有工具會排隊。
這是本次調查最意外的一條,前次掃描與 design memo 都沒有列。
一手證據:
karakeep-app/karakeep#1077(9 comment,52 reaction,CLOSED)— 「Offline cache on Mobile app」。 https://github.com/karakeep-app/karakeep/issues/1077karakeep-app/karakeep#2401(17 comment,20 reaction,2026-01-14)— 自架者離開 LAN 之後,分享選單一直轉。 https://github.com/karakeep-app/karakeep/issues/2401karakeep-app/karakeep#274(3 comment,20 reaction)— 「Cache Hoards while away from server」。合計約 92 reaction。 https://github.com/karakeep-app/karakeep/issues/274obsidianmd/obsidian-clipper#828(0 comment,2026-05-05,OPEN 三個月無人回)— 同一個形狀的靜默資料遺失:「In my case I lost ~10 job-prospect clips this morning before realizing none of them had hit disk. There is no recovery path other than redoing the captures from browser history.」 https://github.com/obsidianmd/obsidian-clipper/issues/828
amem 現況:忽略,而且暴露面比 obsidian-clipper 更大。
background.js 唯一的 alarm 是 amem-bridge-reconnect,只服務 bridge 輪詢,
不是擷取佇列。沒有 retry、沒有 navigator.onLine 判斷、沒有 pending capture 儲存。
amem 的架構讓這件事比競品更常發生。obsidian-clipper 只需要 Obsidian 這個 app
存在。amem 需要 amem mcp serve 這個 daemon 正在跑。daemon 沒跑、剛重開機、
或 crash 了,使用者按下擷取就是失敗。而擷取失敗的當下,使用者已經離開那個頁面了。
痛點 6 — 讀了不回頭的墳場(amem 忽略,但這條被高估了)
陳述:痛感真實且橫跨十年,但它不是使用者會為之投票的痛點。
HN 證據極厚,判定 dominant:
- HN 44066646(rossant,2025-05-22,母 thread 1222 pts / 761 comment)— 「I just exported my data and found 13,000 unread articles out of a total of 34,000.」https://news.ycombinator.com/item?id=44066646
- HN 46880996(al_borland,2026-02-04)— 「They are where my good intentions go to die.」https://news.ycombinator.com/item?id=46880996
- HN 16306040 與 HN 44925963(pixelmonkey,相隔七年半講同一句)— 「I sometimes describe Instapaper as ‘/dev/null for web content’.」 https://news.ycombinator.com/item?id=44925963
- HN 36146108(bsnnkv,2023-05-31)— 「both note apps and read it later queues/apps are where ideas go to die.」https://news.ycombinator.com/item?id=36146108
但 tracker 的 reaction 質量說了另一件事。
resurfacing 相關的三個 issue 合計 41 reaction(karakeep #435 10、#705 16、
#863 15)。同一個 repo 裡,單一個登入擷取 issue(#172)就有 57 reaction,
離線佇列合計 92 reaction。
而且十個 repo 裡沒有任何一個有「我的存檔是墳場」這種 issue。它只以間接形式
出現,例如 karakeep#863:「we keep bookmarking the pages or videos, but we do not
have time to fully read it.」
這對 design memo 是一個修正。 memo 把 #3「write-only graveyard」與 #6「friction」 當成「真正殺死 clipper 的兩個痛」。HN 支持墳場的存在,但 tracker 說使用者不為它 投票,他們為「存不進去」和「找不到」投票。
amem 現況:忽略。 沒有 resurfacing、沒有 index.md、沒有隨機回顧,
## Connections 是空的。但依據上面的證據,這應該排在補完痛點 1 與 2 之後,
不該當頭條。
痛點 7 — 登入頁與付費牆擷取(amem 結構性解決)
陳述:server 端 crawler 永遠看不到登入使用者看到的東西。每個專案都獨立重新發現 唯一解是「從使用者自己的瀏覽器擷取」。這是全調查最高 reaction mass 的主題。
一手證據:
karakeep-app/karakeep#172(52 comment,57 reaction,CLOSED)— karakeep 史上最高 reaction 的 issue,「Local Scraper (Use browser auth)」。 結局是改成相容 SingleFile extension 的 REST endpoint。server 端 crawler 輸給了 瀏覽器 extension。 https://github.com/karakeep-app/karakeep/issues/172karakeep-app/karakeep#414(66 comment,40 reaction,CLOSED)— 存到 cookie 同意視窗而非文章。有使用者的本機 Llama3.2 去摘要了付費牆公告。 https://github.com/karakeep-app/karakeep/issues/414karakeep-app/karakeep#2814(4 comment,2026-05-16,OPEN)— 最關鍵的一筆。 karakeep 已經出貨 client-side crawling,付費牆仍然失敗:「Expected: Full article text is archived. Actual: Only the article preview/teaser is saved.」 他的 workaround 是用 Obsidian Web Clipper 抓,再寫 Python 腳本同步進 karakeep。 https://github.com/karakeep-app/karakeep/issues/2814karakeep-app/karakeep#2885(2026-06-13,維護者本人開)— 「As part of a recent reddit crackdown on crawlers, the endpoint that we were using… is now completely blocked.」https://github.com/karakeep-app/karakeep/issues/2885- 2026 年的新退化不只 Reddit:
#2952archive.is 又要 captcha(2026-07)、#2423Cloudflare 封鎖(2026-01)、#2381JS/cookie 攔截頁被存下來(2026-01)。
amem 現況:結構性解決,這是最該押的一格。
content script 跑在使用者已登入的分頁裡,沒有 cookie 轉移問題,沒有 server 再抓一次
的問題。~/.amem/wiki/ 裡真的躺著 Gmail 與 openreview 登入後才看得到的 capture。
趨勢對 amem 有利。bot wall 在 2026 年系統性收緊,server-side crawler 這條路正在 失效,而「已登入的真實瀏覽器」正在變成唯一可行路徑。
倫理阻力方面,一手來源查到的是零。使用者一律把它當純能力缺口。
最常見的自我正當化是「我付錢看的」(#2236,bronikowski:「when I read something
I paid for」)。
但這一格已經有人在賣了。 見 §3 的 Web2MD。
痛點 8 — 重複擷取(amem 解決,但用資料遺失換)
陳述:沒有工具會在你重存之前告訴你「你已經存過了」。而且有兩位維護者公開宣告 這件事做不到。
obsidianmd/obsidian-clipper#112(8 comment,16 reaction,OPEN)— kepano 是架構性拒絕:「For now there are no plans to create a two-way integration… This is intentional.… I think it poses a security risk.」#323與#521兩個同樣的請求都被 closed as duplicate。三次請求,一次永久拒絕。 使用者 CarcajadaArtificial:「I just want to not end up with duplicate web clippings. That’s all.」https://github.com/obsidianmd/obsidian-clipper/issues/112gildas-lormeau/SingleFile#1642(OPEN)— 維護者:「This is not really possible from a technical point of view, for privacy reasons. Extensions cannot scan folders on the filesytem.」#956同樣答「A database is required… it seems complicated to implement reliably」。 https://github.com/gildas-lormeau/SingleFile/issues/1642karakeep-app/karakeep#486(10 comment,7 reaction,CLOSED)— 十家裡唯一出貨 「已存過」指示器的,花了 13 個月。 https://github.com/karakeep-app/karakeep/issues/486karakeep-app/karakeep#864(3 reaction,OPEN)— 去重不處理結尾斜線,example.com/a與example.com/a/算兩筆。 https://github.com/karakeep-app/karakeep/issues/864
amem 現況:解決了,而且解得比業界好,但代價是痛點 1。
node_id_from_url(clipper_bridge.rs:425-471)做的事比 karakeep 多:
剝除 11 個追蹤參數、小寫 host、去尾斜線、arxiv/github/HN/x.com 各有專屬 id
規則。karakeep #864 抱怨的尾斜線問題,amem 在 url_canonical:521 已經處理。
#633 抱怨的追蹤參數,amem 在 :500-503 已經處理。
這是一個乾淨的勝場。兩位維護者公開說做不到的事,amem 因為有本機 daemon 而做得到。
但實作用覆寫達成去重。 舊內容消失,沒有版本、沒有 hash。所以 amem 把「重複
spam」換成了「靜默資料遺失」。這兩個痛點在 amem 身上是同一行程式碼的兩面
(clipper_bridge.rs:381 的 std::fs::write)。
痛點 9 — 本機 LLM 連線摩擦(amem 部分答)
陳述:「接你自己的模型」是 AI memory 工具的第一大死路。錯誤訊息很籠統, timeout 不可設定。
karakeep-app/karakeep#185(20 comment,OPEN 27 個月)— 標題就是 「How to verify hoarder app is working with the local ollama」。 https://github.com/karakeep-app/karakeep/issues/185karakeep-app/karakeep#424(19 comment,OPEN)— 使用者拿到的 log 是inference job failed: TypeError: fetch failed,沒有 host、沒有 status。 維護者:「Unfortunately the logging does not show where the issue happens.」 https://github.com/karakeep-app/karakeep/issues/424MODSetter/SurfSense有十個 open 的本機 LLM 連線 issue(#1379、#1616、#1550、#587、#517、#1464、#1405、#1394、#1493、#1518)。 https://github.com/MODSetter/SurfSense/issues/1379obsidianmd/obsidian-clipper#515(2 reaction,OPEN)— 「Custom Ollama provider requires non-existent API key」,UI 要一個該 provider 根本沒有的憑證。 https://github.com/obsidianmd/obsidian-clipper/issues/515- karakeep 有五個獨立 open issue 都是同一件事:本機模型比硬寫的 timeout 慢
(
#2770、#2679、#2994、#1129、#1806)。
amem 現況:部分答,而且方向對。
clipper_bridge.rs 的 summarize 有三條路徑:summarize_via_claude(:256)、
summarize_via_codex(:286)、summarize_via_ollama(:314)。有 CLI fallback 是
對的設計,使用者不必先裝 ollama 才能用。這比 SurfSense 與 karakeep 好。
未查證的部分:三條路徑全掛時的錯誤是否對使用者可讀。依據痛點 5 的分析, 擷取失敗沒有佇列,所以錯誤處理路徑值得單獨測。
痛點 10 — 上架被下架的存亡風險(amem 已上架)
陳述:clipper 的死亡證明是 Google 簽的,不是 bug 數量。
deathau/markdownload#378(9 reaction,OPEN)— 該 repo 最高 reaction 的 open issue,標題是「This extension is no longer available because it doesn’t follow best practices for Chrome extensions」。同一件事被獨立提報三次 (#3573 reaction、#3931 reaction),合計 13 reaction。3,997 star 的專案 死於下架。https://github.com/deathau/markdownload/issues/378#393裡有陌生人貼未審核的 fork,下一位留言者回「now the deployments website is unable, the zip is not found」。使用者被推向已經死掉的未簽名 fork。- HN 46880866(2026-02-04)— 「most of the things I come across are dead and gone, or seem abandoned somehow」。2026 年有三個 Show HN 把死亡寫進標題: 「because the others keep dying」(48745735)、「because Mozilla killed it」 (46956985)、「built after Pocket shut down」(48449568)。 https://news.ycombinator.com/item?id=46880866
amem 現況:已上架,前次掃描這一格判錯了。
CWS 公開頁查得到 amem Clipper v0.3.0,更新 2026-08-13,3 users,0 則評論。
llm_wiki 的 extension 至今仍要 chrome://extensions load unpacked。這一格 amem 贏。
順帶一個 null result:「load unpacked / Developer Mode」在十個 repo 裡是零筆
issue。 這不是真實使用者痛點。design memo 用 Developer Mode 摩擦當作不做
chrome.userScripts 的理由,那個結論仍然對(安全面與品牌面成立),但摩擦論據
本身沒有一手支持。
3. Part B — 前次掃描漏掉的競品
只列前次掃描的搜尋角度會漏掉的。星數與日期我親自用 gh api 於 2026-08-17 複驗。
3.1 Web2MD — 架構逐項複製,且已在賣
https://web2md.org/
今天出貨: Chrome extension、16 個站台專用 extractor(含 Reddit、YouTube、
GitHub、arXiv)、npx web2md CLI、npx web2md-mcp-server MCP server、
GPT-4 與 Claude token 計數、watch mode。
Agent Bridge 是核心賣點,官網原文:「Agent Bridge uses your actual Chrome with your cookies and login state — Reddit can’t tell the difference from normal browsing」。
這是 amem 痛點 7 那格優勢的逐字複製,而且已經包裝成產品在賣。 定價 Pro $4.17/月(原價標 $15)。免費層每日 3 次轉換。非開源。
單一勝出軸:CLI 批次同步。 amem 是一頁一頁點,Web2MD 有 web2md sync。
需注意:前次掃描漏掉它,是因為它不在 GitHub 上,搜 repo 搜不到。它也是
web2md.org SEO 內容農場的擁有者,該站產出偽裝成「honest review」的競品評測。
評論挖掘的子代理明確標記了這一點並拒絕採用其內容。
3.2 Letta — SKILL.md 自動生成,全自動且已出貨
https://github.com/letta-ai/letta-code — 3,017 star,v0.30.25 發佈於 2026-08-17(今天)。
我親自讀了 src/agent/subagents/builtin/reflection.md,逐字引用:
name: reflection description: Background agent that reflects on recent conversations to update memory and maintain skills
Skill generation/maintenance — ONLY when the conversation reveals a reusable, durable, multi-step workflow, create or update a skill under
$MEMORY_DIR/skills/.
Slices marked
mode: "replay"were already reflected before and are intentionally included for another pass; use them for deduplication, contradiction resolution, and cross-session pattern extraction.
最後一段特別重要。它殺掉的不只是「自動生成 SKILL.md」,還包括「跨 session 縱向 recurrence detection」這個備援縫。Letta 明文在做跨 session 模式抽取。
它還有 amem 沒設計的東西:skill 生命週期回收。操作集是 update / extend /
deprecate / split / create,寫進 MemFS git repo,有版本。
單一勝出軸:同一個 artifact,更完整的生命週期,已在有資金的產品內出貨。
它缺的是 capture surface。它的語料是對話,不是網頁。
3.3 OpenAI Computer History — 廠商層的 sensor → recurrence → skill
https://learn.chatgpt.com/docs/customization/computer-history
官方文件原文:「Computer History turns your activity across apps and websites into memories and a timeline that ChatGPT and Codex can reference.」以及 「When Computer History notices repeatable work, a timeline entry can suggest a skill or automation.」
這是「跨網站活動 → 偵測重複 → 建議 skill」,由模型廠商出貨。
限制就是 amem 剩下的縫:僅 macOS 桌面版、Pro/Business/Enterprise、預設關閉、 EEA/瑞士/英國不可用。而且它讀 interaction event(點擊、輸入、app 切換), 明確不讀頁面內容本身。
單一勝出軸:通路加作業系統級 sensor。
3.4 其他架構撞車的
| 名稱 | 出貨 | 成熟度(gh api 複驗) | 單一勝出軸 |
|---|---|---|---|
| Rowboat | Mac/Win/Linux 桌面 app,Apache-2.0 | 17,290 star,pushed 2026-08-17 | 論述撞車。README 寫「living Obsidian-style backlinked knowledge graph」「All data is stored locally as plain Markdown」 |
| claude-mem | CLI,v13.15.2 | 90,938 star,pushed 2026-08-17 | 「捕獲即記憶」的心智位置佔有量。無 skill 生成、無網頁擷取 |
| openhuman | dmg/exe/deb/AppImage | 36,323 star,created 2026-02-18 | 發版速度近乎每日 |
| open-knowledge | 桌面 app + CLI,GPL-3.0 | 3,480 star,created 2026-06-03 | 自述「AI-native markdown IDE and LLM wiki」,前端成熟度 |
| screenpipe | 桌面 | 20,980 star,YC S26 | 捕獲面總量最大,隨時可加瀏覽器 lane |
Rowboat 有一點要更正子代理的判讀。它的內建瀏覽器隔離於使用者主瀏覽器, README 明寫使用者必須在裡面重新登入。這是 amem 的優勢,不是 Rowboat 的。
3.5 格式層風險
GoogleCloudPlatform/knowledge-catalog — 8,667 star,created 2026-05-04,pushed 2026-08-15。Google 把「LLM 編譯的 markdown wiki」變成規格。子代理報告 v0.2 加入 trust signals,這直接踩進 amem 的 citation-grounding 差異化。該 v0.2 細節我未親自複驗。建議獨立評估 OKF 相容性。
3.6 負面結果(掃過、確認沒有)
這些空白是 amem 剩餘的可辯護空間。
- MCP 生態沒有 clipper。
punkpeye/awesome-mcp-servers(92,462 star)全文搜 clipper / web clip / clip page 零筆。modelcontextprotocol/serversREADME 無任何 clipper server。 - 大型 memory 玩家都沒有 capture surface。 graphiti、letta、memU、cognee、
txtai、memoripy、memento-mcp 的 repo 樹內都沒有 browser extension 目錄。
mem0ai/mem0-chrome-extension已 archived,最後 push 2026-03-23。 最大玩家退出了瀏覽器擷取賽道。 - browser-control MCP 沒有一個寫 KB。 chrome-devtools-mcp(49,285)、 browser-use(109,480)、playwright-mcp(36,195)、Skyvern、stagehand 全部純自動化、 零持久化。
- 「從我剪過的東西裡找跨篇反覆出現的模式」沒有人做。 搜過
chrome extension SKILL.md、web clipper skill agent、capture to skill、bookmarks to agent skill、recurring patterns into skills全部零相關結果。 - YC 四個 2026 batch 共 638 家,零家做「瀏覽器 extension + 本機 markdown wiki」。 (子代理用 YC 自家 API 分頁,我未複驗)
4. RFC-007 的 moat 判決:GONE
RFC-007 寫「Every clipper competitor stops at storage; nobody closes the loop into agent capability. Step 3 is the moat.」
前次掃描已經把它修正成「moat 不是 step 3 整段,而是 step 3 裡的自動生成那半段」。 那個修正現在也不成立了。 三份證據,全部我親自複驗。
證據一:Anthropic 官方說這件事不需要工具。
https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices 逐字:
Claude models understand the Skill format and structure natively. You don’t need special system prompts or a “writing skills” skill to get Claude to help create Skills. Simply ask Claude to create a Skill and it generates properly structured SKILL.md content with appropriate frontmatter and body content.
護城河不能建在供應商文件標註「此處不需要工具」的動作上。
證據二:Claude Code 二進位檔本身就會自動寫 SKILL.md。
我在本機 /Users/ydwu/.local/share/claude/versions/2.1.233 裡抓到字串,
逐字(去掉 UTF-16 間隔):
If this repo has no project verify skill (
.claude/skills/verify/SKILL.md), that is a reason to run/verify, not to skip it: the run creates that file, saving the working build-and-drive recipe for future sessions.
同一份二進位檔裡 run-skill-generator 出現 7 次,註冊為 bundled skill,
menuDescription 是「Create a skill that knows how to run this project’s app」。
/verify 寫檔是跑 verify 的副作用,使用者沒有要求產生 skill。
這就是無人在迴路的自動蒸餾,而且它是內建的、免費的。
證據三:Letta 已全自動出貨同型機制。 見 §3.2 的逐字引用。
證據四:「X → SKILL.md」已是 commoditized 類別。 星數我逐一用 gh api 複驗:
| 專案 | Star | 做什麼 |
|---|---|---|
| microsoft/SkillOpt | 16,074 | 把 SKILL.md 當可訓練參數優化,含 nightly self-evolution |
| yusufkaraaslan/Skill_Seekers | 14,773 | 18 種來源(docs/repo/PDF/EPUB/YouTube)→ SKILL.md |
| bergside/design-md-chrome | 2,663 | Chrome extension 剪一頁 → 產 SKILL.md |
最後一列要特別看。它已經佔住「Chrome extension 剪頁 → 吐 SKILL.md」這個一模一樣 的手勢。
還剩什麼
一條窄縫,是 feature 不是 moat:沒有人從「跨站累積數週的 clipped web pages」這個 語料出發做蒸餾。
三家的語料都不是網頁。Anthropic 的 capture surface 是螢幕錄影、當下對話、本機 repo。 Letta 的是對話 transcript。OpenAI Computer History 讀 interaction event,明確不讀頁面 內容。design-md-chrome 是一頁換一檔,沒有語料庫概念。
這條縫真實存在,但它撐不起「moat」這個詞。它撐得起一個功能。
建議:不要再把「產生 SKILL.md」寫進對外定位。
5. 建議:第一個該改的東西
先讓 clipper 節點停止破壞證據,再談任何新功能。
理由是這一個改動同時關掉排名第 1、第 2、第 4、第 8 四個痛點的落差, 而且它是 amem 唯一真空象限(chunk 級 hash provenance)的地基。
具體是三件小事,都在 clipper_bridge.rs 一個檔案裡:
- 移除 6000 字截斷,或把完整原文寫進
~/.amem/raw/。 目前截斷後無副本可還原 (:588)。 - frontmatter 加
content_sha256,並對 chunk 逐段 hash。cite.rs已有text_sha256,PDF 路徑在用,clipper 路徑沒接上。 - 重剪時若 hash 改變,另存版本而不是覆寫。 目前
:381直接std::fs::write。
做完這三件,「amem 記錄你真正讀過的東西,而且每一句話都能追回原文」才從願景變成 產品描述。在那之前,這句話的後半是不實陳述。
第二順位是把 refindex 的 BM25 加向量檢索接到 wiki 語料上。程式碼已經寫好了
(refindex.rs 948 行、embed.rs 519 行),只是指向錯的語料。這關掉痛點 2。
第三順位是刪掉三個 stub content script。這是零成本降低權限面(痛點 3), 它們目前換到零功能。
6. 取樣缺口與未查證項目
誠實揭露,這些不要當成已經查過。
- Reddit 完全未取樣。
www.reddit.com、old.reddit.com、api.reddit.com三個 domain 都被 harness 阻擋。r/ObsidianMD、r/PKMS、r/selfhosted、r/DataHoarder、 r/LocalLLaMA、r/ClaudeAI 全部沒有覆蓋。自架與 DataHoarder 族群的觀點缺口很大。 - Chrome Web Store 的 1-3 星評論全部拿不到。 CWS 從瀏覽器核心層禁止 content
script 注入
chromewebstore.google.com(bridge 回"The extensions gallery cannot be scripted."),這條路永久不可行,不是 bug。 WebFetch 只回傳預設「最相關」排序,幾乎全 5 星。受影響最大的是 Recall (getrecall.ai),它是 amem 最直接的 AI-summary 競品,負評完全沒有資料。 - AI 摘要品質主題嚴重取樣不足。 AMO 唯一有 AI 的樣本是 Raindrop 的 2 則。 不要從這 2 則推論任何結論。
- AMO 樣本有 Firefox 結構偏誤。 31 則登入/cookie 抱怨有相當比例是 Firefox 跨站 cookie 預設造成的,不能直接外推到 Chrome。
- 「AI 摘要在筆記庫裡幻覺」找不到第一手案例。 forum.obsidian.md 搜 「AI summary hallucinate」零筆。obsidian-clipper 的 Interpreter 相關 issue 全部是 整合層故障,沒有一筆抱怨輸出內容錯誤。含意:使用者對 agent memory 的幻覺很敏感, 對 clip 時的摘要幻覺還沒有痛感,可能因為原文還在旁邊。
- 付費牆擷取的倫理/ToS 抱怨:零筆。 這是 null result,不是沒查。
- notes/wiki → 自動生成 skills:完全沒有人要求。 使用者的行為是從 session 事後蒸餾 skill(HN 47543139,627 pts / 265 comment,ccosky:「Anytime I do something as a one-off that I know I’ll do in the future, at the end of the session I’ll ask Claude to write a new skill based on what it did」)。 方向與 RFC-007 的假設相反。
- 未親自複驗: OKF v0.2 的 trust signals 細節、YC batch 歸屬與家數、 Minibase 與 LLMnesia 的使用者數、Gemini CLI 那個「require recurrence evidence before extracting skills」PR 的編號。
- amem 三條 summarize 路徑全掛時的錯誤可讀性未測。 依痛點 5 的分析值得單獨測。
RFC-001 — Function-based v0.1 architecture
- Status: Draft
- Authors: @yiidtw
- Created: 2026-05-09
- Related: RFC-002 (Clipper skills catalog UI), RFC-003 (recording orchestration), archived RFC-001 (bridge-first), SPEC.md § Mental model
TL;DR
Ship v0.1 as hardcoded MCP tools and Rust functions, not a skill engine. The three must-have features all have the same shape — agent calls a tool, librarian dispatches to a Rust function, function may bounce DOM-side work through the bridge to amem Clipper. There is no plugin runtime, no DSL, no sandbox. Skill engine is deferred to v0.5+, gated on the rule of three: ship when we have 5+ site adapters, 3+ recording scripts, or external user requests for installable skills.
7.5 days of hardcoded code now beats 13 days of skill-engine plumbing for zero current users. We promote to a real engine the day the third copy of “this looks the same as the previous one” arrives.
Motivation
The previous round of RFCs (archived 2026-05-09) circled around a generic “skills” concept — pluggable units the agent could discover and invoke. The abstraction was attractive on paper (catalog, sandbox, manifest, capability flags) and load-bearing on nothing: we have one user (the operator), zero external testers, zero shipped sites, and a CWS deadline.
Three forces pushed us to lock the architecture as function-based:
- No third instance. Out of the v0.1 must-haves, only the site-adapter
family has any repetition (arxiv, github, hackernews — three sites,
one shape). Recording is one-of-one; bridge-driven Chrome ops are
one-of-one. The rule of three (Refactoring, Fowler) says abstract on the
third occurrence — we have it for adapters and only for adapters, and
even there the shape is “match URL → call extractor,” which is a Rust
matchblock, not a runtime. - Engineering cost is dispositive. A plausible v0.1 skill engine — manifest schema, loader, sandboxed JS runtime in the librarian, capability gating, registry, sidepanel discovery, error surfaces — is ~13 days. A function-based v0.1 is ~7.5 days (3 site adapters at 1d each, recording at 2d, bridge MCP wiring at 1.5d, polish at 1d). The 5.5-day delta is the entire CWS submission window.
- CWS positioning still works without an engine. The Chrome Web Store listing claims “the first agent-callable Chrome skills catalog” — the user-facing surface (RFC-002) renders three hardcoded skills as if they are an installable catalog. v0.1 ships the shape of the product, v0.2 ships the actual installability. The gap is honestly disclosed in-UI (“Custom skills coming v0.2”) so we’re not lying — we’re shipping the minimum that lets the rest of the story land.
The core insight: the operator (and Claude Code, when it’s the operator
acting on the user’s behalf) does not care whether chrome_navigate is
implemented as a Rust function or as a sandboxed JS skill. They care that
the tool exists and works. Build the tools first. Generalise on the third
occurrence.
Proposal
1. Component split — librarian (Rust) vs clipper (JS)
Two binaries, one bridge, one rule for splitting work:
┌────────────────────────────────────────────────────────┐
│ Frontier model (Claude / GPT / Gemini) │
└────────────────────────────────────────────────────────┘
│ MCP (stdio)
▼
┌────────────────────────────────────────────────────────┐
│ amem Librarian (Rust binary, formerly amem-sh) │
│ • All heavy work: capture pipeline, compile, fact- │
│ check, recording orchestration, site adapters, │
│ storage, OCR, screencapture │
│ • Hosts MCP server │
│ • Hosts bridge server (loopback WS 7600) │
└────────────────────────────────────────────────────────┘
│ ws://127.0.0.1:7600
▼
┌────────────────────────────────────────────────────────┐
│ amem Clipper (Chrome MV3 extension, JS) │
│ • DOM-only work: read/write document, observe URL, │
│ capture visible tab, render sidepanel UI │
│ • Holds zero authoritative state (config lives in │
│ librarian's config.toml) │
│ • No fetching, no parsing, no LLM calls, │
│ no file I/O — those live in the librarian │
└────────────────────────────────────────────────────────┘
Rule of thumb: if a step requires the file system, an LLM, ffmpeg, OCR, or a network fetch beyond the active tab, it belongs in the librarian. If it requires a DOM, the active page, or pixel-level capture of the active tab, it belongs in the clipper. Anything else (URL matching, string manipulation, scheduling) belongs in the librarian — the clipper is a sensor, not a brain (per archived RFC-001’s positioning, still in force).
2. Bridge protocol — action / params / id envelope
Single loopback WebSocket on 127.0.0.1:7600. All messages, both directions,
share one envelope:
{
"id": "<uuid-v4>",
"action": "<verb>",
"params": { ... },
"token": "<bridge token from ~/.amem/bridge.token>"
}
Replies use the same shape, with "action": "<verb>_result" and "id"
matching the request. Errors carry {"action":"error","id":"...","error":{ "code","message"}}. The same envelope handles librarian→clipper commands
(e.g. chrome_navigate) and clipper→librarian events (e.g. auto_capture).
Initial verb set for v0.1:
| Direction | Verb | Purpose |
|---|---|---|
| L → C | chrome_navigate | Navigate active tab to URL |
| L → C | chrome_click | Click DOM element by selector |
| L → C | chrome_extract | Read DOM (selector → text/html/attrs) |
| L → C | chrome_screenshot | tabs.captureVisibleTab (best-effort, see §6) |
| L → C | chrome_wait | Wait for selector to appear / URL pattern match |
| L → C | enter_recording_overlay | Toggle “recording” sidepanel state for demos |
| C → L | auto_capture | Sidepanel-toggled URL match fired; please ingest |
| C → L | invoke_skill | User clicked a Run-button in the sidepanel |
| L ↔ C | ping / pong | Liveness |
Security posture inherits from archived RFC-001 unchanged: loopback bind,
origin allow-list (only our extension IDs), 32-byte token in
~/.amem/bridge.token (mode 0600), token rotates on librarian restart.
3. The three must-have features → tool + function map
Each feature is one MCP tool (or a small fixed set), each backed by one Rust function in the librarian. No registry, no plugins.
3a. Claude Code operates already-logged-in Chrome
MCP tools shipped:
chrome_navigate(url: string) -> { status, final_url }
chrome_click(selector: string, opts?: { wait_after_ms }) -> { ok }
chrome_extract(selector: string, kind: "text"|"html"|"attrs") -> { value }
chrome_wait(selector: string, timeout_ms: number) -> { ok }
chrome_screenshot(selector?: string) -> { png_path }
Implementation in librarian:
#![allow(unused)]
fn main() {
// crates/amem-librarian/src/tools/chrome.rs
pub async fn chrome_navigate(url: &str) -> Result<NavigateOut> {
let req = BridgeReq::new("chrome_navigate", json!({ "url": url }));
BRIDGE.send(req).await?.into()
}
}
The librarian is just a thin pass-through here; the clipper does the actual
DOM work in chrome.scripting.executeScript. This is the only path —
agents do not get raw access to the bridge or to the extension. The MCP
boundary is the contract.
Why not use playwright / puppeteer / chrome-devtools-protocol? Because the
operator’s authenticated Chrome (Gmail, LinkedIn, internal tools) is the
whole point. CDP-based tooling either drops the user’s profile (Chrome 136+
--remote-debugging-port + --user-data-dir constraints) or breaks
keychain-backed flows (Chrome for Testing’s ad-hoc signing). Going through
amem Clipper, which lives inside the user’s real Chrome, is the only path
that doesn’t lose login state. (See ~/.claude/CLAUDE.md § Browser
Automation for the gory details.)
3b. URL-mentioned content auto-saves to wiki
MCP tool shipped:
amem_capture_url(url: string, mode?: "auto"|"force") -> {
cite_key, amem_uri, status: "captured"|"already_have"|"unsupported"
}
Two trigger paths feeding the same function:
- Agent-side: any agent (Claude Code, Cursor) that mentions a URL in
its reply calls
amem_capture_url(url)directly via MCP. The librarian matches the URL against the hardcoded adapter table and dispatches. - Browser-side: amem Clipper observes navigation events; if the URL
matches a hardcoded adapter pattern and the user has the matching
skill toggle on (RFC-002), clipper sends
auto_captureover the bridge, which calls the same function.
Hardcoded adapters in crates/amem-librarian/src/adapters/:
#![allow(unused)]
fn main() {
// crates/amem-librarian/src/adapters/mod.rs
pub fn dispatch(url: &Url) -> Option<Box<dyn Adapter>> {
match url.host_str()? {
"arxiv.org" | "www.arxiv.org" => Some(Box::new(arxiv::Arxiv)),
"github.com" => Some(Box::new(github::Github)),
"news.ycombinator.com" => Some(Box::new(hackernews::HackerNews)),
_ => None,
}
}
trait Adapter {
async fn extract(&self, url: &Url) -> Result<CaptureRecord>;
}
}
arxiv.rs, github.rs, hackernews.rs each ~150–250 lines, each a
straightforward fetch + parse. They share a tiny CaptureRecord struct;
they do not share a runtime. When the fourth adapter ships, we revisit
(see §6 “When to revisit”).
3c. Scripted Chrome-only recording
MCP tool shipped:
amem_record_demo(script_path?: string, script_inline?: string)
-> { amem_uri: "amem://recording/<uuid>", mp4_path }
Full mechanics live in RFC-003. From this RFC’s point of view it is one
more Rust function in the librarian — record_demo() — that:
- parses a YAML script,
- uses the bridge to drive the clipper through navigation/click/wait steps,
- simultaneously runs
screencapture -v -l<chromeWindowId>(macOS window-level capture) so the recording covers the whole Chrome window including the clipper’s sidepanel UI, - optionally hands the raw mp4 to video-use for post-processing.
3d. Catalog management (sidepanel needs to list things)
Two more MCP tools so the sidepanel and any agent can ask “what’s available”:
amem_list_skills() -> [{ id, name, description, kind: "auto"|"run", state }]
amem_invoke_skill(id: string, params?: object) -> { result }
For v0.1 these read from a hardcoded Vec<SkillCard> in the librarian.
There is no manifest. There is no sandbox. amem_invoke_skill("cws-demo")
literally calls record_demo(builtin_scripts::CWS_DEMO). The point is the
tool surface: the moment we have a skill engine in v0.5+, these two
tools’ signatures don’t change — only their implementation does. The
agent contract is forward-compatible.
4. MCP tool inventory for v0.1
Total: 6 new tools ship in v0.1 (plus the four existing ones from Day
1: amem_capture, amem_compile, amem_cite, amem_recall).
| Tool | Backing function | Notes |
|---|---|---|
chrome_navigate | tools::chrome::navigate | Bridge pass-through |
chrome_click | tools::chrome::click | Bridge pass-through |
chrome_extract | tools::chrome::extract | Bridge pass-through |
amem_capture_url | adapters::dispatch + capture | Hardcoded site match |
amem_invoke_skill | skills::run | Hardcoded skill table |
amem_list_skills | skills::list | Hardcoded skill table |
amem_record_demo is exposed as a kind:"run" skill via
amem_invoke_skill("cws-demo"), not as a separate top-level tool — keeping
the agent-facing surface tighter and letting RFC-003 own the contract.
5. Why NOT skill engine (yet)
Five reasons, in priority order:
- Rule of three. Of the three v0.1 features, only one (site adapters) has even three instances. The other two are one-of-one. Abstracting now means designing for assumptions we have not yet earned.
- One operator. A skill engine optimises for external authors shipping skills. We have zero external authors. The first user benefitting from sandboxing-vs-trust would be us — and we trust our own code more than we trust a sandbox we just wrote.
- 13 vs 7.5 days. A real skill engine needs: manifest schema with versioning, loader with capability gating, sandboxed JS runtime (boa/quickjs) with controlled host bindings, error surfaces, registry, discovery, version pinning, update flow. None of this work helps the v0.1 user-facing demo.
- CWS deadline lives in this window. Chrome Web Store review can take 2–10 days. Submit narrow + working > submit broad + half-built. The 5.5-day delta is bigger than the review buffer.
- Forward-compatible API.
amem_list_skills/amem_invoke_skillalready shape the surface a future engine will use. We are not painting ourselves into a corner; we are shipping the same MCP signatures that v0.5+ will reuse.
6. When to revisit — the rule of three
Promote to a real engine when any of the following triggers fire:
| Trigger | What it means | What to revisit |
|---|---|---|
| 5+ site adapters | The match on url.host_str() has grown to 5+ arms with similar shape | Extract a SiteAdapter trait + manifest table |
| 3+ recording scripts in production | YAML scripts have proven pattern: navigate, click, capture, repeat | Promote YAML schema to versioned spec; consider script-side templating |
| 1+ external user request | Someone outside yiidtw/ asks “can I write my own skill” | Engine becomes a user-facing feature, not internal cleanup |
| Custom skill in v0.2 plan firms up | RFC-002 promises “Custom skills coming v0.2” — the moment that ships | Engine is the implementation |
Until any trigger fires, the function-based approach is the correct endpoint, not a placeholder.
Privacy
Inherits the archived-RFC-001 posture unchanged:
- All capture data lives in
~/.amem/, never uploaded. - Bridge is loopback only, token-authed, origin-checked.
- The new MCP tools (
chrome_*) execute in the user’s real Chrome under the user’s existing permissions — they cannot access tabs the user is not already authenticated to. - Recording (RFC-003) uses macOS window-level screencapture targeting the Chrome window’s window ID; desktop and other apps are not in frame.
One new consideration: chrome_extract returns DOM content to the
librarian, which may surface to an agent over MCP. This is identical in
sensitivity to today’s amem_capture(url), which already fetches and
parses page content. Document the equivalence in the user guide.
Failure modes
| Mode | Cause | Mitigation |
|---|---|---|
Bridge unreachable when agent calls chrome_* | Librarian not running, or extension not connected | MCP tool returns structured error { code: "BRIDGE_UNAVAILABLE", install_hint }; agent surfaces install CTA |
| Selector not found | Page changed, login required, race | chrome_click / chrome_extract return { ok: false, reason } rather than throw; agent retries with chrome_wait |
| Adapter doesn’t match URL | URL outside hardcoded list | amem_capture_url returns { status: "unsupported" }; agent can fall back to the generic amem_capture (existing Day 1 tool) |
| Skill ID typo | Agent calls amem_invoke_skill("does-not-exist") | Structured error with available_ids list (read from amem_list_skills) |
| Concurrent recording + chrome ops | Recording and other MCP tools racing for the bridge | Recording acquires a recording-mode lock; concurrent chrome_* calls return { code: "RECORDING_IN_PROGRESS" } |
| Multi-tab / multi-window ambiguity | “Active tab” is ambiguous on multi-window setups | Default to focused-window’s active tab; expose tabId param later if it bites |
| Skill state drift between sidepanel and librarian | User toggles in sidepanel while CLI also toggles | Librarian is single source of truth (config.toml); sidepanel reads on every render via amem_list_skills |
Concrete work
In rough order of dependency. Estimates are pessimistic-realistic.
- (
amem-librarian) Bridge envelope + token + origin allow-list — ~1d - (
amem-librarian) MCP wrappers forchrome_navigate/_click/_extract/_wait/_screenshot— ~0.5d - (
amem-clipper) Bridge client + handlers for above verbs viachrome.scripting.executeScript— ~1.5d - (
amem-librarian) Adapters:arxiv.rs,github.rs,hackernews.rs- dispatch table +
amem_capture_urlMCP tool — ~2d
- dispatch table +
- (
amem-librarian) Hardcoded skill catalog (SkillCard, the three v0.1 entries) +amem_list_skills/amem_invoke_skillMCP tools — ~0.5d - (
amem-clipper) Sidepanel “Skills” tab rendering catalog (full detail in RFC-002) — ~1d - (
amem-librarian)record_demo()skeleton stub returningnot implemented— full implementation lives in RFC-003 — ~0.5d - (
docs.amem.sh)guide/skills.mddocumenting the v0.1 hardcoded skills + the v0.2 forward-compat note — ~0.5d
Total: ~7.5d. Compare to ~13d for an equivalent skill-engine v0.1.
Rejected alternatives
- Ship a real skill engine in v0.1. Costs 5.5d we don’t have, optimises for a user we don’t yet have. See §5.
- Ship skill engine manifest only, run hardcoded for v0.1. Tempting middle ground, but commits us to a manifest schema before we know what fields second-and-third skills will need. Premature schema lock.
- Embed skills as JS in the clipper extension. Moves heavy work into the extension (LLM calls, file I/O), which violates the clipper-is-sensor positioning and runs into MV3 storage / CSP limits. The librarian must stay the brain.
- Use playwright / puppeteer / Chrome DevTools Protocol for §3a. Loses the user’s Chrome profile (login state, passkeys, 2FA). The bridge-into-real-Chrome path is non-negotiable.
- Skip the catalog UI, ship MCP tools only. Loses the CWS positioning (“first agent-callable Chrome skills catalog”). The catalog is the user-facing story; without it we’re just another Chrome extension.
Open questions
- Multi-tab targeting. Should
chrome_*tools default to the active tab in the focused window (proposed) or accept an explicittabId? Soft preference: default-active for v0.1, exposetabIdwhen the second use case asks. - Adapter timeouts. What does
amem_capture_urldo if arxiv is slow / down? Soft preference: 30s timeout, returns{ status: "captured_partial", reason: "fetch_timeout" }so the agent can retry later. - MCP tool surface stability. The 6 new tool names are committed. We will add tools post-v0.1 but not rename or remove. The surface is the contract.
- Should
amem_list_skillsincludekind:"hidden"for skills used internally (e.g. by other tools) but not surfaced in the sidepanel? Open. v0.1 ships without; revisit when an internal skill emerges.
Roll-out
- Day 1 (today): this RFC + RFC-002 + RFC-003 land. SUMMARY.md updated.
- Day 2: bridge envelope +
chrome_*MCP tools + clipper handlers. - Day 3: site adapters land.
amem_capture_urlships. - Day 4: skill catalog + sidepanel UI (RFC-002 implementation).
- Day 5: recording skeleton lands (RFC-003 implementation).
- Day 6: end-to-end CWS demo recorded by amem itself, dogfooded.
- Day 7: CWS submission. Listing copy claims “first agent-callable Chrome skills catalog” honestly — three skills shipped, custom-skills disclosure visible.
- Post-v0.1: track the rule-of-three triggers in §6. When any fires, open RFC-00X for the skill engine.
RFC-002 — amem Clipper skills catalog UI (v0.1)
- Status: Draft
- Authors: @yiidtw
- Created: 2026-05-09
- Related: RFC-001 (function-based v0.1), RFC-003 (recording orchestration), archived RFC-001 (bridge-first), SPEC.md § Mental model
TL;DR
The amem Clipper sidepanel grows a second tab — Skills — that renders the v0.1 catalog as if it were an installable marketplace. v0.1 ships three hardcoded skills behind that UI: Auto-capture arxiv, LinkedIn inbox glance, CWS demo recording. The skill cards look identical to what v0.2’s actually-installable skills will look like, with one honest difference: a “Custom skills coming v0.2” disclosure underneath the catalog. CWS positioning leans on this UI: the listing claims “the first agent-callable Chrome skills catalog,” which is true — there is a catalog, each entry is agent-callable via MCP, and the catalog will accept custom entries in v0.2. We are shipping the shape of the product before the generality.
Motivation
RFC-001 commits to function-based v0.1 — hardcoded MCP tools in the librarian, no skill engine. That decision optimises engineering throughput to the CWS deadline. It does not, by itself, give us a CWS listing or a narrative.
The sidepanel UI is where the strategic positioning lives:
- Without a catalog UI, we are submitting “another Chrome extension that captures pages and integrates with an MCP server.” Crowded category; forgettable listing.
- With a catalog UI, the listing reads: “amem Clipper turns your browser into an agent-callable skills runtime. Three skills shipped — arxiv auto-capture, LinkedIn inbox glance, demo recording. Custom skills coming v0.2.” This is the differentiator. Nobody else is shipping Chrome skills the agent can list and invoke over MCP.
The catalog UI also pulls future weight: the v0.2 install flow is identical in structure to the v0.1 hardcoded path (skill card → toggle → bridge command). Users who learn the v0.1 model carry the mental model forward unchanged.
Proposal
1. Sidepanel tab strip
The sidepanel grows from a single capture surface to a two-tab strip:
┌────────────────────────────────────┐
│ amem Clipper ⚙ │
├────────────────────────────────────┤
│ [ Captures ] Skills │ ← tabs, [bracketed] = active
├────────────────────────────────────┤
│ │
│ (tab content) │
│ │
└────────────────────────────────────┘
- Captures — the existing capture log (recent items, filters, search). Default tab on first launch.
- Skills — the new catalog tab. The shape this RFC is about.
State is stored in chrome.storage.local; users land on whichever tab
they last closed.
2. Skill card schema (UI-level, not manifest)
Each skill renders as one card. The schema below is the render contract
between the librarian (which owns the catalog truth) and the sidepanel
(which renders it). It is not a manifest on disk in v0.1; it is the
shape returned by amem_list_skills over the bridge.
type SkillCard = {
id: string; // stable, e.g. "arxiv-autocapture"
name: string; // "Auto-capture arxiv"
description: string; // one-liner, ~80 chars
icon: string; // emoji or built-in icon name
kind: "auto" | "run"; // toggle vs button
state: SkillState;
badges?: string[]; // e.g. ["builtin", "v0.1"]
};
type SkillState =
| { kind: "auto"; enabled: boolean } // for auto skills
| { kind: "run"; busy: boolean; last_run?: string } // for run skills
| { kind: "error"; message: string };
Card layout:
┌─────────────────────────────────────────────────────┐
│ 📄 Auto-capture arxiv [ ●─── ] │
│ Captures every arxiv abstract you open. │
│ builtin · v0.1 │
└─────────────────────────────────────────────────────┘
- Top-right control depends on
kind:kind:"auto"→ toggle switch, bound to theenabledbooleankind:"run"→ “Run” button, disabled whilebusy=true, last-run timestamp under the buttonkind:"error"→ red text + “Retry” button
- Badges render as small tags at the bottom of the card.
- Cards are not reorderable in v0.1 (catalog order is hardcoded). v0.2 may add user pinning.
3. The three v0.1 skills
3a. Auto-capture arxiv (id: "arxiv-autocapture", kind: "auto")
Toggle-on means: any time the user navigates to an arxiv URL matching
arxiv.org/abs/* or arxiv.org/pdf/*, the clipper sends an auto_capture
event over the bridge with the URL; the librarian dispatches to the arxiv
adapter (RFC-001 §3b) and ingests.
Default: off. Auto-capturing every arxiv abstract is opinionated; users opt in.
UI behaviour:
- Toggle ON → small toast “Auto-capture armed for arxiv.org”
- On a successful capture → unobtrusive notification dot on the Clipper toolbar icon for ~5s
- On a duplicate (already captured) → silent (don’t spam)
- On error → clipper toolbar icon shows red dot, click for details
Implementation note: the toggle state lives in chrome.storage.local for
fast content-script reads; the librarian’s config.toml mirrors it as the
authoritative copy. Sidepanel reads truth from amem_list_skills on
every render; toggle writes go through amem_invoke_skill("arxiv- autocapture", { enabled: true }).
3b. LinkedIn inbox glance (id: "linkedin-inbox", kind: "run")
User clicks Run → librarian uses chrome_navigate to open
linkedin.com/messaging/, then chrome_extract to read the visible
inbox previews, then renders a digest in the sidepanel: who messaged,
unread count, first line of each thread.
This skill exists for two reasons:
- Demo value. It demonstrates that the agent can operate the user’s already-logged-in Chrome — the v0.1 differentiator. Without a visible flagship for that capability, the CWS reviewer has nothing concrete to evaluate.
- Honest utility. The operator actually uses this. It is not a throwaway skill written for the demo.
UI behaviour:
- Click Run → button disables, spinner, “Reading inbox…”
- Success → digest renders inline below the card; expandable
- Error (e.g. logged out) → “LinkedIn requires sign-in. Please open linkedin.com and sign in, then retry.”
3c. CWS demo recording (id: "cws-demo", kind: "run")
User (or the agent) clicks Run → librarian invokes record_demo() against
the built-in CWS demo script (RFC-003). The sidepanel shows a recording
state: live elapsed time, current step, “Stop” button.
This is the meta-skill: amem records its own CWS demo by driving its own
extension. The recording captures the entire Chrome window including the
sidepanel UI itself, which is why we need window-level macOS screencapture
rather than chrome.tabCapture (full reasoning in RFC-003 §5).
UI behaviour:
- Click Run → card expands with live status (current step, elapsed time)
- “Stop” cancels gracefully, finalises whatever was captured
- On finish → card shows
amem://recording/<uuid>link + path to mp4 - The recorded mp4 lands in
~/.amem/recordings/<uuid>.mp4
4. Auto-fire mechanism for URL-matched skills
The auto-capture flow lives in two places:
- Content script (clipper) — observes URL changes via
chrome.webNavigation.onCommitted. Pattern matches against the active set ofkind:"auto"skills withenabled:true. Patterns live inchrome.storage.local, synced from the librarian via the bridge on connect. - Background service worker (clipper) — receives match events,
forwards to bridge as
auto_capture.
Why store patterns in chrome.storage.local rather than asking the
bridge per navigation:
- Latency. URL change → bridge round-trip → match would add 50–200ms; unacceptable for ambient capture.
- Resilience. If the bridge briefly disconnects, the toggle behaviour shouldn’t change; the next reconnection re-syncs state.
Sync protocol on bridge connect:
clipper → librarian: { action: "skills_subscribe" }
librarian → clipper: { action: "skills_state",
params: { skills: [SkillCard, ...] } }
(thereafter, librarian pushes "skills_state" on any change)
Drift detection: every amem_list_skills MCP call also pushes the
current state to the clipper, so any out-of-band CLI toggle propagates.
5. Bridge command flow for Run-skills
User clicks Run on a kind:"run" skill:
sidepanel UI background.js librarian
│ user clicks Run │ │
├─ "invoke_skill",───────────▶│ │
│ id: "linkedin-inbox" } │ │
│ ├─ WS send ───────────────▶│
│ │ │
│ │ amem_invoke_skill("linkedin-inbox")
│ │ │
│ │ │ ┌─────────────┐
│ │ │ │ runs the │
│ │ │ │ skill → │
│ │ │ │ chrome_* │
│ │ │ │ over bridge │
│ │ │ └─────────────┘
│ │ ◀─ chrome_navigate /
│ │ chrome_extract pulses
│ │ │
│ │◀─ "skills_state",────────│
│ │ { busy: false, │
│ │ last_run: ... } │
│ "skills_state" forwarded │ │
│◀────────────────────────────┤ │
│ re-render card │ │
The skill execution itself is just an MCP tool call (amem_invoke_skill).
The librarian is the orchestrator; the clipper is the executor for DOM
side-effects. The sidepanel only kicks the kickoff and re-renders state.
6. CWS listing positioning
Listing copy (proposed for store page):
amem Clipper The first agent-callable Chrome skills catalog.
amem Clipper turns your browser into a runtime for Chrome skills your AI agent can list and invoke over MCP. Three skills ship today:
- Auto-capture arxiv — every paper you open lands in your local wiki
- LinkedIn inbox glance — your agent can read your inbox without you opening the tab
- CWS demo recording — amem records its own demos by driving its own extension
Custom skills coming v0.2.
Pairs with
amem-librarian, the local Rust binary that runs on your machine. Your data stays on your disk.
Three claims worth defending:
- “First agent-callable Chrome skills catalog.” True if “catalog”
means “a list of skills the agent can enumerate and invoke.” We have
amem_list_skillsandamem_invoke_skillover MCP; that is the catalog interface. v0.2 adds the user-installs-their-own dimension. The claim is honest with the v0.2 disclosure intact. - “Three skills ship today.” True; all three are real and tested.
- “Your data stays on your disk.” True; storage layout in
~/.amem/is unchanged; bridge is loopback only.
7. Honest disclosure: “Custom skills coming v0.2”
Below the catalog, a fixed footer:
┌─────────────────────────────────────────────────────┐
│ ✨ Custom skills coming v0.2 │
│ Bring your own scripts. ==AmemSkill== headers │
│ will install via drag-and-drop or URL paste. │
│ Read the v0.2 design intent → │
└─────────────────────────────────────────────────────┘
The “Read the v0.2 design intent” link goes to docs.amem.sh/skills/v0.2,
which renders the next section (§8) for transparency.
8. v0.2 design intent — Tampermonkey-style headers
When v0.1 hits a rule-of-three trigger (per RFC-001 §6) we ship a real skill engine. Sketch of the v0.2 user-facing format:
// ==AmemSkill==
// @id youtube-autocapture
// @name Auto-capture YouTube
// @description Saves every YouTube video you watch to your wiki
// @kind auto
// @match https://www.youtube.com/watch*
// @capability bridge:chrome_extract
// @capability librarian:capture
// @version 1
// ==/AmemSkill==
export async function onMatch({ url, ctx }) {
const title = await ctx.chrome.extract("h1.ytd-watch-metadata", "text");
await ctx.librarian.capture(url, { title });
}
Key design choices, locked-in for forward-compat:
- Header format borrowed from Tampermonkey/Greasemonkey. Familiar to anyone who’s written userscripts; no new format to learn.
- Capabilities are explicit. Every skill declares which librarian and bridge verbs it touches. Unknown capabilities = install-time rejection.
- Two execution targets. A skill is either DOM-side (runs in clipper) or librarian-side (runs in Rust via JS sandbox). The header decides.
- Install paths. Drag-and-drop a
.amemskill.jsonto the sidepanel, paste a URL pointing to one, or import from a community registry once one exists. - Same MCP surface.
amem_list_skillsandamem_invoke_skillkeep their v0.1 signatures. Custom skills appear in the same catalog, flagged withbadges: ["custom"].
This section is design intent, not commitment. Numbers can change. What is locked-in is the v0.1 MCP surface; v0.2 will not re-shape it.
Privacy
- The catalog UI itself sends no telemetry. Skill card render data comes exclusively from the loopback bridge.
- Auto-fire patterns (URL globs) live in
chrome.storage.localper-user, per-profile. They are not synced with Chrome Sync (we explicitly opt out by not declaringstorage.syncpermission). - Run-skill invocations log to the librarian only (
~/.amem/skills.log, one line per invoke with timestamp, skill id, outcome). Off by default; enable with[skills] log = trueinconfig.toml. - The “LinkedIn inbox glance” skill reads DOM content from the user’s own logged-in tab. It does not exfiltrate anywhere except the librarian’s local storage. The inbox digest is not auto-captured to the wiki — it renders inline only.
Failure modes
| Mode | Cause | Mitigation |
|---|---|---|
| Bridge disconnects mid-render | Librarian crashed or restarted | Cards render in kind:"error" state with “Reconnect” affordance; on reconnect, skills_state re-syncs and cards refresh |
| Auto-skill toggle drift | User toggles in CLI (amem skills enable …) while sidepanel is open | Sidepanel listens for skills_state push; re-renders |
| Run-skill hangs | Skill awaits a selector that never appears | Run buttons have a 60s soft timeout; user sees “Taking longer than usual…” + “Stop” |
| Pattern match fires on wrong URL | Glob too loose | v0.1 patterns are tightly scoped (e.g. arxiv.org/abs/* not *arxiv*); we err narrow |
| Sidepanel renders empty | amem_list_skills returned an empty array | Show “Skills service unavailable. Is amem-librarian running?” with install link |
| LinkedIn UI changes | LinkedIn redesigns inbox | Skill returns kind:"error" with selector-not-found; we patch the selector and ship a librarian update — no extension update needed (selectors are server-side) |
| Recording skill UI conflict | User clicks Run on cws-demo while another chrome_* call in flight | Recording acquires librarian-wide lock; concurrent invocations get RECORDING_IN_PROGRESS (RFC-001 §failure modes) |
Concrete work
In rough order:
- (
amem-clipper) Sidepanel tab strip +Captures/Skillsshell — ~0.5d - (
amem-clipper) Skill card component (auto + run + error variants) — ~1d - (
amem-clipper) Bridge subscribe /skills_statehandler + drift reconciliation — ~0.5d - (
amem-librarian) HardcodedSKILLStable +amem_list_skills/amem_invoke_skillMCP tools (lives in RFC-001 §3d but the wiring to the catalog UI is here) — ~0.5d - (
amem-clipper) Auto-fire content-script — URL pattern matcher,chrome.storage.localcache,auto_captureemitter — ~1d - (
amem-clipper) “Custom skills coming v0.2” footer + link — ~0.25d - (
docs.amem.sh)skills/v0.2.mdpage documenting the design intent header — ~0.5d - (
amem-clipper) CWS listing assets: screenshots of the catalog, listing copy per §6, demo gif (recorded by the cws-demo skill itself) — ~0.5d
Total: ~4.25d (overlaps with RFC-001 §concrete-work step 6, which budgets 1d for the same UI work — net new is ~3.25d.)
Rejected alternatives
- Ship without a “Skills” tab; make capture-only the v0.1 surface. Loses the CWS positioning. We are submitting at the same time as a hundred other capture extensions; we need the catalog story to stand out.
- Render skills as a flat list of MCP tools. Technically accurate, user-hostile. “MCP tool” is jargon; “skill” maps to mental models from Tampermonkey, App Store, Raycast extensions, etc.
- Make custom skills installable in v0.1 by accepting hardcoded patches. Doable but every patch is a librarian release; no actual install flow; misleading. Better to be honest and ship the disclosure.
- Hide the “v0.2 coming soon” disclosure. Considered for marketing cleanliness; rejected for trust. Every user opening the sidepanel will immediately wonder “can I write my own?”; pretending otherwise breeds cynicism.
- Use
manifest_version/ install registry now to keep “skills” honest. Equivalent to building a skill engine; rejected per RFC-001 §5.
Open questions
- Card density. Three skills look fine; do we need pagination / search at 10+? Soft preference: defer to v0.2 when skill count is user-driven.
- Skill icons. Use emoji per skill (📄, 💼, 🎬) or commission custom SVGs? Soft preference: emoji for v0.1 (zero design cost), SVGs if a CWS reviewer flags emoji as low-effort.
- Sidepanel width on narrow screens. The skill cards assume ~360px; Chrome sidepanel can be narrower. Verify on 1280-wide laptop.
- Should
amem_list_skillsfilter by current tab URL? I.e. only show “LinkedIn inbox” when the user is on linkedin.com? Considered; rejected for v0.1 — the catalog is supposed to feel like a marketplace shelf, not a context menu. - Run-skill audit. Do
kind:"run"invocations need a confirmation dialog (“Run LinkedIn inbox glance?”) or fire immediately? Soft preference: immediate for v0.1, confirmation if user feedback says surprising.
Roll-out
- v0.1 (this round): three hardcoded skills, sidepanel UI shipped, CWS submission. The catalog is real but not extensible.
- v0.1.x patches: site adapters and skill cards iterate based on actual usage. Each adapter or card is a librarian release; the extension changes only when the card schema changes (rare).
- v0.2: skill engine ships per the design-intent sketch (§8). Custom skills install via drag-and-drop. Catalog UI gains an “Install” affordance; existing cards keep working unchanged.
- v0.3+: community skills registry, skill versioning, signed skills. Out of scope here; track in a future RFC.
RFC-003 — Recording skill orchestration (v0.1)
- Status: Draft
- Authors: @yiidtw
- Created: 2026-05-09
- Related: RFC-001 (function-based v0.1), RFC-002 (Clipper skills catalog), SPEC.md § “amem is its own best demo,” archived guide/self-recording.md (superseded by this RFC)
TL;DR
Replace the existing chrome.tabCapture self-recording skeleton with a
scripted, librarian-driven recording pipeline. A YAML script declares a
sequence of Chrome operations; the librarian drives Chrome through the
bridge (which executes them in amem Clipper) while simultaneously running
macOS window-level screencapture -v -l<windowID> against the Chrome
window. Output: an mp4 at ~/.amem/recordings/<uuid>.mp4 plus an
amem://recording/<uuid> URI. Optional handoff to video-use for
post-processing (transcribe, cut filler, captions).
The whole thing surfaces as one MCP tool — amem_record_demo — and one
skill card (cws-demo, RFC-002 §3c). v0.1 ships one built-in script
(the CWS demo). Two more (feature-update demo, tutorial template) follow
when the third script is requested.
Why librarian-side and not extension-side: chrome.tabCapture cannot
record the sidepanel UI itself (sidepanels are out-of-tab surfaces),
which is exactly the surface our CWS demo needs to show. Window-level
macOS capture fixes that and only that.
Motivation
amem inherits crossmem’s principle — “a product that is its own best
demo” (SPEC.md § README, also docs/src/guide/self-recording.md). We
record our own marketing video by driving our own extension. This is real
load-bearing infrastructure, not a marketing gimmick: every CWS
re-submission, every feature announcement, every tutorial benefits.
The Day 1 skeleton (docs/src/guide/self-recording.md) used
chrome.tabCapture from an offscreen document. Three problems made us
abandon it for v0.1:
chrome.tabCapturecannot capture sidepanels. It captures the tab’s rendered area; the sidepanel is outside that surface. The single most important thing our demo must show — the skills catalog sidepanel — is invisible to tabCapture. We could render the catalog inside a tab as a workaround, but then we are demoing a fake.- No driving model. The skeleton assumed “an orchestrator” sends
start_recordingand “drives the extension UI.” There is no orchestrator design — just a hand-wave. v0.1 must ship a real one. - MV3 service worker lifecycle is hostile to long recordings. Service workers can be evicted under memory pressure; offscreen documents help but add coordination complexity. Doing this in the librarian (a long-running Rust process) is simpler and more reliable.
Moving the recorder to the librarian means it can use macOS-native
screencapture (window-level, captures everything in the Chrome window
including all of Chrome’s chrome) and orchestrate the demo via the bridge.
Same component split as RFC-001: librarian = brain, clipper = sensor.
Proposal
1. YAML script schema
Scripts live in two places:
crates/amem-librarian/builtin_scripts/*.yaml— built-in templates (CWS demo, etc.), compiled into the binary~/.amem/recordings/scripts/*.yaml— user scripts (v0.1 supports these but does not have a UI for managing them; CLI only)
Schema:
# ~/.amem/recordings/scripts/cws-demo.yaml
name: "CWS demo (v0.1)"
duration_target: 45s # advisory; total run time including waits
window:
app: "Google Chrome"
match: "active" # or { title_regex: "..." } for multi-window
record_cursor: true # passed to screencapture
output:
format: mp4
resolution: source # or 1080p, 720p (downscale via ffmpeg)
filename: "cws-demo-{ts}.mp4"
steps:
- id: open-arxiv
type: navigate
url: "https://arxiv.org/abs/1706.03762"
wait_for: "h1.title"
- id: capture-arxiv
type: caption
text: "amem captures any arxiv paper you open"
duration: 3s
- id: open-skills
type: click_selector
selector: "[data-amem-tab='skills']"
wait_after: 500ms
- id: show-skills
type: caption
text: "Three skills shipped — and your agent can call any of them"
duration: 4s
- id: click-linkedin
type: click_selector
selector: "[data-amem-skill='linkedin-inbox'] button.run"
wait_for: "[data-amem-skill='linkedin-inbox'] .digest"
- id: settle
type: wait
duration: 2s
- id: terminal-finish
type: terminal
cmd: "amem recall --json 'attention is all you need'"
show: "stdout"
duration: 4s
2. Step types
Each step is one variant; steps execute sequentially (no v0.1 parallelism).
| Type | Purpose | Driver |
|---|---|---|
navigate | Load URL in active tab | chrome_navigate over bridge |
click_selector | Click an element | chrome_click |
wait | Pause for N ms / N s | librarian sleep |
wait_for (also a field on other steps) | Block until selector / URL appears | chrome_wait |
extract_dom | Read DOM, save to script vars (for later assertions/captions) | chrome_extract |
caption | Render an overlay caption for N seconds | librarian draws into a borderless overlay window |
terminal | Run a binary command, optionally render its stdout in an overlay | librarian process + overlay |
The terminal step is interesting: it lets a recording show CLI usage
alongside the browser. The librarian opens a small floating window (via
its own UI process, not the Chrome window) that renders the command and
its stdout in monospace; screencapture picks up the overlay because we
target the Chrome window but composite the overlay on top of it before
each frame. Implementation detail: macOS lets us position a borderless
NSWindow above the target Chrome window and screencapture -l<chromeWinId>
will include it (Quartz compositing, same as visible-tab capture).
If the overlay approach proves unreliable across macOS versions, fallback
v0.1: split the recording — capture Chrome window for browser steps,
capture full screen for terminal steps, stitch in post via video-use.
Decision deferred until we hit a Sonoma/Sequoia regression in testing.
caption steps render a similar borderless overlay at the bottom-center
of the Chrome window with a translucent black bar.
3. macOS implementation — window-level screencapture
Core command:
screencapture -v -l <chromeWindowId> -V <duration_seconds> output.mov
-v— start recording immediately (no UI)-l <windowId>— target a specific window. We get this fromCGWindowListCopyWindowInfo(.optionOnScreenOnly)filtered to bundle identifiercom.google.Chrome(or Brave / Arc / Edge variants — we hardcode the major Chromium IDs).-V <seconds>— fixed duration. v0.1 sets this toduration_target + 10s buffer; we kill the process early if all steps finish before the timer.- Output
.mov; converted to.mp4viaffmpeg -i in.mov -c:v libx264 -crf 20 out.mp4post-recording.
Why window-level not full-screen:
- Privacy: full-screen captures the user’s desktop background, dock, notifications, other apps. Window-level captures only the Chrome window’s pixel rect.
- Aesthetics: window-level avoids us cropping the recording in post.
- Privacy posture matches the rest of amem (“data stays local; we don’t capture what we don’t need”).
Permission flow: macOS requires Screen Recording permission. On
first invoke, the librarian prompts the user via TCC. If denied, we
return { code: "SCREEN_RECORDING_DENIED", how_to_fix } so the agent
can surface the system-settings link.
Linux / Windows: out of scope for v0.1. The recording skill is
gated to macOS in v0.1 (#[cfg(target_os = "macos")]); on other OSes
the skill renders as kind:"error" with “Recording requires macOS in
v0.1.” Linux pipeline (probably wf-recorder for Wayland +
scrot/ffmpeg-x11grab for X11) is a follow-up RFC.
4. video-use integration (optional post-processor)
The raw mp4 from §3 is usable as-is — but for marketing-grade output we
want transcription, filler-word cutting, captions, and consistent
encoding. video-use is the operator’s existing tool for that; integrating
is a one-liner:
#![allow(unused)]
fn main() {
// crates/amem-librarian/src/skills/record_demo.rs
async fn record_demo(script: &Script) -> Result<RecordingOutput> {
let raw_mov = run_screencapture(&script.window, script.duration_target).await?;
let raw_mp4 = ffmpeg_convert(&raw_mov).await?;
let final_mp4 = if script.post_process.unwrap_or(false) {
video_use::process(&raw_mp4, &script.post_options).await?
} else {
raw_mp4
};
Ok(RecordingOutput {
amem_uri: format!("amem://recording/{}", uuid),
mp4_path: final_mp4,
})
}
}
post_process: false is the v0.1 default (raw recording). Setting
post_process: true in the YAML opts in. video-use is treated as an
optional dependency: if it isn’t installed, the field is ignored with a
warning.
The script’s post_options mirror video-use’s CLI flags one-to-one (so
the integration stays a thin shim, not a redesign).
5. amem_record_demo MCP tool
Exposed indirectly through amem_invoke_skill("cws-demo") per RFC-001
§3d, but also as a direct tool for ad-hoc agent use:
amem_record_demo(
script_path?: string, // path to YAML
script_inline?: string, // YAML literal (mutually exclusive with script_path)
output_dir?: string, // defaults to ~/.amem/recordings/
)
-> {
amem_uri: "amem://recording/<uuid>",
mp4_path: "/Users/.../<uuid>.mp4",
duration_s: number,
steps_run: number,
}
Errors:
| Code | Meaning |
|---|---|
SCREEN_RECORDING_DENIED | macOS TCC denied screen recording |
BRIDGE_UNAVAILABLE | Cannot reach amem Clipper to drive Chrome |
WINDOW_NOT_FOUND | No Chrome window matched script.window |
STEP_FAILED | Some step failed; partial recording saved with partial: true |
RECORDING_IN_PROGRESS | Another recording is active; refuse to start a second |
UNSUPPORTED_OS | Not macOS in v0.1 |
The tool is blocking: it returns when the recording finishes (typically 30–90s for v0.1 scripts). MCP clients already render “long-running tool call” affordances.
6. Use cases
The same pipeline serves three concrete needs, each justifying ship in v0.1:
6a. CWS promo video (the SPEC’s “own best demo” principle)
The CWS listing needs a 30–60 second demo. We script it (§1), the librarian records it, video-use polishes it, we upload. When v0.1.x patches change the UI, we re-run the same script. The demo never goes stale.
6b. Feature update demos (post-v0.1)
When v0.2 ships custom-skill installation, we want a demo of that flow. Same pipeline, new YAML script. v0.2 itself is a script.
6c. Tutorials
docs.amem.sh user guide pages can embed inline mp4s recorded from
canonical YAML scripts in crates/amem-librarian/builtin_scripts/. When
the UI changes, regenerate.
7. Why the librarian records, not the clipper
The cleanest framing of this RFC’s central decision:
| Option | Records sidepanel UI? | Privacy? | Lifecycle? | Cross-OS path? |
|---|---|---|---|---|
A. chrome.tabCapture (extension) | ❌ No (sidepanel is out-of-tab) | ✅ tab-only | ⚠️ MV3 service worker eviction risk | ✅ identical everywhere |
B. getDisplayMedia (extension) | ✅ User picks the window | ✅ user opt-in per recording | ⚠️ same MV3 risks | ✅ uniform |
| C. macOS window screencapture (librarian) | ✅ entire Chrome window | ✅ window-only, no desktop leak | ✅ Rust process is long-lived | ❌ macOS-only v0.1 |
Option B is the second-best choice; we considered it seriously. We rejected it for v0.1 because:
getDisplayMediarequires a user picker dialog every recording — incompatible with a “click Run, get a recording” agent-driven flow.- The MV3 service worker would still need to coordinate with the librarian for the script driver, doubling the moving parts.
- We want the librarian to own the timeline anyway — it already owns the script, the bridge, the storage, the post-process step. Adding the recording itself keeps responsibility in one place.
Option C costs us OS portability in v0.1 — accepted tradeoff. Linux / Windows recording lands when there is a non-macOS user.
Privacy
- Window-level screencapture: only the Chrome window’s pixels are captured. Desktop, dock, menu bar, other apps, notifications — all outside the frame.
- The recorded mp4 lives in
~/.amem/recordings/<uuid>.mp4. Never uploaded by the librarian. The user canamem upload(future) or manually drag-drop to a CWS listing. - macOS Screen Recording permission is requested on first invoke; the user can revoke at any time in System Settings.
- Caption / terminal overlays render the script’s literal text; no enrichment, no LLM rewriting (v0.1).
video-use(when used) is a local binary; transcription runs locally via whisper. No cloud calls in the default config.- The script itself is recorded into the mp4 metadata’s
commentfield (so anyone with the mp4 can see what was scripted). The user can disable withmetadata: falsein the YAML.
Failure modes
| Mode | Cause | Mitigation |
|---|---|---|
| Screen Recording permission denied | First-run user hasn’t granted TCC | Tool returns SCREEN_RECORDING_DENIED with open System Settings → Privacy → Screen Recording instruction; sidepanel card shows the same |
| Chrome window not found | User has no Chrome window open, or app is Brave/Arc | YAML window.app allow-list extended to common Chromium variants; tool errors with WINDOW_NOT_FOUND listing detected windows |
| Step times out | Selector never appears | Step times out at wait_for budget (default 30s); recording stops; partial: true flag in result so user knows |
| Bridge disconnects mid-recording | Librarian-clipper WS drops | Recording continues (it’s screencapture, not bridge-driven pixels), but subsequent steps can’t drive Chrome; we abort gracefully and save what we have |
screencapture produces 0-byte file | Known macOS bug in some Sonoma builds when target window is fully occluded | Pre-check: bring Chrome to front + verify visibility before starting; document the Apple bug in docs/troubleshooting.md |
| Multiple recordings requested concurrently | Two agents both call amem_record_demo | Librarian-wide recording lock; second call gets RECORDING_IN_PROGRESS |
| Output mp4 huge | High-resolution display, long demo | Default resolution: source for v0.1; users can opt to 1080p / 720p in YAML; ffmpeg downscale handled in §4 |
| Terminal overlay flickers / drops frames | macOS compositor under load | Documented; user can switch to “split capture + post-stitch” mode |
| Wrong Chrome window picked (multi-window) | User has 3 Chrome windows; we pick wrong | YAML window.match accepts title_regex; default is “active window” which uses the focused one at recording start |
Concrete work
In rough order:
- (
amem-librarian) YAML script parser + schema validator — ~0.5d - (
amem-librarian) macOS window-id resolver via QuartzCGWindowListCopyWindowInfo(FFI throughcore-graphicscrate) — ~0.5d - (
amem-librarian)screencapturedriver: spawn, monitor, kill-early on completion — ~0.5d - (
amem-librarian) Step executor: bridge-drivennavigate/click_selector/wait_for/extract_dom(mostly reuses RFC-001 bridge wrappers) — ~1d - (
amem-librarian) Caption / terminal overlay window (NSWindow with borderless Cocoa view) — ~1.5d (riskiest step; if compositing misbehaves, fall back to post-stitch mode and trim to ~0.5d) - (
amem-librarian)ffmpegmp4 conversion + optionalvideo-usehandoff — ~0.5d - (
amem-librarian)amem_record_demoMCP tool surface + structured errors — ~0.5d - (
amem-librarian) Built-incws-demo.yamlscript + smoke test on the actual amem Clipper UI — ~1d (this is the dogfood — script will reveal UI bugs) - (
docs.amem.sh) Replaceguide/self-recording.mdcontent with pointers to this RFC + how-to for users — ~0.25d
Total: ~6.25d (5d if overlay step falls back to post-stitch).
Rejected alternatives
- Keep
chrome.tabCapturefrom offscreen documents. Cannot capture sidepanel; demo would have to fake the catalog UI. Disqualifying. - Use
getDisplayMediafrom the extension. User picker dialog every time; service worker lifecycle headaches; doubles the moving parts. See §7 Option B. - Run a generic OBS / ffmpeg pipe and let the user start/stop manually. Loses the “agent-driven scripted demo” capability that makes self- recording load-bearing. We’d be back to manual demo production.
- Embed a video-use clone in the librarian. video-use is its own product with its own scope; reimplementing transcription / cut-filler in amem is out of scope. Optional handoff is the right boundary.
- Ship Linux / Windows recording in v0.1. Tripled scope for an audience we don’t currently have. macOS-only is acceptable v0.1 posture.
- Record from a JS-only stack (puppeteer + ffmpeg-screen). Loses the user’s Chrome profile and login state (per RFC-001 §3a / browser automation rules). Non-starter.
- Skip captions / terminal overlays. Dropping captions makes the resulting mp4 unsuitable for CWS listings without manual editing, which defeats the whole “own best demo” pipeline. Worth the implementation cost.
Open questions
- Overlay window rendering technology. Native NSWindow + Cocoa view
(proposed) or a small SwiftUI helper app the librarian shells out to?
Soft preference: Cocoa from Rust via the
cocoa/objc2crates; SwiftUI helper if FFI gets miserable. - Caption style. Plain translucent black bar with white text v0.1; do we offer themes / fonts? Soft preference: defer to v0.2; one hardcoded style in v0.1.
- Whether to embed the YAML script in the mp4 metadata by default.
Useful for reproducibility, mildly leaky for privacy if the user
shares the mp4. Soft preference: opt-in (
metadata: truein YAML). - video-use handoff API stability. v0.1 calls it as a binary subprocess; if video-use grows a Rust crate API later, switch over. Not blocking.
- Recording lock granularity. Per-librarian (current proposal) or per-script-id? Soft preference: per-librarian (simpler; scripted recording is sequential by nature).
- Should the sidepanel show a live preview of the recording? Tempting but adds complexity (need to read the in-progress mp4 or echo screencapture frames). Defer to v0.2.
Roll-out
- Day 4 of v0.1 build: YAML parser + window resolver + screencapture driver land. Smoke test: record a 5-second blank capture.
- Day 5: step executor + bridge wiring. Smoke test: 10-second scripted recording opens arxiv, clicks something, finishes.
- Day 6: caption / terminal overlay. Risk day; fallback path ready.
- Day 7 (CWS submission day): full
cws-demo.yamlruns, video-use polishes, we have the listing video. Submit. - Post-v0.1: track real script usage. When the third user-script exists (per RFC-001 §6 rule-of-three trigger), re-evaluate whether YAML schema needs versioning, whether scripts need a UI, whether Linux support belongs in v0.2.
- Future: cloud-render mode (a remote machine runs the script and returns the mp4) for users without macOS. Out of scope here; tracked in a future RFC.
RFC-004 — Reference self wiki
- Status: Draft
- Authors: @yiidtw
- Created: 2026-05-14
- Supersedes: archive/003-claim-grounding.md (broader fact-check pipeline; cut)
- Numbering note: originally drafted as RFC-003 (unarchive of the claim-grounding RFC) but renumbered to 004 because
003-recording-orchestration.mdlanded on main first.
TL;DR
When the model is talking with the user, it should pull cites from ~/.amem/wiki/
on its own. Today the amem_recall and amem_cite MCP tools exist but the
model only calls them when explicitly told. This RFC closes that gap with one
prompt resource and one convenience tool.
Scope is small on purpose: just self-reference. No web dial-out, no LLM verifier, no clipper overlay, no iOS keyboard. Those were in the archived 003 and got cut because none of them needed to ship before the basic loop works.
Motivation — where the gap actually is
Anthropic already does grounding for things they can see:
| What | Citation source |
|---|---|
| Web search tool | Web pages |
| Citations API | Documents you pass in context |
| claude.ai Projects + Files | Files you uploaded to their cloud |
What none of those touch: markdown files you wrote on your own disk.
Anthropic can’t see ~/.amem/wiki/. MCP is the seam they left for it — and
amem already exposes amem_recall + amem_cite over MCP.
The remaining gap is behavioural, not technical:
Today: user says "what did I read on transformers"
→ model speculates from training data
→ only calls amem_recall if user types "@amem" or asks explicitly
Wanted: model auto-calls amem_recall when the topic is something the user
might have captured, attaches amem:// cite when it hits, says
nothing extra when it misses.
Proposal
1. MCP system_prompt resource
amem mcp serve exposes one resource:
URI: amem://system/wiki-grounding
Mime: text/plain
Body:
When the user asks about a paper, dataset, technical spec, talk, or
any source-able fact they might have captured, call `amem_recall`
with the topic's key terms BEFORE answering from training data.
- If a hit is returned: phrase the answer in terms of the wiki entry
and append a cite "[<cite_key>](amem://<cite_key>)" so the user can
click through. If they want BibTeX/APA/MLA, call `amem_cite`.
- If no hit: answer normally from training knowledge. Do NOT invent
a cite_key. Optionally suggest: "I don't see this in your wiki —
want me to amem_capture <url>?"
Skip recall for: code questions, logistics, jokes, opinions, the
user's own preferences, very-well-known facts (e.g. "Python is dynamically typed").
MCP-aware clients (Claude Code, Cursor, Cline, Zed) merge resource content into the session system prompt automatically. No client patches.
2. New tool: amem_ground(query)
Single round-trip alternative to recall→cite chaining:
amem_ground(query: string, limit?: int) -> {
hits: [{
cite_key: "vaswani2017attention",
title: "Attention Is All You Need",
amem_uri: "amem://vaswani2017attention",
excerpt: "...",
bibtex: "@article{vaswani2017attention, ...}"
}],
inline_md: "[Vaswani et al. 2017](amem://vaswani2017attention)"
}
Useful when the model knows it’ll cite (e.g. user asked a “what did the paper say” question). Saves one MCP round-trip vs. recall+cite separately.
3. amem:// URI scheme
Stable, filesystem-independent reference:
amem://<cite_key> # whole wiki entry
amem://<cite_key>#chunk=<n> # specific chunk (post-v0.1)
Plus a CLI handler so links in chat are clickable from terminal:
amem open amem://vaswani2017attention
→ opens ~/.amem/wiki/1776567380_vaswani2017attention.md in $EDITOR
What’s intentionally NOT in this RFC
These were in archived 003 and get pushed out:
| Cut | Why |
|---|---|
| Trust list + arxiv/wiki dial-out fetch | amem capture <url> already exists; user-triggered capture is enough until v0.2 |
| LLM-verify step (claim ↔ evidence) | Over-engineering before the basic recall loop is proven |
amem-clipper typing observer / contradiction toast | Adds 3rd UI surface; ship reference-self alone first |
amem audit chat-history | After-the-fact, not in-conversation |
| iOS keyboard extension | Apple keyboard sandbox is brutal; far future |
default_action = block modes | Paternalistic; warn-only is fine for v0.1 |
If reference-self works and users want more, those come back in 004/005/… in their own RFCs, scoped tightly.
Verification — how we know it works
Two metrics from dogfooding (target: 1 week, 50+ conversations):
- Call rate: of conversations that mention a topic the user has in wiki,
what % auto-trigger
amem_recall? Target: ≥60%. Below 30% = prompt too weak; strengthen. Above 90% in conversations without relevant wiki content = noisy; soften. - Hit rate: of
amem_recallcalls, what % return ≥1 hit? Target: ≥40%. Lower means model is searching too vaguely; refine the prompt’s “key terms” guidance.
Both numbers come from logging in amem-librarian MCP server (already logs
recall calls; just need a daily summary).
Open questions
- Does MCP
system_promptresource actually flow into Claude Code’s prompt? Need to confirm against current Claude Code MCP behaviour. If resources don’t auto-merge, fallback is to bake the rule into each tool’sdescriptionfield (every tool tells the model when to call it). - Empty-result nudge: should
amem_recallreturn{hits: [], suggest_capture: "<url>"}when it detects a URL-shaped query? Soft yes — turns misses into capture opportunities. - Stale wiki entries: if you captured arxiv v1 in 2024 and the paper updated to v3, the cite is technically wrong. Out of scope for v0.1; flag for later.
Roll-out
| Day | Work | Owner |
|---|---|---|
| 1 | amem mcp serve exposes amem://system/wiki-grounding resource | amem-librarian |
| 1 | Add amem_ground tool wrapping recall+cite | amem-librarian |
| 2 | amem open amem://... CLI resolver | amem-librarian |
| 3 | Dogfood — count call/hit rates over 1 week | — |
| 7+ | If metrics look good: ship; if not: tune prompt and repeat | — |
No new infra. No client patches. Doesn’t block CWS submission of amem-clipper.
Rejected alternatives
- Auto-recall on every model turn — too noisy; would fire on “thanks” and “no”
- Background daemon scanning Claude transcripts — privacy + complexity for marginal gain
- Docs telling users to type
@amem— that’s the current state; doesn’t work because nobody remembers - Force model to cite something even on misses — turns into hallucinated cite_keys; worse than silence
- Ship as part of amem-clipper sidepanel UI — the gap is in Claude Code / Cursor / Cline, not the browser; clipper sidepanel doesn’t help
RFC-006 — Agent-driven file upload in logged-in Chrome
- Status: Draft
- Authors: @yiidtw
- Created: 2026-05-22
- Related: RFC-001 (function-based v0.1), RFC-002 (Clipper skills catalog), amem-librarian#3 (window-only capture)
TL;DR
Today no MCP-driven path lets a Claude/Cursor/Cline agent upload a file in the user’s real, logged-in Chrome. Three reasons:
- JS-initiated
<input type="file">click → Chrome blocks the native picker (or opens it but JS can’t see/select) - crossmem bridge has no
execute_script/evalaction (verified 2026-05-22), so we can’t even inject the DataTransfer workaround through it - Playwright / CDP routes are banned by
CLAUDE.md(debug-port issues)
This RFC adds file upload as a first-class amem capability by giving
amem-clipper its own bridge to amem-librarian (separate from crossmem) and
landing one new MCP tool: chrome_upload_file(selector, path).
Differentiation framing: when this ships, amem-clipper is the first non-debug-port stack that lets an agent complete a file upload on any logged-in site (CWS dashboard, GitHub PR attachments, Notion image upload, Slack file drops). Closing this hole is one of the largest practical agent gaps in 2026; TARS / Connector / chrome-devtools MCP all hit the same wall.
What “file upload” actually needs
Standard web upload flow:
1. user clicks <button> / <label for=fileinput>
2. browser opens native file picker
3. user selects file(s)
4. <input type=file>.files is set
5. page reads .files, kicks off XHR/fetch
Steps 1, 5 are normal DOM operations. Step 2–4 happen inside a sandbox JS can’t reach. The workaround the web has used for ~a decade:
// in page context
const dt = new DataTransfer();
dt.items.add(file); // file is a File object
input.files = dt.files;
input.dispatchEvent(new Event('change', { bubbles: true }));
This BYPASSES the picker entirely. The page’s own change-handler runs as if
the user picked the file. Works on >95% of normal <input type=file> sites.
Sites with custom drag-drop-only zones may need a drop event variant.
Why this can’t run through crossmem today
crossmem bridge actions (verified 2026-05-22 via curl /command):
✅ navigate, click, type, wait, extract, screenshot, summarize,
tab_info, ping
❌ execute_script, eval, inject_script, capture_page,
chrome_runtime_send, fetch_resource
Without an execute_script verb, the agent can’t push the DataTransfer
snippet into the page. To stay within CLAUDE.md rules (no Playwright, no
debug port), the only options are:
A. Fork crossmem to add execute_script — out of scope (third party)
B. Grow amem-clipper into its own bridge endpoint — chosen
C. Use Native Messaging — feasible but adds OS-specific plist/registry
plumbing; left as v0.2 hardening (see §Hardening)
Architecture (chosen path: B)
Add a second bridge daemon, embedded inside amem mcp serve, scoped to
amem-specific commands. crossmem stays in charge of general agent
computer-use; amem owns this new lane.
┌────────────────────────────────────────────┐
agent (claude)──►│ amem-librarian │
│ stdio MCP server │
│ ┌─────────────────────────────────────┐ │
│ │ clipper_bridge (NEW) │ │
│ │ tokio HTTP server on │ │
│ │ 127.0.0.1:7601 │ │
│ │ ├─ POST /command (agent ↔ daemon) │ │
│ │ └─ GET /poll (ext ↔ daemon) │ │
│ └────────────────┬────────────────────┘ │
└────────────────────┼───────────────────────┘
│ long-poll JSON
┌────────────────────▼───────────────────────┐
│ amem-clipper (Chrome MV3 extension) │
│ background.js — poll loop │
│ └─ on "upload_file": │
│ chrome.scripting.executeScript({ │
│ target:{tabId:active}, │
│ func: dataTransferInject, │
│ args:[selector, base64, mime, name] │
│ }) │
└────────────────────────────────────────────┘
Why a separate port from crossmem (7600 → 7601):
- amem can be installed without crossmem and still work
- crossmem can be uninstalled without breaking amem
- Daemons stay single-purpose: cross-extension fan-out vs amem-specific verbs
- Avoid editing third-party code we don’t own
Wire protocol
POST /command (agent → daemon):
{
"id": "<uuid>",
"action": "upload_file",
"params": { "selector": "input[type=file]",
"path": "/Users/me/file.mp4" }
}
Daemon reads the file, base64-encodes, enqueues:
{
"id": "<uuid>",
"action": "upload_file",
"params": {
"selector": "input[type=file]",
"fileName": "file.mp4",
"mimeType": "video/mp4",
"base64": "AAAAFGZ0eXBpc..."
}
}
GET /poll?since=<lastId> (extension → daemon, long-poll up to 30s):
returns next pending command or 204 on timeout.
POST /ack (extension → daemon):
{ "id":"<uuid>", "success":true, "error":null, "data":{...} }
Daemon returns the ack back to the original POST /command caller.
DataTransfer injection snippet (runs in page context)
(selector, base64, mime, name) => {
const bin = atob(base64);
const buf = new Uint8Array(bin.length);
for (let i = 0; i < bin.length; i++) buf[i] = bin.charCodeAt(i);
const file = new File([buf], name, { type: mime });
const input = document.querySelector(selector);
if (!input) return { ok: false, error: 'selector not found' };
const dt = new DataTransfer();
dt.items.add(file);
input.files = dt.files;
input.dispatchEvent(new Event('input', { bubbles: true }));
input.dispatchEvent(new Event('change', { bubbles: true }));
return { ok: true };
};
Failure modes & mitigations
| Mode | Cause | Mitigation |
|---|---|---|
Selector matches <button> not <input> | Common — many sites hide the real input | Resolver tries selector → if not input[type=file], walks up to <label for> / <form> / aria-controls to find the real input. Documented in tool description. |
| Site uses drag-drop only zone | No <input type=file> to set | v0.1 fails fast with “no compatible input”. v0.2 may dispatch a synthetic drop event with the file in DataTransfer. |
| Site validates with a custom event listener | Most use change; some only input | Snippet dispatches both. |
| File >100MB | base64-over-localhost is slow/memory hungry | Cap at 50MB in v0.1; bigger files return {error: "file too large; use v0.2 chunked path"}. |
| Multiple file inputs on page | Wrong one selected | User must provide a specific selector. Document :nth-of-type patterns. |
File path outside $HOME | Surprising | Resolve path; refuse if outside $HOME unless --allow-system-paths is set. |
| Extension not connected | amem-clipper not installed / not polling | Daemon returns 502 after 10s with “amem-clipper not connected — load the extension”. |
What’s NOT in this RFC
- Native Messaging variant — cleaner long-term, postponed to v0.2 once cross-platform installer (mac/win/linux) is built. See §Hardening.
- Drag-drop only sites — niche, defer to v0.2
- Chunked uploads — same; v0.1 caps at 50MB total
- Folder uploads —
webkitdirectoryinputs — defer - Cross-frame uploads — iframes that own the input — defer
MCP tool surface
One new tool registered with the existing amem mcp serve:
chrome_upload_file(selector: string, path: string) -> string
description:
Upload a local file to a logged-in Chrome page via amem-clipper.
The selector should target a standard <input type="file"> or a
parent <label for=...>/<button> that maps to one. Path must be
inside $HOME unless --allow-system-paths is set in amem config.
Returns:
success: <fileName> uploaded into <selector>on okError: <reason>on failure (selector miss, file too large, extension not connected, etc.)
Hardening (post-v0.1, separate RFCs)
- Native Messaging variant — replace the loopback HTTP poll with a NM port. Faster, no port choice/conflict, survives reboots. Requires per-platform manifest install.
- Drop-zone fallback — synthesize
dropevent with DataTransfer for sites that don’t expose an<input type=file>. - CWS-submission skill — built on top of
chrome_upload_file, automates the 4 file-pickers in the CWS dashboard flow (issue #amem-hq/11). - File chunking — for >50MB uploads, slice base64 across multiple poll messages reassembled in the extension before injection.
Roll-out
| Day | Work | Owner |
|---|---|---|
| 1 | This RFC merged | amem-hq |
| 1 | amem-librarian clipper_bridge module + MCP tool stub | amem-librarian |
| 2 | amem-clipper background poll + DataTransfer injection | amem-clipper |
| 2 | End-to-end smoke test against a public <input type=file> page | both |
| 3 | Document in docs/guide/file-upload.md + bake into CWS-submission skill | amem-hq |
Nothing in this RFC blocks the current CWS amem-clipper submission (that’s manual on the 4 file pickers); but shipping this RFC means the next extension submission could be one MCP call end-to-end.
Rejected alternatives
- Add
execute_scriptto crossmem — out of scope; crossmem is a third-party project we shouldn’t fork unilaterally - Use chrome-devtools MCP (CDP) — banned by CLAUDE.md, debug-port issues
- Browser-use / TARS visual route — they hit the same native-picker wall
- Server-side preview + manual user step — not agent automation, defeats the point
- Ask the user to drag-drop into a sidepanel — friction; sidepanel-only inputs don’t help when the upload form is on the target site
Open questions
- Should the daemon also handle download mirrors (
amem_download(url, path))? Likely yes — symmetric verb, same protocol shape. Out of scope for this RFC. - Long-term: does amem-clipper replace crossmem for our users, given the new bridge channel? Initial answer: no, parallel — crossmem keeps its generic bridge role.
- Multi-window Chrome: if user has two Chrome windows, which gets the upload? v0.1 picks the active tab of the focused window. Documented.
RFC-007 — Skill distillation: download → wiki → Claude skill
- Status: Draft
- Authors: @yiidtw
- Created: 2026-07-12
- Related: SPEC.md “Pipeline” section, RFC-004 (reference self wiki), amem-clipper#3 (Paste Board)
TL;DR
amem Clipper’s positioning is a three-step pipeline: download the resource the user is reading/watching (web page, video) → compile it into an Obsidian vault (personal wiki, knowledge graph) → distill recurring procedural knowledge into Claude skills.
Steps 1–2 ship today. This RFC defines step 3: when a cluster of wiki nodes qualifies for distillation, how the compile works, and the guardrails that keep the skill list from drowning in junk.
Corrected 2026-08-11 (competitor scan). The original claim here — “every clipper competitor stops at storage; nobody closes the loop into agent capability” — is now half wrong, and the surviving half is narrower:
- Closing the loop to agent access is a red ocean. LLM Wiki ships a
bundled MCP server plus a published
llm_wiki_skillinstallable withnpx skills add; SiYuan, Karakeep, and basic-memory all serve MCP too. Claiming nobody reaches agent capability is false. - Compiling captured knowledge into new skills is still unclaimed. Every
one of those, verified by reading their source, consumes hand-written
SKILL.mdfiles — none generates one from what the user captured.
So the moat is not “step 3”; it is the automatic generation half of step 3. Write it that way externally, because the wider claim does not survive contact with LLM Wiki.
Three-tier semantics
| Tier | Kind of knowledge | Agent usage |
|---|---|---|
| wiki node | declarative — “what is X” | storage substrate |
| MCP recall | reference — search + cite with provenance | on-demand lookup |
| skill | procedural — “how to do X” | auto-triggered reflex |
Distillation is selective. Of ~100 captures, maybe 3 clusters represent a
repeatable workflow worth compiling; the other 97 stay reference material
served via amem_recall. Compiling everything would pollute skill discovery
and destroy trigger accuracy.
Qualification criteria (what makes a cluster skill-worthy)
A node cluster qualifies for distillation when ALL of:
- Procedural — the nodes describe steps/commands/decision rules, not just facts. Heuristic: imperative verbs, numbered steps, code blocks with commands (not just definitions).
- Recurring — signal that this workflow repeats:
- ≥3 captures in the same topic cluster, OR
- the same nodes surfaced in ≥3 distinct
amem_recallqueries, OR - the user explicitly says so.
- Self-contained — the distilled skill can execute from its own text + linked wiki nodes, without the original page being live.
v1 trigger is explicit only: amem compile-skill <cluster> (MCP tool +
CLI). Auto-suggestion (“these 4 nodes look like a workflow — distill?”) is a
later phase; auto-compilation without a human in the loop is a non-goal.
Compile output
~/.claude/skills/<slug>/SKILL.md # or project .claude/skills/ when scoped
- Frontmatter
name+descriptionfollow skill-creator conventions (description states WHEN to trigger, with 中英 keywords the user actually says). - Body: distilled procedure, NOT a paste of the source nodes.
- Provenance footer is mandatory: wikilinks back to the source
~/.amem/wiki/<node_id>.mdnodes + original URLs. A skill whose sources died should be auditable and re-verifiable (amem_factchecktie-in).
Guardrails
- Skill budget — warn when distilled skills exceed ~20; force review of the least-triggered before adding more.
- No secrets — compile refuses content matching credential patterns;
vault references (
vault get KEY) instead of literals. - Eval before install — run the
eval/skill-creatorgrader on the generated SKILL.md; below-threshold output lands as a draft in the wiki, not in~/.claude/skills/. - Idempotent recompile — re-running on the same cluster updates the existing skill (matched by provenance), never duplicates.
Video capture (step 1 scope note)
“Download” for video = transcript + key frames, not the media file: storage, copyright, and yt-dlp maintenance all argue against full downloads, and the agent consumes the text layer anyway. Full-media archival is a non-goal.
Acceptance
-
amem compile-skill <cluster>MCP tool + CLI verb - Qualification check (procedural + recurring + self-contained) with human-readable rejection reasons
- SKILL.md output with provenance footer, eval-gated install
- Idempotent recompile on provenance match
- E2E: 3 captures on one workflow → compile → new skill triggers in a fresh Claude Code session
Non-goals
- Auto-compiling every capture into a skill (junk-pollution failure mode)
- Full video/media archival
- Distilling from sources the user hasn’t captured (that’s the frontier model’s own knowledge, not amem’s)
RFC-001 — Bridge-first architecture + settings/feature-flags
- Status: Draft
- Authors: @yiidtw
- Created: 2026-04-21
- Supersedes: SPEC.md § Roadmap Day 3 item “standalone mode”
- Related: RFC-002 (RSS), amem-sh
youtubepipeline
TL;DR
Reposition the Chrome extension as amem Clipper — a sensor for the
agent in the browser. The native amem binary is the brain; sensors don’t
function without it. Drop standalone mode from the roadmap: the binary is a
hard prerequisite. When the bridge is unreachable or a feature flag is off,
the relevant UI is grayed out in place with a single-click enable/install
affordance, not hidden. Heavy capture features (YouTube, RSS, Drive) are
feature-flagged in ~/.amem/config.toml and lazy-loaded on demand.
Positioning and naming
amem is the brand for agent memory. The CLI (amem-sh) is the brain —
sync, compile, recall, fact-check. The Chrome extension is a sensor that
feeds the brain when the user is in the browser. They are not peers; the
sensor depends on the brain entirely.
Name: amem Clipper. Inherits the Evernote Web Clipper lineage (users know
what “clipper” means), avoids collision with the internal “bridge” process
name (amem-bridge server on WS 7600), and keeps the brand “amem” attached
to the core (the brain) rather than to a single sensor.
Product model:
| Layer | Name | Role |
|---|---|---|
| Brain | amem-sh (CLI + library + MCP) | The product. Storage, compile, recall, fact-check, agent API. |
| Sensor | amem Clipper (Chrome extension) | Browser sensor — surfaces capture moments to the brain. Cannot function without the brain. |
| Sensor | amem Pockist (iOS native, shipped 2026-04-23) | Mobile sensor — share-sheet, OCR, place extraction. Same dependency on the brain. |
| Infra | amem-bridge (WS 7600 process) | Loopback IPC between sensors ↔ brain. |
This positioning reshapes every downstream decision:
- “Does the sensor need feature X?” → Only if feature X is a capture moment. Storage, recall, fact-check always live in the brain.
- “What happens with no binary?” → The sensor is dark. A sensor without a brain is a window with no eye behind it. Frontload install; don’t half-ship.
- “Where do settings live?” → In
~/.amem/config.toml, owned by the brain. Sensors render them via a bridge RPC; they never hold authoritative state.
Motivation
The original SPEC envisioned two extension modes:
| Mode | Day | Requirement |
|---|---|---|
| Bridge | 2 | Native amem binary + WS 7600 |
| Standalone | 3 | No native binary; cloud sync via Drive |
After shipping the YouTube pipeline we hit three things that make standalone look worse than we expected:
- Standalone is structurally crippled. Compile requires Ollama, transcription requires whisper-rs — neither runs in a Chrome MV3 extension. A standalone capture can store URLs but cannot compile, so the wiki never builds. Agent-side MCP is also dark because MCP is native-only.
- Feature cost is real. YouTube alone needs yt-dlp (~20 MB), ffmpeg (~60 MB), whisper model (75 MB–3 GB depending on size). Bundling all of this into default install breaks “offline-first with zero cloud dependencies by default” by shifting the pain from network to disk.
- Dual code paths are a maintenance tax. Standalone + bridge would mean two storage backends, two capture pipelines, two sets of bugs. The SPEC’s principle 4 (“complement aide, don’t duplicate”) applies to our own internals too.
Meanwhile, bridge mode is already the richer experience. A 30-second curl amem.sh/install | sh is less friction than a crippled standalone fork.
Proposal
1. Bridge is the only mode — UI degrades by graying out
Clipper on cold-start pings ws://127.0.0.1:7600/status. Based on the response, individual UI regions render as enabled, gray-disabled with a one-click enable, or gray-disabled with install CTA:
| State | UI for capture-web | UI for YouTube | UI for RSS | Global banner |
|---|---|---|---|---|
| Bridge unreachable | gray, tooltip “amem not running” | gray | gray | “Install amem → curl amem.sh/install | sh” with copy button |
| Bridge OK, all features off | enabled | gray, inline “Enable YouTube (~95 MB)” button | gray, inline “Enable RSS” button | — |
| Bridge OK, YT enabled | enabled | enabled | gray + enable | — |
| Bridge OK, all on | enabled | enabled | enabled | — |
Rationale: grayed controls are discoverable (user sees the feature exists, understands why it’s off) and honest (no hidden states). All-or-nothing install cards punish curious first-run users; per-feature gray-out is the “sensor goes dark when disconnected from the brain” pattern the positioning promises.
2. Feature flags live in ~/.amem/config.toml
Clipper renders a Settings page that maps to keys in this file via a new bridge RPC (settings_get / settings_set). The file is the single source of truth — both CLI and Clipper read/write the same keys. Clipper holds no authoritative config of its own, consistent with the peripheral positioning.
# ~/.amem/config.toml
version = 1
[features]
youtube = false # enables YT capture + compile (lazy-downloads yt-dlp + whisper model)
rss = false # enables RSS subscription ingestion (see RFC-002)
drive = false # enables Google Drive backup (Day 3)
[youtube]
whisper_model = "tiny.en" # tiny.en | base.en | small.en | medium.en
[bridge]
host = "127.0.0.1" # MUST be loopback (see Security)
port = 7600
token_file = "~/.amem/bridge.token"
3. Lazy-load on feature enable
Enabling a flag from Clipper or CLI triggers a setup routine:
amem youtube setup # CLI: downloads yt-dlp + whisper model + checks ffmpeg
amem rss setup # (RFC-002)
amem drive setup # (Day 3)
Clipper setup button → bridge RPC feature_setup({name}) → server runs the corresponding amem <name> setup, streams progress back over WS so Clipper can show a progress bar inline next to the (still grayed) control.
Graceful degradation in core flows. If a user runs amem capture <youtube-url> when features.youtube = false, the CLI prints:
YouTube capture is not enabled. To turn it on:
amem youtube setup
This will download yt-dlp (~20 MB) and the tiny.en whisper model (~75 MB).
The MCP tool amem_capture returns an analogous structured error, so agents can surface it to their user.
4. Bridge auto-start
On first install, amem install (the curl|sh script) registers a per-user background service:
- macOS:
launchctluser agent (~/Library/LaunchAgents/sh.amem.bridge.plist) - Linux:
systemctl --userunit (~/.config/systemd/user/amem-bridge.service) - Windows: Task Scheduler at-logon task (deferred; Windows support is a follow-up)
Goal: after first-run setup, the bridge is as available as Ollama is today — it just runs.
Security
Bridge-always means a persistent localhost WebSocket. Three defences, all MUST land before the “always on” posture ships:
| Defence | Mechanism |
|---|---|
| Loopback binding | Server binds 127.0.0.1, never 0.0.0.0. Reject --bind CLI flags that widen this. |
| Origin allow-list | WS handshake rejects connections whose Origin: header is not chrome-extension://<amem-clipper-prod-id> (production ID) or chrome-extension://<amem-clipper-dev-id> (dev build). |
| Token auth | On bridge start we mint a 32-byte random token to ~/.amem/bridge.token (mode 0600). Extension retrieves it via native-messaging handshake at install time. All WS messages must carry {"token": "..."} in their envelope. Tokens rotate on bridge restart. |
Threat model
| Threat | Impact | Mitigation |
|---|---|---|
| Another local program connects to WS | Could trigger amem_capture → write files to ~/.amem/raw/ | Token auth + origin check kill 99% of this |
| Disk-fill DoS via repeated capture | Fill user’s disk | Rate-limit captures per minute; refuse when ~/.amem/ exceeds configurable quota |
| Malicious browser extension connects as us | Impersonates our extension_id | Chrome refuses to forge Origin: for a different extension_id |
| RCE via yt-dlp / ffmpeg CVE | Arbitrary code execution | Use pinned versions, track security advisories; same posture as Ollama |
| Prompt injection in captured content | Poisons MCP amem_recall output | Same risk as today’s arxiv/PDF pipeline; not new from bridge-always |
Net risk: slightly higher than CLI-only (persistent WS endpoint exists), lower than an HTTP server accepting remote connections. Comparable to VS Code’s language-server loopback.
Migration
- SPEC.md §Roadmap: strike Day 3 “standalone mode”; add “bridge security hardening + install polish” and “Drive backup” (Drive stays).
- amem-clipper (repo renamed from
amem-extension2026-04-21): delete any standalone-only code paths (none should exist yet — Day 2 skeleton is bridge-only; this is a no-op today). Rename the product surface to “amem Clipper” in README, store listing, manifestname, and UI chrome. - docs:
guide/extension.mdrenamed toguide/clipper.md; its “Standalone vs Bridge” section is being rewritten as “How install works”. - README (amem-hq): clarify install-first story on all public pages, label the repo as “amem Clipper (Chrome MV3 extension)”.
Rejected alternatives
- “Pure cloud standalone” — extension + Drive only, no native binary. Breaks offline-first and agent-MCP. Also introduces OAuth complexity earlier than Day 3.
- Bundle everything in the default install. Ships ~3 GB of whisper models most users never use. Opposite of the lazy-load principle.
- Run whisper.wasm in the browser. Early 2026 performance is still 3–10× slower than native for
base.en; model download in the extension also hits MV3 storage limits.
Concrete work
See GitHub issues linked from this RFC.
amem youtube setupsubcommand + graceful-degradation prompt incapture(amem-sh)- Bridge: loopback binding +
Origin:check + token auth (amem-sh) - Bridge: auto-start service installers (macOS launchd, Linux systemd) (
amem-sh) - Clipper: bridge status probe, per-region gray-out UI, Settings page backed by
config.tomlvia bridge RPC (amem-clipper) - Clipper: rename product surface to “amem Clipper” (manifest, store listing copy, in-UI strings) — repo slug already renamed to
amem-clipper2026-04-21 (amem-clipper) - SPEC.md + docs: drop standalone, document bridge-first (
amem-hq)
Open questions
- Do we treat Ollama as a similarly lazy-loaded feature? Arguably yes — PDF compile also blocks without it. Worth a follow-up RFC if so.
- Drive backup (Day 3): should it require Pro/auth once shipped, or stay free? Product decision, not in scope here.
RFC-002 — RSS / Atom subscription management
- Status: Draft
- Authors: @yiidtw
- Created: 2026-04-21
- Related: RFC-001 (feature flags,
amem rss setup)
TL;DR
Add a first-class subscription layer on top of the existing capture pipeline. Users amem sub add <feed> to follow a source (arxiv category, blog, YouTube channel); a polling loop fetches the feed, dedups by GUID, and routes each new item through the existing capture + (optional) compile flow. MCP exposes amem_subscribe / amem_sub_list so agents can manage the user’s reading queue.
Motivation
amem’s capture flow is reactive: it only runs when a human (or agent) hands it a URL. That makes it useless for tracking ongoing sources:
- Following arxiv
cs.CLas new papers drop - Karpathy / Willison / lesswrong blog posts
- A YouTube channel’s new uploads (YouTube publishes per-channel RSS natively)
Every knowledge worker we’ve talked to does some version of this manually today — Feedly/NetNewsWire for reading, then copy-paste URLs into whatever capture tool they use. amem can collapse both steps.
This also inverts the self-recording story: amem is designed for you to produce content into. RSS lets other people’s content flow in on the same rails, so the wiki grows continuously rather than only after active capture.
Proposal
1. Subscription storage
# ~/.amem/subscriptions.toml
version = 1
[[subscription]]
id = "arxiv-cs-cl"
url = "http://export.arxiv.org/rss/cs.CL"
title = "arXiv cs.CL (Computation and Language)"
auto_compile = false # capture-only by default; compile is opt-in
poll_minutes = 60
enabled = true
added_at = "2026-04-21T00:00:00Z"
last_polled = "2026-04-21T00:30:00Z"
[[subscription]]
id = "3b1b-channel"
url = "https://www.youtube.com/feeds/videos.xml?channel_id=UCYO_jab_esuFRV4b17AJtAw"
title = "3Blue1Brown"
auto_compile = true # small channel, OK to auto-transcribe
poll_minutes = 240
enabled = true
2. Dedup ledger
~/.amem/subscriptions/
ledger.jsonl # append-only, one JSON object per seen item
state/{sub_id}/last_etag # HTTP caching
Each ledger line:
{"sub_id":"3b1b-channel","guid":"yt:video:aircAruvnKk","captured_at":"2026-04-21T00:30:00Z","cite_key":"3blue1brown2017neural"}
Dedup is GUID-based. If a feed republishes an item (edit, repost), the existing capture wins; we don’t re-download.
3. CLI
amem sub add <url> [--auto-compile] [--poll-minutes N] [--title "..."]
amem sub list [--json]
amem sub remove <id>
amem sub enable|disable <id>
amem sub fetch [<id>] # one-shot poll, honours etag
amem sub daemon # long-running poller (used by service unit)
amem rss setup # install daemon (macOS launchd / Linux systemd user)
4. Poll algorithm
For each enabled sub whose now - last_polled >= poll_minutes:
- GET feed with
If-None-Match: {last_etag}andIf-Modified-Since: {last_polled} - 304 → update
last_polled, skip - 200 → parse via
feed-rs, iterate items - For each item not in ledger:
- Route to existing
cite::cmd_capture(item.link)(auto-picks arxiv / PDF / YouTube based on URL) - If
auto_compile = true→ also call the appropriatecmd_compile - Write ledger line
- Route to existing
- Update ledger + state
Failures per-item don’t block the rest of the feed. Aggregate failures re-queue with exponential backoff (15 min → 2 h cap).
5. MCP surface
amem_sub_add(url, auto_compile?) -> sub_id
amem_sub_list() -> [{ id, title, last_polled, enabled, ... }]
amem_sub_remove(id)
amem_sub_fetch(id?) -> {fetched: N, captured: M, errors: [...] }
This lets an agent maintain its own research feed without a human in the loop: “follow every arxiv paper that cites Vaswani 2017” becomes a single MCP call.
6. Gated behind features.rss
Disabled by default. amem rss setup enables it, installs the daemon, and writes features.rss = true to ~/.amem/config.toml (per RFC-001).
Non-goals
- Rich reader UI. amem is not Feedly. Reading lives in the wiki +
amem recall. If people want visual unread counts, that belongs in an extension page, not the core. - OPML import on day 1. Easy add later; skip for MVP to keep surface small.
- Arbitrary scheduling cron.
poll_minutesis enough; cron-syntax scheduling is out of scope. - Podcast audio-only feeds. These would need whisper anyway — treat them as RFC-002b when YouTube pipeline is stable on more models.
Risks
| Risk | Mitigation |
|---|---|
| A popular arxiv category fills disk (dozens of papers/day) | poll_minutes default 120 + per-feed disk quota + user confirmation on first-time auto_compile = true |
| Feed publisher rate-limits us | Honour Retry-After, respect 429; back off to 6 h for repeat offenders |
| Duplicate captures when arxiv updates a paper’s version | Keep first ingest; subsequent versions append a note to the existing wiki entry rather than creating a new cite_key |
| RSS spec is loose — malformed feeds break parser | feed-rs handles common variants; log + skip malformed entries, do not abort the poll |
Concrete work
- Rust crate additions:
feed-rs = "2",toml_edit = "0.22"(config writes preserve comments) (amem-sh) amem subsubcommand family (amem-sh)amem rss setupinstaller andamem sub daemonlong-runner (amem-sh)- MCP tools (
amem-sh) - SPEC.md: add
subscriptionto the storage layout section (amem-hq) - Docs: new
guide/subscriptions.mdpage (amem-hq)
Extension UI for managing subscriptions is deferred — CLI first.
RFC-003 — Claim-grounding & fact-check pipeline
- Status: Draft
- Authors: @yiidtw
- Created: 2026-04-28
- Related: RFC-001 (bridge-first), RFC-002 (RSS), RFC-004 (planned: arxiv/wiki capture extension)
TL;DR
Fact-check is the core of amem, not a feature added on top. This RFC specifies the loop that makes amem distinct from every other “personal RAG” or “agent memory” product: every factual claim a user (or an agent speaking for them) makes is grounded against (1) the user’s own captured wiki and (2) a user-curated trust list of authoritative sources, and the result is written back so next time it’s a local hit.
Two automatic behaviours:
- Cite from your own wiki —
amem_recallis invoked on the claim’s key terms; if there’s a hit, the claim is annotated withamem://<id>and the evidence chunk. - Dial out to fetch + verify — when there’s no local hit,
amem_factcheckfetches from the user’s trust list (arxiv, wikipedia, specific blogs they trust), runs a verification step (LLM compares claim ↔ retrieved evidence), and captures the result back into the wiki.
Net effect: the things you assert in conversation become traceable to either your own captured knowledge or to sources you have explicitly chosen to trust. Whatever can’t be traced is surfaced as opinion.
Motivation — what amem actually is
amem is the brand for agent memory. amem-sh (the CLI) is the brain; amem Clipper and amem Pockist are sensors that feed the brain through the browser and the phone. Frontier models access the brain through MCP.
A “memory” that just stores and recalls is half a product. The other half — and the one that nobody else is doing well — is grounding: when the agent speaks for the user, every factual claim must trace to either captured content or a trusted source. Without that, the memory is just a fancier ChatGPT context window.
Two unsolved problems amem-grounding addresses:
- Agents speaking for the user invent plausible facts unless rigorously prompted. Even good models hallucinate at ~5–30% on long-tail factual claims.
- Humans speaking for themselves repeat half-remembered facts without knowing they’ve contradicted something they read last month.
amem already owns the storage layer (sensors capture; brain compiles). This RFC adds the outbound retrieval loop that closes the system.
Existing tools don’t fill this gap:
| Tool | Why it doesn’t ground claims |
|---|---|
| Glasp / Readwise | Capture + highlight only; no retrieval at speak-time |
| Notion AI / Mem.ai | Cites within their own product; doesn’t help when typing in Slack |
| Perplexity / Claude Web Search | Web-wide ranking, not the user’s trust list, no persistence |
| ChatGPT memory / Claude memory | Black-box server-side, no source surfaced, no user-curated audit |
| Letta / Mem0 | Agent memory with no grounding obligation; cites are not enforced |
amem is uniquely positioned because it already owns the user’s local corpus, already runs a sensor stack feeding it, and already exposes everything via MCP.
Proposal
1. New MCP tool: amem_factcheck
amem_factcheck(claim: string, context?: string) -> {
verdict: SUPPORTED | CONTRADICTED | NO_EVIDENCE,
evidence: [{ source_uri, snippet, captured_at, similarity }],
amem_uri: "amem://<id>"
}
Pipeline:
claim
├─ amem_recall(claim) ──────────────────► local hit?
│ │
│ ┌──────┴──────┐
│ ▼ ▼
│ HIT MISS
│ │ │
│ │ ▼
│ │ walk trust list
│ │ (priority order)
│ │ └─► fetch via adapter
│ │ └─► amem_capture
│ │ (becomes local
│ │ hit next time)
│ │
│ └─► LLM verify:
│ "Does <evidence> support
│ <claim>?"
│ → SUPPORTED / CONTRADICTED
│ / NO_EVIDENCE
│
└────────────────────────────────────────► return verdict + evidence + uri
amem_recall (already shipped Day 1) is the local search; amem_factcheck
adds the dial-out + verify steps.
2. amem:// URI scheme
Every cite uses one canonical format:
amem://<item_id> # whole captured item
amem://<item_id>#chunk=<n> # specific chunk
Why a custom scheme rather than https://...:
- Stable across
~/.amem/filesystem reorganisations - Works offline (the URI resolves to local content first, with
https://fallback for external clients) - Identifies content uniquely whether the user clipped it or it was pulled via factcheck
A small CLI handler resolves amem://... to the corresponding local file or
chunk for human-readable preview (amem open amem://abc#chunk=3).
3. Trust list in ~/.amem/config.toml
[grounding]
enabled = true
default_action = "warn" # warn | block | log
[[grounding.sources]]
name = "arxiv"
url_pattern = "arxiv.org/abs/*"
priority = 100 # higher = checked first
[[grounding.sources]]
name = "wikipedia-en"
url_pattern = "en.wikipedia.org/wiki/*"
priority = 80
[[grounding.sources]]
name = "fed-h15"
url_pattern = "federalreserve.gov/datadownload/Output.aspx?rel=H15"
priority = 90
[[grounding.sources]]
name = "user-rss" # see RFC-002
priority = 50
Design stance: amem does not ship an opinionated default trust list.
Users curate their own. This is a deliberate epistemic position — the tool
helps you ground claims in sources you chose, not in sources the tool’s
author chose for you. (amem grounding init --opinionated can offer a starter
list for first-runs who don’t want to think about it.)
4. Surfaces (phased)
| Phase | Surface | Effort | Effect |
|---|---|---|---|
| A | Agent system prompt rule via MCP system_prompt resource | 1 day | Every MCP-aware client (Claude Code, Cursor, …) auto-cites when speaking for the user |
| B | amem_factcheck tool implementation (trust list + dial-out + verify) | 3 days | Real grounding capability available to any caller |
| C | amem clipper input overlay on claude.ai / chatgpt.com / slack.com / gmail.com | 1 week | Catches human-typed claims, not just agent-generated ones |
| D | Weekly amem audit chat-history + iOS keyboard extension | TBD | Post-hoc + cross-device coverage |
Phase A in detail (ship-this-week)
amem mcp serve registers a system_prompt MCP resource:
Resource URI: amem://system/grounding-rule
Content:
When generating any factual claim about the world (not opinions, not
questions, not logistics), you MUST first call `amem_recall(claim)`.
- If a result is returned, append `[src: amem://<id>]` to the sentence.
- If no result is returned, call `amem_factcheck(claim)`.
- If factcheck returns NO_EVIDENCE, prefix the sentence with
"Without source: ".
- If factcheck returns CONTRADICTED, stop and surface the contradiction
to the user before proceeding.
Logistics, opinions, jokes, and the user's own preferences do not need
citations.
MCP-aware clients (Claude Code, Cursor, Claude.ai mobile) read this resource and merge it into their system prompt automatically. Zero new infra; zero client work.
Phase B — amem_factcheck implementation
Per-source adapter strategy is intentionally minimal in v1:
#![allow(unused)]
fn main() {
// Hardcode the two common adapters. Generalise to a trait only after
// the third source ships (rule of three).
async fn factcheck(claim: &str) -> Verdict {
if let Some(hit) = amem_recall(claim).await? {
return verify_with_llm(claim, &hit);
}
for source in trust_list_sorted_by_priority() {
match source.name.as_str() {
"arxiv" => arxiv_factcheck(claim).await,
"wikipedia-en" => wiki_factcheck(claim).await,
_ => generic_http_factcheck(&source, claim).await,
}
}
Verdict::NoEvidence
}
}
Both adapters reuse code from RFC-004 (arxiv/wiki capture extension), which ships alongside.
Phase C — clipper overlay
The amem Clipper extension (already has content scripts on claude.ai /
chatgpt.com / gemini.google.com for autosave, per current manifest)
acquires a new responsibility: debounced typing observer.
UI sketch:
[user typing in claude.ai compose box]
└─ debounce 800ms
└─ extract latest sentence ending in . ! ?
└─ background.js → bridge POST /factcheck
└─ if CONTRADICTED → side-panel toast:
┌────────────────────────────────────┐
│ ⚠ "GPT-4 hits 50% on ARC-AGI" │
│ Your wiki [Chollet 2024]: 30% │
│ [insert correct] [I have a new src]│
│ [different metric] [skip] │
└────────────────────────────────────┘
The overlay is non-blocking; the user can ignore it. It’s a hint, not a gate.
Phase D — audit + iOS
After-the-fact: amem audit chat-history --since=last-week ingests Slack /
iMessage / Gmail exports (where the user has authorised), runs each message
through amem_factcheck, and produces a markdown report grouped by:
- ✅ supported (cited from your wiki or trust list)
- ⚠ contradicted (you said X, your wiki says Y)
- ❓ unverified (no evidence in your trust list)
iOS keyboard extension: same idea as Phase C, but in a system keyboard. Painful to build (Apple keyboard sandbox is nasty), so deferred until A/B/C have shown value.
Privacy
- Trust list ≠ “the web”:
amem_factcheckonly contacts hosts the user has explicitly listed. There is no fallback to a generic web search. - Outbound traffic is logged: every dial-out is recorded in
~/.amem/grounding.logwith timestamp, host, and claim hash. - Claims never leave the loopback bridge unless they’re being sent to a
trust-list source. The local LLM verify step uses the locally configured
Claude API key (or a local model if configured) — same posture as existing
amem captureenrichment. - No federation by default: the captured factcheck results stay in
~/.amem/. They are not shared with any community cache.
Failure modes
| Mode | Cause | Mitigation |
|---|---|---|
| Over-citation | Casual chat triggers factcheck on every “hi” | Phase A prompt rule explicitly excludes “logistics, opinions, jokes”. Phase C overlay only fires on full-sentence claims. |
| Latency | factcheck takes 1–3s; agent feels slow | (1) cache verified claims by content hash; (2) run async, let agent draft sentence first then annotate; (3) skip factcheck on claims < 5 words |
| Source bias | User’s trust list is itself wrong | This is on the user. amem does not adjudicate truth, it just enforces traceability. (Documented as a feature, not a bug.) |
| Contradiction noise | Old captured content disagrees with current consensus | Surface both with timestamps; let user mark old item as superseded_by |
| Evidence ≠ support | LLM verify says SUPPORTED but it’s a misread | Show the evidence snippet next to the claim so user can sanity-check the reasoning step |
| Privacy regression | dial-out reveals reading habits to source | Documented in trust list config; user can disable per-source or globally |
Concrete work
In rough order:
- (
amem-sh) MCPsystem_promptresource with the grounding rule (Phase A) - (
amem-sh)amem://URI scheme + resolver CLI - (
amem-sh)amem_factchecktool with hardcoded arxiv + wiki adapters (Phase B) - (
amem-sh)~/.amem/config.toml[grounding]section + parser - (
amem-clipper) input observer + side-panel toast (Phase C) - (
amem-sh)amem audit chat-history(Phase D, slack/imessage/gmail importers) - (
amem-pockist) iOS keyboard ext (Phase D, much later)
Rejected alternatives
- “Build a generic SourceAdapter trait + trust-graph DSL” — premature abstraction. Hardcode arxiv + wiki; revisit on the third source.
- “Use Perplexity / Claude Web Search as the dial-out” — abandons the
user-curated trust list, which is the whole point. The web-search products
are a different product category (entropy reduction over global web), not
what
amemis for (entropy reduction over your chosen sources). - “Push factcheck verdicts to a community cache” — privacy mess + trust mess + premature for a pre-CWS product. Federation is RFC-005-or-later.
- “Block sending when CONTRADICTED” — too paternalistic.
default_action = "warn". Block is opt-in for users who want stricter discipline.
Open questions
- Should Phase A’s prompt rule include language for how to phrase uncertainty (e.g., “according to my wiki, …”) vs. leaving that to the client? Soft preference: prescribe a template, since otherwise every client prints citations differently.
- Should
amem_factcheckblock on contradiction in the verdict, or always return both verdict + evidence and let the caller decide UX? Soft preference: always return both; UX gating is a client concern. - Trust list versioning: when the user changes their trust list, do previously stored factcheck results need re-evaluation? Open.
- Should agent-generated claims that get
NO_EVIDENCEbe auto-filed somewhere (~/.amem/unverified/) for the user to either capture-the-source or reject-as-opinion later? Probably yes, but UX needs design.
Roll-out
- Day 8: ship Phase A (1 day of work).
amem mcp serveexposesclaim-grounding-ruleresource. amem-using devs immediately notice their agent starts citing. - Day 9–11: ship Phase B.
amem_factchecklands. - Day 12+: Phase C in
amem-clipper, scheduled after CWS approval + sidepanel UI work. - Phase D: open-ended, after first 5 external testers report grounding is the feature they actually use.
RFC-005 — Keyboard-driven region capture (multi-source)
- Status: Draft
- Authors: @yiidtw
- Created: 2026-05-06
- Related: RFC-001 (bridge), RFC-003 (grounding — this RFC supplies the citable units), RFC-007 (planned: macOS Accessibility hint mode)
TL;DR
Make every visual region a user can point at — a figure on a webpage, a
table on a PDF, a frame in a design tool — addressable by a stable
amem:// URI. Reference once, reference forever, even after the source
artifact is renamed, mutated, or deleted by an agent.
Two sources ship together in v0:
- Chrome hint mode — keyboard hotkey on amem Clipper draws
vimium-style hints over
<figure>,<img>,<table>, headings, and selectable blocks; user types the hint; region is captured + assignedamem://<uuid>. - Sioyek annotation watcher — amem-sh tails Sioyek’s local SQLite db;
every keyboard rectangle-mark (
rmode) the user makes auto-becomes anamem://<uuid>capture without changing Sioyek’s UX.
Unified output: every captured region surfaces in the user’s session with a
short alias like @fig-A that paste-able into agent prompts. The agent
reads the region image + metadata via the existing amem_recall MCP tool.
Motivation — the fig 4 → fig 2 problem
A scenario the operator hits weekly:
User: "Claude, change fig 4's color to blue."
Claude: <does the modification, but renames it to fig 2 in the process>
User: "fig 4 is now where? was it deleted?"
Claude: <confused about identifier history>
User: Spends 5 minutes verbally re-anchoring "the figure showing X-axis Y"
…or gives up and re-screenshots.
The bug isn’t the agent. The bug is that fig 4 is a name binding, not
an identity reference. As soon as the agent mutates the underlying
artifact, the binding breaks but the user’s mental model still points at
“that thing.”
The solved-problem analogues live elsewhere in software:
- Git’s commit hash (content-addressable; renaming the branch doesn’t invalidate the commit)
- Notion’s per-block immutable IDs
- Roam’s
((block-ref)) - amem’s existing
amem://<uuid>(already content-addressable for whole-document captures)
This RFC extends amem:// from document-grade addressing to
region-grade addressing — and gives the user a keyboard-only path to
mint such IDs from anywhere they read.
The killer secondary effect: every amem:// minted this way is automatically
a citable unit for RFC-003 fact-check. Grounding gets concrete targets
(“amem://b3f29... shows X”) rather than vague text recalls.
Proposal
1. amem:// URI extended for regions
Existing scheme:
amem://<item-id> # whole captured document
amem://<item-id>#chunk=<n> # specific text chunk
Add region addressing:
amem://<region-id> # always a 16-char base32 of UUID
# mints a NEW item, not a sub-ref of an
# existing one — content-addressable, so a
# region you saved last week resolves to
# the same id even if the source is gone
A region is a first-class amem item with these required fields:
# example: ~/.amem/raw/regions/b3f29x7q4t8m6w2n.toml
uri = "amem://b3f29x7q4t8m6w2n"
short = "fig-A" # session-scoped alias, may collide across sessions
source_app = "amem-clipper" # or "sioyek", "macos-ax", ...
captured_at = 2026-05-06T11:23:08Z
# What was captured
image_path = "raw/regions/b3f29x7q4t8m6w2n.png"
ocr_text = "Figure 4: Self-attention scores across heads…"
# Where it came from (best-effort, may be partial for some sources)
source_url = "https://arxiv.org/pdf/1706.03762"
source_doc = "Attention Is All You Need (Vaswani 2017)"
source_page = 7
source_bbox = [120, 340, 480, 620] # x, y, w, h in source coordinates
source_dom = "main > article > figure:nth-of-type(4)" # if web
# Optional, for grounding
parent_item = "amem://efb1...." # if region was extracted from an existing
# whole-doc capture
The point is: the URI alone is enough to recover the image + context even after the source is offline / mutated / removed. amem caches the image bytes locally; the source metadata is provenance, not the source of truth.
2. New MCP tool: amem_capture_region
amem_capture_region(target?: TargetSpec, mode?: "hint" | "wait")
-> { uri: "amem://...", short: "fig-A", image_path, source_meta }
TargetSpec :=
| { app: "chrome", tab: "active" | <tabId> }
| { app: "sioyek", current_doc: true }
| (omitted) // amem decides based on which source is "live"
Behaviour by mode:
mode: "hint"(default for Chrome) — agent calls the tool; bridge sendsenter_hint_modeto amem Clipper; user types a hint; the call blocks until the user picks (or 60s timeout aborts).mode: "wait"(default for Sioyek) — agent calls the tool; amem waits for the next new entry in Sioyek’s db (or 60s timeout); returns the region that landed.
Both are blocking from MCP’s perspective — clients (Claude Code, Cursor, etc.) already render “tool running” UI for long-running calls. The user isn’t surprised; they expect the tool to wait for their selection.
3. Source 1 — Chrome hint mode (amem-clipper)
Keyboard flow:
User in Chrome: presses ⌥G (or invoked via amem_capture_region MCP call)
↓
Content script enumerates candidate elements:
• <figure>, <img>, <table>, <video>, <pre>, <code>
• <h1>–<h4>
• Block elements with computed area > 8000 px²
• PDF.js text-layer divs (when PDF rendered in Chrome)
↓
Each element gets a 2-letter hint label: aa, ab, ac, …
↓
Overlay rendered at element's top-left corner with high z-index
↓
User types: "ac"
↓
Content script captures the region:
• Screenshot via chrome.tabs.captureVisibleTab + crop to bbox
• Plus: outerHTML of the element (DOM range archive)
• Plus: element's computed CSS rect, page URL, page title
↓
Posts to bridge: { action: "region_captured", payload: <region-record> }
↓
amem-sh receives, mints amem://<uuid>, stores PNG + metadata, returns short alias
Hint allocation: depth-first DOM traversal, skipping invisible elements (zero size or display:none). Two-letter hints support up to 676 elements per page; in practice ~50–100 visible candidates is plenty.
Activation: extension’s existing capture button gets a long-press / shift
modifier for “hint mode,” or a new dedicated icon. The hotkey ⌥G is a
reserved keyboard shortcut declared in manifest.json.
4. Source 2 — Sioyek annotation watcher (amem-sh)
Sioyek already provides keyboard-driven rectangle selection via the r
command — and it stores the result in a stable SQLite database. This RFC
does not require changes to Sioyek; we just listen.
~/.config/sioyek/local.db ← Sioyek writes (highlights, bookmarks, marks)
│
▼ fsnotify
amem-sh sioyek-watcher loop:
every change event ↓
▼
SELECT * FROM highlights
WHERE creation_time > last_seen_creation_time
▼
for each new row:
• read (begin_x, begin_y, end_x, end_y, page, document_path, type, creation_time)
• render that bbox of <document_path>:<page> via mupdf/pdfium → PNG
• OCR the rendered region via Vision/tesseract → text
• mint amem://<uuid>, store under raw/regions/
• update last_seen_creation_time
end
▼
emit MCP notification (if any client subscribed) so the agent can refresh
Trade-offs:
- ✅ Zero Sioyek modification, zero plugin maintenance
- ✅ Keyboard-only (Sioyek’s
rmode already is) - ✅ Cross-platform (SQLite is portable; mupdf works on Mac/Linux)
- ⚠️ Couples to Sioyek’s db schema. We pin
~/.config/sioyek/local.dbat SQLite version we tested with; on Sioyek upgrade we re-validate. - ⚠️ User has to be using Sioyek for this source to fire — non-Sioyek PDF users go via Chrome PDF viewer (see RFC-005 §3 / “PDF in Chrome” path)
5. Short alias system (@fig-A)
The full URI amem://b3f29x7q4t8m6w2n is unwieldy in chat. amem maintains
a per-session alias table:
~/.amem/state/aliases.json
{
"session_id": "2026-05-06-am",
"started_at": "...",
"aliases": {
"fig-A": "amem://b3f29x7q4t8m6w2n",
"fig-B": "amem://4ek38u2t1y9q5w0v",
"tbl-A": "amem://nz0p7qm6r3s4d2f8"
}
}
Alias scheme:
- Prefix indicates type:
fig-(figure / image),tbl-(table),txt-(text block),pg-(whole page screenshot),dom-(DOM range) - Suffix is alphabetic, in capture order this session
- Resets at session start (user can
amem alias persistto keep them)
CLI surface:
amem aliases # list current session aliases
amem alias persist # promote session aliases to permanent
amem alias resolve @fig-A # → amem://b3f29x7q4t8m6w2n
amem alias forget @fig-A # remove from current session
In agent conversations, the amem_recall tool already resolves amem://
URIs. Aliases get resolved client-side by amem-sh before the MCP call:
the user pastes @fig-A, the agent’s amem_recall sees the full URI.
6. Normalized capture record (cross-source schema)
All sources produce records of the same shape, regardless of provenance:
#![allow(unused)]
fn main() {
struct RegionCapture {
uri: String, // "amem://<uuid>"
short_alias: Option<String>, // "fig-A"
captured_at: DateTime,
// Always present
image_path: PathBuf, // PNG, full-resolution
image_bytes: u64, // for budget tracking
// Optional but encouraged
ocr_text: Option<String>,
image_caption: Option<String>, // alt text, figure caption, etc.
// Source provenance (one of these blocks present)
source: SourceProvenance,
}
enum SourceProvenance {
Web {
url: String,
title: String,
dom_path: String,
bbox_css: [f32; 4],
outer_html: Option<String>, // archived for posterity
},
Sioyek {
document_path: String,
document_hash: String, // for tracking renames / moves
page: u32,
bbox_pdf: [f32; 4], // PDF coordinates
document_title: Option<String>,
},
MacOSAX { // RFC-007, placeholder
bundle_id: String,
window_title: String,
bbox_screen: [f32; 4],
},
}
}
Every consumer (amem_recall, the side panel, amem cite) handles a
RegionCapture regardless of source. New sources just add new
SourceProvenance variants.
Privacy
- All captures land only in
~/.amem/raw/regions/on the user’s machine. Never uploaded. - Source provenance metadata is verbose by design (DOM path, PDF coordinates) so the user can audit. Verbosity stays local.
- The Sioyek watcher reads
~/.config/sioyek/local.dbonly; it never writes to it. Sioyek’s own data integrity is unaffected. - amem Clipper’s hint mode uses the same
chrome.tabs.captureVisibleTabpermission already declared. No new permissions. - Region OCR runs locally (Vision on macOS, tesseract on Linux). No cloud OCR.
Failure modes
| Mode | Cause | Mitigation |
|---|---|---|
| Hint mode times out (60s) | User got distracted, never picked | MCP tool returns a structured timeout error; agent can retry or apologise |
| User picks two hints simultaneously (race) | Multiple keypresses queued | Last-arrived wins; emit a debug log |
| Sioyek db schema changes after upgrade | Apple/Sioyek release breaks SQL | Watcher pins schema version, gracefully disables itself with a warning if the schema doesn’t match; user gets amem doctor sioyek to update |
| Same source region captured twice | User picks the same hint or makes the same Sioyek annotation | Content-hash dedupe; second capture gets the same amem://uri as the first |
| Region image too large (e.g., a full-screen 4K screenshot) | High-resolution monitor + lazy hint | Cap raw image at 4MB; downscale rest; preserve original under raw/regions/full/ if the user wants it |
| OCR misses non-Latin scripts | Default tesseract lang | Both Vision (macOS, multi-language) and tesseract are configured for zh-Hant, zh-Hans, en, ja to match amem-pockist |
Concrete work
In rough order of dependency:
- (
amem-sh) Extendamem://URI parser to accept region IDs (compat with existing item IDs — same UUID space) - (
amem-sh) AddRegionCapturedata model + storage layout underraw/regions/ - (
amem-sh) Addamem_capture_regionMCP tool (blocking, mode parameterised) - (
amem-sh) Addamem aliasCLI subcommands + session alias state - (
amem-bridge) Addenter_hint_mode+region_capturedverbs to bridge protocol - (
amem-clipper) Add hint overlay content script (~3 days; see §3) - (
amem-clipper) Wire ⌥G hotkey + side-panel “hint mode” toggle - (
amem-sh) Add Sioyek watcher (~/.config/sioyek/local.dbpoller + PDF region renderer via mupdf) - (
amem-sh) OCR pipeline (Vision + tesseract fallback) for region images - (
docs.amem.sh) User guide page: “Capturing regions for AI references”
Estimated total: two weeks of focused work, parallelisable across amem-sh / amem-bridge / amem-clipper. Sioyek source ships in week 1 (simpler — db reads only); Chrome hint mode ships in week 2.
Rejected alternatives
- “Write a Sioyek Lua plugin to bind a hotkey to amem capture” —
duplicates Sioyek’s built-in
rmode, requires per-Sioyek-version maintenance, and breaks if Sioyek loses Lua support. The fs-watch approach decouples completely. - “Mouse-driven rectangle selector” — operator’s stated preference is hand-stays-on-keyboard. Mouse mode could be a v2 nice-to-have but isn’t the design center.
- “Just screenshot and let the agent OCR/describe” — loses stable identity. Two screenshots of the same figure get different IDs; agent can’t track “the same thing” across calls.
- “Push everything through Chrome only (drop Sioyek)” — operator uses Sioyek for daily PDF reading. Forcing them into Chrome is a UX regression for the work Sioyek is good at.
- “Build native macOS overlay with Accessibility API now” — would cover all native apps generically, but is multi-week macOS work and is unjustified before Chrome + Sioyek prove the model. Deferred to RFC-007.
Open questions
- Alias naming policy — should aliases be
fig-Astyle (semantic prefix + letter), or pure@a1,@a2, …? Soft preference: semantic prefix; users immediately know@fig-≠@tbl-. But it requires classification at capture time (heuristic on element tag / PDF caption detection). - Cross-session alias persistence — should aliases auto-persist if the user uses them in a chat that the agent also persists into amem (closing the loop)? Probably yes; soft preference: aliases referenced in any captured conversation get auto-persisted.
- Multi-monitor / multi-window — when the user has two Chrome windows
on different monitors, which one gets the hint overlay? Soft
preference: only the focused window. If the agent calls
amem_capture_region(target={app: chrome, tab: ...})with no tab specified, route to the active window’s active tab. - PDF.js inside Chrome — should the hint mode treat PDF text-layer divs as targets (currently yes per §3), or should it route to the Sioyek source instead? The user’s choice of viewer is the answer; if the PDF is in Chrome, hint mode handles it.
- Should
amem_factcheck(RFC-003) accept a region URI as the claim context? Soft preference: yes —amem_factcheck(claim, context_uri: "amem://<region>")makes grounding richer.
Roll-out
- Week 1 — RFC-005 implementation kickoff. Sioyek watcher first
(cleanest path, no extension changes). Ships behind feature flag
features.sioyek_capture = true. - Week 2 — amem Clipper hint mode lands. Side panel shows live
capture log. Ships behind
features.region_capture = trueuntil stable. - Week 3 — wire
amem_factcheckto accept region URIs (RFC-003 Phase B integration). Now grounding loop is end-to-end. - Post-CWS — promote both flags from beta to default. Update
docs.amem.sh/clipperand add a newdocs.amem.sh/sioyekpage. - Future — RFC-007 macOS Accessibility hint mode, when first user asks “I want this for Sketch / Figma desktop / Adobe.”
Privacy Policy
The policy for amem Clipper and the amem CLI lives at:
https://yiidtw.dev/projects/amem-clipper/policy
That is the URL registered with the Chrome Web Store, and it is the only copy.
Why it is not here
A privacy policy is a URL a store holds on file, and it has to outlive the
product’s branding, its domain, and whether a given repo is public this month.
This one was served from docs.amem.sh, built out of a private repo — and when
that repo went private on 2026-08-22 the page’s own header links, its contact
channel, and four links in its body all became 404s at once. The document that
tells a user how to reach you should not depend on the thing it is documenting.
Keeping a second copy here would recreate the failure this project keeps
running into: CWS_LISTING.md sat at v0.2 for months while the live store text
moved on, and was read as authoritative anyway. A pointer cannot go stale in
that way.