RFC-003 — Claim-grounding & fact-check pipeline
- Status: Draft
- Authors: @yiidtw
- Created: 2026-04-28
- Related: RFC-001 (bridge-first), RFC-002 (RSS), RFC-004 (planned: arxiv/wiki capture extension)
TL;DR
Fact-check is the core of amem, not a feature added on top. This RFC specifies the loop that makes amem distinct from every other “personal RAG” or “agent memory” product: every factual claim a user (or an agent speaking for them) makes is grounded against (1) the user’s own captured wiki and (2) a user-curated trust list of authoritative sources, and the result is written back so next time it’s a local hit.
Two automatic behaviours:
- Cite from your own wiki —
amem_recallis invoked on the claim’s key terms; if there’s a hit, the claim is annotated withamem://<id>and the evidence chunk. - Dial out to fetch + verify — when there’s no local hit,
amem_factcheckfetches from the user’s trust list (arxiv, wikipedia, specific blogs they trust), runs a verification step (LLM compares claim ↔ retrieved evidence), and captures the result back into the wiki.
Net effect: the things you assert in conversation become traceable to either your own captured knowledge or to sources you have explicitly chosen to trust. Whatever can’t be traced is surfaced as opinion.
Motivation — what amem actually is
amem is the brand for agent memory. amem-sh (the CLI) is the brain; amem Clipper and amem Pockist are sensors that feed the brain through the browser and the phone. Frontier models access the brain through MCP.
A “memory” that just stores and recalls is half a product. The other half — and the one that nobody else is doing well — is grounding: when the agent speaks for the user, every factual claim must trace to either captured content or a trusted source. Without that, the memory is just a fancier ChatGPT context window.
Two unsolved problems amem-grounding addresses:
- Agents speaking for the user invent plausible facts unless rigorously prompted. Even good models hallucinate at ~5–30% on long-tail factual claims.
- Humans speaking for themselves repeat half-remembered facts without knowing they’ve contradicted something they read last month.
amem already owns the storage layer (sensors capture; brain compiles). This RFC adds the outbound retrieval loop that closes the system.
Existing tools don’t fill this gap:
| Tool | Why it doesn’t ground claims |
|---|---|
| Glasp / Readwise | Capture + highlight only; no retrieval at speak-time |
| Notion AI / Mem.ai | Cites within their own product; doesn’t help when typing in Slack |
| Perplexity / Claude Web Search | Web-wide ranking, not the user’s trust list, no persistence |
| ChatGPT memory / Claude memory | Black-box server-side, no source surfaced, no user-curated audit |
| Letta / Mem0 | Agent memory with no grounding obligation; cites are not enforced |
amem is uniquely positioned because it already owns the user’s local corpus, already runs a sensor stack feeding it, and already exposes everything via MCP.
Proposal
1. New MCP tool: amem_factcheck
amem_factcheck(claim: string, context?: string) -> {
verdict: SUPPORTED | CONTRADICTED | NO_EVIDENCE,
evidence: [{ source_uri, snippet, captured_at, similarity }],
amem_uri: "amem://<id>"
}
Pipeline:
claim
├─ amem_recall(claim) ──────────────────► local hit?
│ │
│ ┌──────┴──────┐
│ ▼ ▼
│ HIT MISS
│ │ │
│ │ ▼
│ │ walk trust list
│ │ (priority order)
│ │ └─► fetch via adapter
│ │ └─► amem_capture
│ │ (becomes local
│ │ hit next time)
│ │
│ └─► LLM verify:
│ "Does <evidence> support
│ <claim>?"
│ → SUPPORTED / CONTRADICTED
│ / NO_EVIDENCE
│
└────────────────────────────────────────► return verdict + evidence + uri
amem_recall (already shipped Day 1) is the local search; amem_factcheck
adds the dial-out + verify steps.
2. amem:// URI scheme
Every cite uses one canonical format:
amem://<item_id> # whole captured item
amem://<item_id>#chunk=<n> # specific chunk
Why a custom scheme rather than https://...:
- Stable across
~/.amem/filesystem reorganisations - Works offline (the URI resolves to local content first, with
https://fallback for external clients) - Identifies content uniquely whether the user clipped it or it was pulled via factcheck
A small CLI handler resolves amem://... to the corresponding local file or
chunk for human-readable preview (amem open amem://abc#chunk=3).
3. Trust list in ~/.amem/config.toml
[grounding]
enabled = true
default_action = "warn" # warn | block | log
[[grounding.sources]]
name = "arxiv"
url_pattern = "arxiv.org/abs/*"
priority = 100 # higher = checked first
[[grounding.sources]]
name = "wikipedia-en"
url_pattern = "en.wikipedia.org/wiki/*"
priority = 80
[[grounding.sources]]
name = "fed-h15"
url_pattern = "federalreserve.gov/datadownload/Output.aspx?rel=H15"
priority = 90
[[grounding.sources]]
name = "user-rss" # see RFC-002
priority = 50
Design stance: amem does not ship an opinionated default trust list.
Users curate their own. This is a deliberate epistemic position — the tool
helps you ground claims in sources you chose, not in sources the tool’s
author chose for you. (amem grounding init --opinionated can offer a starter
list for first-runs who don’t want to think about it.)
4. Surfaces (phased)
| Phase | Surface | Effort | Effect |
|---|---|---|---|
| A | Agent system prompt rule via MCP system_prompt resource | 1 day | Every MCP-aware client (Claude Code, Cursor, …) auto-cites when speaking for the user |
| B | amem_factcheck tool implementation (trust list + dial-out + verify) | 3 days | Real grounding capability available to any caller |
| C | amem clipper input overlay on claude.ai / chatgpt.com / slack.com / gmail.com | 1 week | Catches human-typed claims, not just agent-generated ones |
| D | Weekly amem audit chat-history + iOS keyboard extension | TBD | Post-hoc + cross-device coverage |
Phase A in detail (ship-this-week)
amem mcp serve registers a system_prompt MCP resource:
Resource URI: amem://system/grounding-rule
Content:
When generating any factual claim about the world (not opinions, not
questions, not logistics), you MUST first call `amem_recall(claim)`.
- If a result is returned, append `[src: amem://<id>]` to the sentence.
- If no result is returned, call `amem_factcheck(claim)`.
- If factcheck returns NO_EVIDENCE, prefix the sentence with
"Without source: ".
- If factcheck returns CONTRADICTED, stop and surface the contradiction
to the user before proceeding.
Logistics, opinions, jokes, and the user's own preferences do not need
citations.
MCP-aware clients (Claude Code, Cursor, Claude.ai mobile) read this resource and merge it into their system prompt automatically. Zero new infra; zero client work.
Phase B — amem_factcheck implementation
Per-source adapter strategy is intentionally minimal in v1:
#![allow(unused)]
fn main() {
// Hardcode the two common adapters. Generalise to a trait only after
// the third source ships (rule of three).
async fn factcheck(claim: &str) -> Verdict {
if let Some(hit) = amem_recall(claim).await? {
return verify_with_llm(claim, &hit);
}
for source in trust_list_sorted_by_priority() {
match source.name.as_str() {
"arxiv" => arxiv_factcheck(claim).await,
"wikipedia-en" => wiki_factcheck(claim).await,
_ => generic_http_factcheck(&source, claim).await,
}
}
Verdict::NoEvidence
}
}
Both adapters reuse code from RFC-004 (arxiv/wiki capture extension), which ships alongside.
Phase C — clipper overlay
The amem Clipper extension (already has content scripts on claude.ai /
chatgpt.com / gemini.google.com for autosave, per current manifest)
acquires a new responsibility: debounced typing observer.
UI sketch:
[user typing in claude.ai compose box]
└─ debounce 800ms
└─ extract latest sentence ending in . ! ?
└─ background.js → bridge POST /factcheck
└─ if CONTRADICTED → side-panel toast:
┌────────────────────────────────────┐
│ ⚠ "GPT-4 hits 50% on ARC-AGI" │
│ Your wiki [Chollet 2024]: 30% │
│ [insert correct] [I have a new src]│
│ [different metric] [skip] │
└────────────────────────────────────┘
The overlay is non-blocking; the user can ignore it. It’s a hint, not a gate.
Phase D — audit + iOS
After-the-fact: amem audit chat-history --since=last-week ingests Slack /
iMessage / Gmail exports (where the user has authorised), runs each message
through amem_factcheck, and produces a markdown report grouped by:
- ✅ supported (cited from your wiki or trust list)
- ⚠ contradicted (you said X, your wiki says Y)
- ❓ unverified (no evidence in your trust list)
iOS keyboard extension: same idea as Phase C, but in a system keyboard. Painful to build (Apple keyboard sandbox is nasty), so deferred until A/B/C have shown value.
Privacy
- Trust list ≠ “the web”:
amem_factcheckonly contacts hosts the user has explicitly listed. There is no fallback to a generic web search. - Outbound traffic is logged: every dial-out is recorded in
~/.amem/grounding.logwith timestamp, host, and claim hash. - Claims never leave the loopback bridge unless they’re being sent to a
trust-list source. The local LLM verify step uses the locally configured
Claude API key (or a local model if configured) — same posture as existing
amem captureenrichment. - No federation by default: the captured factcheck results stay in
~/.amem/. They are not shared with any community cache.
Failure modes
| Mode | Cause | Mitigation |
|---|---|---|
| Over-citation | Casual chat triggers factcheck on every “hi” | Phase A prompt rule explicitly excludes “logistics, opinions, jokes”. Phase C overlay only fires on full-sentence claims. |
| Latency | factcheck takes 1–3s; agent feels slow | (1) cache verified claims by content hash; (2) run async, let agent draft sentence first then annotate; (3) skip factcheck on claims < 5 words |
| Source bias | User’s trust list is itself wrong | This is on the user. amem does not adjudicate truth, it just enforces traceability. (Documented as a feature, not a bug.) |
| Contradiction noise | Old captured content disagrees with current consensus | Surface both with timestamps; let user mark old item as superseded_by |
| Evidence ≠ support | LLM verify says SUPPORTED but it’s a misread | Show the evidence snippet next to the claim so user can sanity-check the reasoning step |
| Privacy regression | dial-out reveals reading habits to source | Documented in trust list config; user can disable per-source or globally |
Concrete work
In rough order:
- (
amem-sh) MCPsystem_promptresource with the grounding rule (Phase A) - (
amem-sh)amem://URI scheme + resolver CLI - (
amem-sh)amem_factchecktool with hardcoded arxiv + wiki adapters (Phase B) - (
amem-sh)~/.amem/config.toml[grounding]section + parser - (
amem-clipper) input observer + side-panel toast (Phase C) - (
amem-sh)amem audit chat-history(Phase D, slack/imessage/gmail importers) - (
amem-pockist) iOS keyboard ext (Phase D, much later)
Rejected alternatives
- “Build a generic SourceAdapter trait + trust-graph DSL” — premature abstraction. Hardcode arxiv + wiki; revisit on the third source.
- “Use Perplexity / Claude Web Search as the dial-out” — abandons the
user-curated trust list, which is the whole point. The web-search products
are a different product category (entropy reduction over global web), not
what
amemis for (entropy reduction over your chosen sources). - “Push factcheck verdicts to a community cache” — privacy mess + trust mess + premature for a pre-CWS product. Federation is RFC-005-or-later.
- “Block sending when CONTRADICTED” — too paternalistic.
default_action = "warn". Block is opt-in for users who want stricter discipline.
Open questions
- Should Phase A’s prompt rule include language for how to phrase uncertainty (e.g., “according to my wiki, …”) vs. leaving that to the client? Soft preference: prescribe a template, since otherwise every client prints citations differently.
- Should
amem_factcheckblock on contradiction in the verdict, or always return both verdict + evidence and let the caller decide UX? Soft preference: always return both; UX gating is a client concern. - Trust list versioning: when the user changes their trust list, do previously stored factcheck results need re-evaluation? Open.
- Should agent-generated claims that get
NO_EVIDENCEbe auto-filed somewhere (~/.amem/unverified/) for the user to either capture-the-source or reject-as-opinion later? Probably yes, but UX needs design.
Roll-out
- Day 8: ship Phase A (1 day of work).
amem mcp serveexposesclaim-grounding-ruleresource. amem-using devs immediately notice their agent starts citing. - Day 9–11: ship Phase B.
amem_factchecklands. - Day 12+: Phase C in
amem-clipper, scheduled after CWS approval + sidepanel UI work. - Phase D: open-ended, after first 5 external testers report grounding is the feature they actually use.