Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

OpenWiki / llm-wiki-agent interop research for amem-librarian

Research date: 2026-08-10. Read-only study. Both repos cloned to /private/tmp/claude-501/-Users-ydwu-claude-projects-amem-hq/09fb9b1e-7475-4775-aad3-0eeb124d78aa/scratchpad/.


0. Headline finding

OpenWiki does not have a format of its own. It emits Google’s Open Knowledge Format (OKF). The interop question is therefore not “should amem talk to OpenWiki” but “should amem emit OKF”. Those are different decisions with different risk profiles, and the second one is the one worth making.

Second finding: OpenWiki shipped a personal mode with Gmail / X / Slack / Notion / Hacker News / web-search connectors writing to ~/.openwiki/wiki. That is the same product shape as amem — sensors feeding a local agent wiki. OpenWiki is now a competitor, not just an interop target. Details in §6.


1. Repos found

OpenWikillm-wiki-agent
URLhttps://github.com/langchain-ai/openwikihttps://github.com/SamurAIGPT/llm-wiki-agent
AuthorBrace Sproul (Head of Applied AI, LangChain) + LangChainAnil Chandra Naidu Matcha (SamurAIGPT)
Versionnpm openwiki 0.3.1no release; pyproject says 0.1.0
LanguageTypeScript (Node ≥22), React/Ink CLIPython 3.10–3.13
LicenseMITMIT
First commit2026-06-222023-04-21 (repurposed later)
Last commit2026-08-072026-08-03
Commits total255102
Commits last 30d16411
Contributors60+; top 3 = Colin Francis 139, bracesproul 63, Brace Sproul 5110; top 3 = Anil 32(+9+3), watsonk1998 23, Tony Lin 17
Bus factor~2–3 (LangChain-backed, corporate)~2 (single maintainer + one heavy contributor)

The ai.miraheze.org background page’s claim of >10k stars is consistent with the repo’s activity, though star count is not verifiable from a clone.

Engine

OpenWiki is a DeepAgents documentation agent (deepagents + @langchain/*). Model providers are pluggable and numerous — OpenAI (default gpt-5.6-terra), Anthropic, Gemini/Vertex, Bedrock, Copilot, OpenRouter, Baseten, Fireworks, Nebius, NVIDIA, Ollama, LM Studio, any OpenAI-compatible gateway. Keys live in ~/.openwiki/.env. BYO key; no hosted LangChain service.

llm-wiki-agent has two execution paths: standalone Python scripts calling litellm (LLM_MODEL env, default claude-3-5-sonnet-latest), or — the headline mode — no API key at all, just Claude Code / Codex / Gemini CLI reading CLAUDE.md / AGENTS.md / GEMINI.md and driving the workflow with its own file tools.


2. Storage layout

amem (current, on disk)

~/.amem/
  raw/          -> symlink to /Volumes/YDExtSSD/amem/raw
  recordings/   -> symlink to external SSD
  wiki/         34 flat .md files, no subdirectories

Flat namespace, node id encoded in filename with a colon: url:d5ffad083d4a.md, arxiv:1706.03762.md, x:AnthropicAI.md. Namespace counts: url: 30, arxiv: 1, x: 1, plus 2 legacy files.

~/.amem/index.md does not exist, although SPEC.md §“Storage layout” lists it as “auto-maintained”. No log.md. No config.toml either.

amem has two coexisting node schemas in one directory, which matters for any mapping work:

Schema A — legacy PDF/arxiv compile pipeline (1776567380_vaswani2017attention.md):

cite_key: vaswani2017attention
title: "Attention Is All You Need"
authors: [ ... ]
year: 2017
arxiv_id: "1706.03762"
raw: "/Users/ydwu/.amem/raw/1776567380_vaswani2017attention.pdf"
pdf_sha256: "bdfaa68d…"
chunks: 15

Body: # Title → ## Citations (APA/MLA/Chicago/IEEE/BibTeX) → ## Chunks with per-chunk ### p1c1 (p.1) + `sha256:3b3e5b…`.

Schema B — clipper node (the format the task description gives, 32 of 34 files):

id: url:d5ffad083d4a
type: "page"
title: "ForMACE Lab"
url: https://…
host: "formace-lab.zulipchat.com"
first_seen_at: 2026-07-28T09:48:19.062Z
captured_at: [ …, … ]      # list, grows on re-capture
tldr: "…"
tags: [amem-clipper]

Body: # Title → > tldr → **Source**: <url> → ## Content (verbatim extracted text) → ## Links → ## Connections (empty in v0.2; the file literally carries <!-- amem-clipper v0.2 leaves this empty; v0.3 populates with [[<node_id>]] wikilinks. -->).

Only Schema A carries SHA-256. The clipper nodes — 32 of 34 — have no hash at all, so the “provenance / SHA-256 excerpts” differentiator currently exists only on the legacy arxiv/PDF path.

OpenWiki

Two modes, two roots (src/config/openwiki-home.ts, src/okf/index-sync.ts:66):

code mode:      <repo>/openwiki/          (virtual root /openwiki)
personal mode:  ~/.openwiki/wiki/         (virtual root /)

~/.openwiki/                    mode 0700, ACL-restricted on Windows
  wiki/                         the knowledge bundle
  connectors/<id>/
    config.json                 never contains raw secrets, only env var names
    state.json                  {version, lastRunAt, latestIds, runs[]}
    raw/                        raw dumps + manifests
    logs/
  conversation_history/
  skills/                       bundled SKILL.md dirs, synced on each run
  .env                          provider keys
  install-id                    telemetry install id

Nested directories are first-class: /sources/gmail.md, /topics/ai-research.md. Every directory gets its own generated index.md.

llm-wiki-agent

Repo-rooted, not $HOME-rooted (tools/_utils.py):

raw/          immutable source documents, never modified
wiki/
  index.md    catalog, updated every ingest
  log.md      append-only chronological record
  overview.md living synthesis across all sources
  sources/    one page per source document      (kebab-case.md)
  entities/   people, companies, projects       (TitleCase.md)
  concepts/   ideas, frameworks, methods        (TitleCase.md)
  syntheses/  saved query answers
graph/        graph.json, graph.html, .cache.json, .refresh_cache.json
tools/        ingest/query/heal/health/lint/build_graph/refresh/…

One file per role, not per source: a single ingest fans out into one source page plus N entity pages plus M concept pages.


3. File format

OKF — the actual standard

Spec: GoogleCloudPlatform/knowledge-catalog/okf/SPEC.md. The published spec is now v0.2. OpenWiki still emits and hardcodes v0.1 (src/okf/index-sync.ts renderIndex() writes okf_version: "0.1"; src/agent/prompts/personal.ts:97 instructs “MUST follow the Google Knowledge Catalog OKF v0.1 schema”).

OKF v0.2 frontmatter:

FieldStatusShape
typeREQUIREDshort string, uncontrolled vocabulary (“BigQuery Table”, “Playbook”, “Reference”)
titlerecommendedstring
descriptionrecommendedone-sentence summary, retrieval-optimized
resourcerecommendedcanonical URI of underlying asset
tagsrecommendedlist of strings
sourcesoptionallist of {resource(req), id, title, author, usage_count, last_modified}
usage_windowoptional{from, to}
generatedoptional{by: <actor>, at: <ISO8601>}
verifiedoptionallist of {by: <actor>, at: <ISO8601>}
statusoptionaldraft | stable | deprecated (default stable)
stale_afteroptionalYYYY-MM-DD
extensionsallowedany additional keys

Actor convention (§7): <producer>/<version> for tools, human:<id> for people, process:<id> for automation. Consumers classify trust by the human: prefix.

Structural rules:

  • Reserved files index.md and log.md are not concepts; everything else .md is a concept document.
  • okf_version may appear only in the bundle-root index.md — the sole place frontmatter is allowed in an index.
  • Cross-links are standard markdown links, not wikilinks: [label](/tables/customers.md) bundle-absolute, or [label](./other.md) relative.
  • Conformance (§11): parseable YAML + non-empty type + reserved-file rules. Consumers MUST NOT reject a bundle for unknown keys, unknown type values, broken links, or missing index.md.
  • Extensions (§10): “Consumers SHOULD preserve unknown keys when round-tripping and MUST NOT reject documents with unrecognized fields.”

v0.1 → v0.2 breaking changes (§13.1): timestamp → generated.at; body # Citations section → frontmatter sources.

OpenWiki’s validator (src/okf/frontmatter.ts) enforces exactly the v0.1 subset: type required; type|title|description|resource|timestamp must be non-empty strings when present; tags must be a list of non-empty strings. It knows two of its own extension fields, openwiki_generated and openwiki_translation_pending.

llm-wiki-agent

Schema lives in CLAUDE.md (also AGENTS.md, GEMINI.md) — the agent reads it as instructions. It is prose convention, not a validator; nothing in tools/ enforces frontmatter shape.

title: "Page Title"
type: source | entity | concept | synthesis   # closed vocabulary
tags: []
sources: []          # list of source slugs (NOT OKF's structured sources)
last_updated: YYYY-MM-DD

Source pages additionally use date: YYYY-MM-DD and source_file: raw/….

Body convention for a source page: ## Summary → ## Key Claims → ## Key Quotes → ## Connections → ## Contradictions. Domain templates exist for diary/journal and meeting notes.

Links are [[WikiLinks]] — tools/_utils.py:extract_wikilinks() is re.findall(r"\[\[([^\]]+)\]\]", content). Resolution is by filename stem, case-insensitive, ignoring directories (tools/ingest.py:validate_ingest, tools/lint.py:page_name_to_path). Obsidian-compatible by design, with a documented vault-symlink pattern in the README.

wiki/index.md is a hand-structured catalog with fixed sections (Overview / Sources / Entities / Concepts / Syntheses); new entries are string-spliced under the right heading by tools/ingest.py:update_index().

wiki/log.md: append-only, ## [YYYY-MM-DD] <operation> | <title>, deliberately grep-parseable (grep "^## \[" wiki/log.md | tail -10). Operations: ingest, query, health, lint, graph.


4. Update loop, dedupe, provenance

OpenWiki

  1. openwiki personal --init / --update, or openwiki ingest <connector|all>.
  2. Connector tools (deterministic, no model) fetch and write ~/.openwiki/connectors/<id>/raw/ plus a manifest, and update state.json with lastRunAt / latestIds / a runs[] summary. Incrementality is per-connector watermarks — latestIds, commit SHAs, Notion object ids + last-edited timestamps + content hashes.
  3. migrateWikiToOkf() runs before the agent, normalizing every concept page’s frontmatter so the agent sees a conformant wiki.
  4. The DeepAgents agent reads raw dumps and writes/edits wiki pages.
  5. synchronizeWikiIndexes() regenerates every directory’s index.md deterministically.
  6. validateWikiInternalLinks() stamps broken links inline.
  7. Mermaid fences validated; failures downgraded to text fences with a reason comment.
  8. Post-run snapshot; unchanged runs record no new metadata (“no-op runs are free”).

Dedupe is agent-mediated — the model decides whether to edit an existing concept page or create a new one. There is no content-hash dedupe at the wiki layer and no stable node-identity field.

Provenance: source URI in resource; connector raw dumps retained on disk; state.json run history. No hashes in frontmatter. No verbatim archive of the source in the wiki page — pages are agent-synthesized prose. OKF v0.2’s sources[] / generated / verified families are not emitted, because OpenWiki targets v0.1.

llm-wiki-agent

python tools/ingest.py <file> (or “ingest raw/x.md” to Claude Code):

  1. Auto-convert non-.md via markitdown (20+ formats incl. pdf/docx/pptx/ xlsx/epub/ipynb and wav/mp3 audio transcription).
  2. sha256(source_content, truncate=16) computed and printed.
  3. Build context = index.md + overview.md + 5 most-recently-modified source pages (that’s the contradiction-detection window).
  4. One LLM call returns a strict JSON envelope: title, slug, source_page, index_entry, overview_update, entity_pages[], concept_pages[], contradictions[], log_entry.
  5. Write pages, splice index, append log.
  6. Post-ingest validation: broken [[wikilinks]] + pages missing from index.md, both printed.

Dedupe/staleness: tools/refresh.py re-hashes each source_file and compares to graph/.refresh_cache.json; only changed sources are re-ingested (--force overrides). tools/build_graph.py caches by SHA-256 so only changed pages get re-processed by the semantic pass.

Provenance: source_file: points at the immutable raw/ document, date:, sources: [] slug list, plus the hash cache. Hashes are truncated to 16 chars and stored in a cache file, not in the page frontmatter — so a page alone is not verifiable.


5. Entry points and MCP posture

OpenWikillm-wiki-agentamem
Installnpm i -g openwikigit clone + pip install -r requirements.txtRust CLI + MCP server
Serves MCPNoNo (zero MCP references in the repo)Yes — amem_capture/compile/cite/recall/…
Consumes MCPYes — src/connectors/mcp-client.ts, full JSON-RPC over stdio + HTTP, openwiki_list_mcp_tools / openwiki_call_mcp_tool, backend: "mcp-http" | "mcp-stdio"Non/a
Agent integrationwrites AGENTS.md/CLAUDE.md blocks between <!-- OPENWIKI:START/END --> markersis driven by Claude Code/Codex/Gemini via CLAUDE.md + .claude/commands/*.md slash commandsMCP tools
API keyrequired (any of ~15 providers)optional — none needed in Claude Code moden/a
Skillsbundles skills/mermaid-diagrams, skills/write-connector; syncBundledSkills() installs into ~/.openwiki/skills.claude/commands/wiki-{ingest,query,lint,graph}.mdRFC-007 distillation (not yet shipped)

This asymmetry is the interop lever. OpenWiki is an MCP client with a generic mcp-stdio / mcp-http connector backend and an McpReadOnlyOperation config. amem is an MCP server. An OpenWiki user can therefore configure amem as a connector today, with zero code in either project — no format conversion involved. That is a distribution channel, not a schema problem.


6. Competitive read (unrequested but load-bearing)

OpenWiki personal mode overlaps amem’s core pitch: local markdown wiki, private (~/.openwiki, mode 0700), BYO key, sensors pulling from the user’s tools, agent-maintained, MIT. With LangChain’s distribution and 164 commits/month it will not stay behind.

Where amem still differs, on the evidence:

  • Real browser sensor. OpenWiki’s web reach is Tavily search and public APIs. It cannot see a logged-in page. amem Clipper runs in the user’s real Chrome session — Zulip DMs, X timeline, paywalled reading. url:d5ffad083d4a (a Zulip login-gated page) is exactly the capture OpenWiki structurally cannot make.
  • iOS share-sheet sensor. No equivalent.
  • Verbatim archive. amem keeps ## Content — the actual extracted text. OpenWiki keeps only agent-written synthesis in the wiki (raw dumps live separately under connectors/*/raw/ and are not the wiki page).
  • Hash-level provenance + citation formatting. OpenWiki has neither; llm-wiki-agent has truncated hashes in a side cache.
  • Fact-check. Neither project has any grounding/verification tool. amem_factcheck (RFC-003) has no counterpart in either codebase.
  • Skill distillation. OpenWiki ships skills to its own agent; neither project compiles user knowledge into reusable agent skills. RFC-007 is still unique.

Where amem is behind: index generation, link inference, staleness policy, audit log, schema self-healing, multi-format ingest, graph visualization. Section 9 lists what to take.


7. Field-by-field schema mapping

amem clipper node (Schema B) as the source of truth.

amem fieldOKF v0.2 targetllm-wiki-agent targetVerdict
id (url:d5ffad083d4a)none — keep as extension amem_id; optionally sources[].idfilename stemlossy. OKF has no document-identity field; identity is the path. Colons in filenames also need sanitizing.
type: "page"type (required)type: sourcelossless but degraded. amem’s type encodes media kind; OKF’s encodes concept kind; llm-wiki’s encodes role. "page" is legal OKF but carries almost no routing signal — remap to Web Page / Paper / Profile.
titletitletitlelossless
urlresourcesource_file (local path)lossless to OKF. lossy to llm-wiki, which expects a repo-relative raw path, not a URL.
hostnone — extension amem_hostnonemissing both; trivially re-derivable from resource.
first_seen_atnone — extension, or sources[].last_modified (date-only)date (date-only)lossy: both flatten ISO-8601 datetime to a date, and neither has “first seen” semantics.
captured_at[] (list)nonenonelossy — worst case. Re-capture history is a list; OKF’s nearest list-valued time field is verified[], which means something else. Encode as extension amem_captured_at.
tldrdescription## Summary body sectionlossless to OKF (rename). Structural move for llm-wiki.
tagstagstagslossless
## Content (verbatim)no standard sectionno standard sectionlossy. Both target formats expect synthesized prose. Survives only as an unrecognized body section — which conformant consumers must tolerate, but no consumer will understand.
## Links——lossy, low value
## Connections [[node_id]][label](/node_id.md) markdown links[[PageName]]lossless to llm-wiki (same syntax, resolved by stem). Requires rewrite for OKF — wikilinks are not OKF edges. Currently empty at v0.2, so this is a forward-looking cost, not a migration cost.

Legacy Schema A (arxiv/PDF nodes):

amem fieldOKF v0.2llm-wiki-agentVerdict
cite_keysources[].idsource sluglossless-ish
authors[]sources[].author (single actor)—lossy — OKF’s author is one actor per source entry, not a list.
year, arxiv_idextension fields—lossy
raw (local path)sources[].resource (bundle-relative or absolute)source_filelossless
pdf_sha256none — extensioncache onlymissing from both standards. amem-unique.
per-chunk sha256:nonenonemissing from both. amem-unique; the substrate amem_cite depends on.
## Citations (APA/BibTeX/…)v0.1 had # Citations; v0.2 retired it in favour of sources—churn risk: the one body section OKF standardized was removed in the very next version.
chunks: 15extension—lossy

Fields amem should adopt from OKF, not just map to:

OKF v0.2 fieldWhy amem wants it
verified: [{by, at}]The standard’s designated slot for verification events. amem_factcheck output belongs here: verified: [{by: "amem/factcheck@0.3", at: …}]. amem would be the first producer populating the trust family with real verification rather than self-attestation.
generated: {by, at}Distinguishes clipper-captured from agent-synthesized nodes; human: prefix convention gives free provenance classing.
status, stale_afteramem has captured_at[] but no staleness policy; stale_after is a re-capture trigger.
sources[]Structured multi-source provenance — a node captured from 3 URLs is currently inexpressible in amem’s single url field.

8. Interop options

Effort: S. Recommendation: reject.

Mechanically near-free, but it is not a read: OpenWiki mutates the directory it is pointed at.

  • migrateWikiToOkf() (src/okf/index-sync.ts) walks every .md and calls normalizeConceptContent(), which rewrites any page lacking a usable type — replacing its frontmatter with a minimal derived block and stamping openwiki_generated: true. amem’s Schema A nodes have no type field at all, so every legacy arxiv/PDF node would have its cite_key, authors, year, arxiv_id, raw, pdf_sha256, and chunks deleted on first run. Only openwiki_translation_pending is on the preservation list.
  • synchronizeWikiIndexes() writes an index.md into every directory.
  • validateWikiInternalLinks() rewrites files to stamp broken links.
  • The agent then edits page bodies as prose — destroying ## Content verbatim text, which invalidates every chunk sha256.

Also: personal mode’s wiki root is hardcoded to ~/.openwiki/wiki (openWikiLocalWikiDir), so “pointing” means symlinking that path at ~/.amem/wiki — full write exposure, no read-only mode.

llm-wiki-agent is gentler (its WIKI_DIR is repo-relative and its tools are opt-in), and its [[wikilink]] syntax already matches amem. But it expects wiki/{sources,entities,concepts,syntheses}/ subdirectories and a hand-structured index.md; a flat directory of url:*.md files yields “unindexed page” warnings for all 34 nodes and an empty graph.

Secondary blocker for both: colons in filenames. url:d5ffad083d4a.md is illegal on Windows, ambiguous in some markdown link resolvers, and gets percent-encoded to url%3Ad5ffad083d4a by OpenWiki’s index generator (encodeURIComponent).

Effort: M (~2–4 days for a conformant v0.2 bundle).

Target OKF v0.2 the specification, and note in the docs that OpenWiki currently reads v0.1. Write to a fresh directory (~/.amem/export/okf/), never in place.

Transform:

  1. tldr → description; url → resource; title, tags pass through.
  2. type: "page" → a real concept type derived from the id namespace (url:→Web Page, arxiv:→Paper, x:→Social Profile).
  3. Emit generated: {by: "amem-clipper/0.2", at: <latest captured_at>}.
  4. Emit sources: [{resource: <url>, id: <node_id>, last_modified: <first_seen_at date>}]; for Schema A add author and the raw path, and carry pdf_sha256 as an extension.
  5. Preserve everything unmappable under amem_* extension keys (amem_id, amem_host, amem_captured_at, amem_pdf_sha256, amem_chunks) — §10 guarantees consumers must not reject them.
  6. Sanitize filenames: url:abc.md → web-page/abc.md, i.e. use the namespace as a directory. Kills the colon and gives OKF the nested structure it expects.
  7. Rewrite ## Connections [[id]] → [title](/web-page/abc.md).
  8. Generate root index.md with okf_version: "0.2" and per-directory indexes — reuse the algorithm in §9.1.
  9. Emit log.md from captured_at history.

What breaks / what you accept: the export is a projection, not the node. Chunk-level sha256 blocks and ## Content have no OKF home and survive only as extension body — legal, but no OKF consumer will act on them. Re-importing an OKF bundle edited elsewhere is explicitly out of scope (see (c)).

Cost is bounded because it is pure output — nothing in amem’s write path changes, nothing in amem’s schema is held hostage to OKF’s version churn, and if OKF v0.3 breaks again you edit one exporter.

(c) Bidirectional sync

Effort: L. Recommendation: reject, and say so in the docs.

Five independent reasons, any one sufficient:

  1. No stable identity in OKF. Identity is the file path. amem’s id is the join key and has no standard home, so a page renamed by the other tool is unmatchable on the way back.
  2. The other writer is an LLM. OpenWiki’s agent rewrites body prose non-deterministically. Round-tripping mutates ## Content, which invalidates every chunk sha256 — the exact substrate amem_cite and amem_factcheck stand on. Sync would corrupt the differentiator.
  3. Lossy in both directions. §7 shows captured_at[], chunk hashes, and verbatim content have no OKF representation; a round trip cannot restore what the projection dropped.
  4. No merge model. Two writers, no vector clocks, no CRDT, no conflict surface. Last-writer-wins over a knowledge base is silent data loss.
  5. Version skew is already real. OpenWiki writes v0.1 while the spec is v0.2, so a bidirectional bridge must translate between two versions of a moving format on every hop.

(d) Ship amem as an OpenWiki MCP connector — the cheap win

Effort: S (documentation only).

OpenWiki already speaks MCP as a client (src/connectors/mcp-client.ts, backend: "mcp-stdio" | "mcp-http", readOnlyOperations). amem already serves MCP. So an ~/.openwiki/connectors/amem/config.json pointing at amem mcp serve with readOnlyOperations: [{type: "tool", name: "amem_recall"}] makes amem a source OpenWiki reads from — no format conversion, no exporter, no schema commitment.

Strategically this is the right shape: amem is the sensor-and-provenance layer; OpenWiki becomes one more consumer of amem_recall. It also inverts the competitive dynamic — instead of amem exporting into LangChain’s format, LangChain’s tool queries amem’s API.


9. Ideas worth stealing

9.1 Deterministic per-directory index generation — openwiki/src/okf/index-sync.ts:synchronizeWikiIndexes() / synchronizeDirectory() / renderIndex(). Zero LLM calls; reads title + description from each page’s frontmatter; sorts by href; writes only when the rendered content differs from what’s on disk, so scheduled runs don’t churn git. amem’s SPEC promises ~/.amem/index.md and it does not exist — this is a ~100-line port and amem already has the two fields it needs (title, tldr).

9.2 Normalize-before-agent, never reject — openwiki/src/okf/frontmatter.ts:normalizeConceptContent() + migrateWikiToOkf(). A non-conformant page is repaired deterministically (derive title from first H1, else filename) and stamped openwiki_generated: true; the agent later sees that flag and enriches it. PRESERVED_EXTENSION_FIELDS carries control markers across the rebuild. Directly applicable to amem’s own Schema A/Schema B split and to the v0.2→v0.3 ## Connections migration: repair silently, mark for enrichment, never fail a run.

9.3 Broken-link stamping instead of failing — openwiki/src/agent/wiki-link-validator.ts:validateWikiInternalLinks(). Broken links get an inline <!-- openwiki: broken internal link … --> comment; existing stamps are cleared at the start of each pass so they never accumulate, and a later run finds the comment and repairs the href. A self-healing loop that degrades instead of erroring.

9.4 Degrade-and-repair for generated artifacts — same pattern for Mermaid (src/mermaid/validate.ts, README §Diagrams): an invalid diagram becomes a plain text fence with a reason comment rather than a broken render; the next update finds it and fixes it. Quality recovers over successive runs instead of requiring a correct one-shot.

9.5 The health/lint cost boundary — llm-wiki-agent/tools/health.py vs tools/lint.py, boundary documented as a table in CLAUDE.md. health is deterministic, zero LLM calls, free, run every session (empty/stub files, index sync, log coverage). lint is semantic, costs tokens, run every 10–15 ingests (orphans, broken links, contradictions, gaps). “Run health first — linting an empty file wastes tokens.” This is the correct shape for an amem doctor, and the explicit run-order rule is the valuable half.

9.6 Two-pass graph with typed edges — llm-wiki-agent/tools/build_graph.py. Pass 1 parses [[wikilinks]] → EXTRACTED edges (deterministic). Pass 2 asks the model for implicit relationships → INFERRED (with confidence) or AMBIGUOUS. Louvain community detection clusters topics; SHA-256 cache means only changed pages are re-inferred; output is a self-contained graph.html. amem’s ## Connections is empty at v0.2 — this is a ready-made v0.3 design, and the EXTRACTED/INFERRED distinction keeps model guesses auditable rather than laundering them into the same link namespace.

9.7 Contradiction flagging at ingest time — llm-wiki-agent/tools/ingest.py, the contradictions[] field of the JSON envelope, checked against index.md + overview.md + the 5 most recent source pages. Cheap precursor to amem_factcheck: catch conflicts when writing, not at query time. Their own README frames it as the RAG differentiator (“Contradictions surface at query time (maybe)” vs “Flagged at ingest time”).

9.8 Append-only, grep-parseable log — llm-wiki-agent/wiki/log.md, ## [YYYY-MM-DD] <operation> | <title>, designed so grep "^## \[" wiki/log.md | tail -10 is the read API. log.md is also an OKF reserved filename, so adopting it is free conformance. amem has no audit trail today.

9.9 Refresh-by-hash staleness — llm-wiki-agent/tools/refresh.py re-hashes each source_file, compares against graph/.refresh_cache.json, re-ingests only what changed. amem already stores captured_at[] but has no policy that consumes it; combine with OKF stale_after for a re-capture trigger.

9.10 Multi-format ingest via markitdown — llm-wiki-agent/tools/ingest.py:convert_to_md() gets pdf/docx/pptx/xlsx/ html/epub/ipynb and wav/mp3 transcription from one dependency. amem’s Pockist share-sheet would inherit a large format surface for very little code.

9.11 Secrets by reference, never by value — openwiki/src/connectors/: connector config.json stores env var names; values live only in ~/.openwiki/.env; ~/.openwiki is mode 0700 with Windows ACL restriction (src/platform/windows-acl.ts). Matches RFC-007’s “no secrets, use vault references” guardrail and shows the config-file shape.

9.12 .openwikiignore as a read boundary — gitignore syntax; when active, it filters filesystem discovery and restricts shell execute. The README is careful about what it does not promise (“does not guarantee a topic is never mentioned, since the agent may still infer an ignored area from other allowed evidence”). Good model for an .amemignore on the clipper, and good copy for honest scoping language.


10. Risks

Schema churn — high. OKF v0.1 → v0.2 shipped two breaking changes (timestamp→generated.at, # Citations→sources) inside roughly a month, and one of them retired the only body section the format had standardized. OpenWiki has not caught up: it validates and emits v0.1 today. Anything amem builds against “OpenWiki’s format” is targeting a stale snapshot of a moving spec. Mitigating factors: type is the sole required field, and §11 forbids consumers from rejecting unknown keys — so a minimal export (type + title + description + resource + tags) is very unlikely to break, while the rich provenance families are where churn will bite.

Velocity churn — high for OpenWiki. 164 commits in 30 days, 255 total since 2026-06-22, with active refactors landing in the exact modules that matter here (#611 reorganized the CLI, #513 restructured the repo into domain directories, #371 added the link validator). Any code-level coupling will rot fast. Format-level coupling will not.

Bus factor. OpenWiki ~2–3 with LangChain behind it — low abandonment risk, high direction risk (it is a company’s product and will follow the company’s roadmap). llm-wiki-agent ~2, single-maintainer, 11 commits in the last month and most of them cosmetic star-history chores — treat as a design reference, not a dependency.

License — no obstacle. Both MIT. amem’s repos are private/proprietary; MIT permits use and derivation with attribution, so reading their code for ideas is fine and vendoring a file is fine with the notice retained. Implementing OKF creates no license relationship at all: a data format is not copyrightable, and OKF is published by Google as an open spec. The clean rule: implement the format, do not vendor the code.

Coupling cost. Option (b) is one output-only module with no upstream dependency — its failure mode is “the export is stale”, which is cheap. Options (a) and (c) put a competitor’s LLM agent inside amem’s write path; their failure mode is silent corruption of the provenance data that amem_factcheck is supposed to stand on. That asymmetry is the whole decision.

Strategic risk. Emitting OKF is also a small act of standard adoption in LangChain-and-Google’s direction. It is worth it — the format is genuinely better specified than anything amem would invent, and verified/sources give the fact-check story a standard vocabulary — but amem’s internal schema should stay amem’s, with OKF as an export target only.


11. Recommendation

Do (b) + (d): a one-way amem export --okf targeting OKF v0.2, plus a documented recipe for registering amem as an OpenWiki MCP connector.

Rationale:

  • The real standard is OKF, not OpenWiki. Build against the spec; treat OpenWiki as one consumer that happens to lag at v0.1.
  • One-way export is bounded, reversible, and touches nothing in amem’s write path. If OKF v0.3 breaks, one module changes.
  • OKF’s verified: [{by, at}] and sources[] give amem_factcheck a standard vocabulary to publish into. amem would be the first producer filling the trust family with actual verification rather than self-attestation — that is a positioning asset, not just plumbing.
  • The MCP connector path costs a documentation page and inverts the dependency: LangChain’s tool queries amem, rather than amem exporting into LangChain’s world.

Order of work: steal 9.1 (index generation) and 9.5 (health/lint split) first — both are pure wins independent of any interop decision, and 9.1 closes a gap between SPEC.md and reality. Then the exporter. Then the MCP connector doc.

What amem should NOT do:

  1. Do not let OpenWiki write to ~/.amem/wiki/. migrateWikiToOkf() strips frontmatter from any page without a type — that is every legacy arxiv/PDF node, including pdf_sha256 and chunks.
  2. Do not build bidirectional sync. §8(c).
  3. Do not adopt OKF as amem’s internal schema. It has no field for captured_at[], no field for chunk hashes, and no document identity — the three things amem’s provenance model is built on. Export to it; don’t live in it.
  4. Do not target OpenWiki’s v0.1 dialect. Emit v0.2 and let OpenWiki catch up; v0.1 consumers tolerate the extra keys by §11 conformance anyway.
  5. Do not vendor their code. Reimplement the patterns; the value is in the design, and both codebases are moving too fast to track.
  6. Do not chase connector parity (Gmail/Slack/Notion/X). That is LangChain’s strength and a treadmill. amem’s edge is the logged-in browser session and the iOS share sheet — sensors OpenWiki structurally cannot build — plus fact-check and skill distillation, which neither project has attempted.