Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

amem Clipper + Librarian — 使用者痛點 vs 產品現況

研究日期:2026-08-17。唯讀研究,未修改任何 repo,未發任何 issue/PR。

方法:先親自讀 amem 原始碼與 ~/.amem/ 實際產出,再挖四類一手來源。 一手來源為 GitHub issue tracker(10 個 repo)、HN Algolia API、forum.obsidian.md Discourse API、Firefox AMO ratings API。前一份競品掃描(2026-08-11)的結論視為已知, 不重推。


0. 結論先講

amem 對「登入頁擷取」這個品類最大痛點答得很好,對「存下來的東西是否可信」答得最差。

第二點是致命的,因為它正是 amem 對外的定位句。SPEC 說 fact-check 是核心, 定位句說「每一句話都能追回原文」。實際出貨的 clipper 路徑做不到: 14 個節點在 6000 字截斷、沒有任何內容 hash、沒有 raw 備份、重剪直接覆蓋舊檔。

第三點是 RFC-007 的 moat 已經消失。不是變弱,是消失。Anthropic 官方文件明寫 產生 SKILL.md 不需要工具,Claude Code 二進位檔本身就會自動寫 SKILL.md。


1. amem 現況查核(我親自讀碼,非讀 SPEC)

這張表是後面所有判斷的地基。每一列都有檔名行號。

項目SPEC/定位的說法程式碼實際行為證據
內容 hash「每一句話都能追回原文」clipper 節點完全沒有 hash 欄位amem-librarian/src/clipper_bridge.rs:570-597 render_wiki_md 的 frontmatter 只有 id/type/title/url/host/first_seen_at/captured_at/tldr/tags
hash 涵蓋率—2/35 個 wiki 節點有 hash,都是 legacy PDF~/.amem/wiki/ 實測;30 個 url:* 節點與 1 個 clip:* 節點皆無
verbatim 全文「raw immutable → wiki」三層body 截斷在 6000 字,尾巴接 …(truncated)clipper_bridge.rs:588
截斷實際比例—14/30 個 clipper 節點已被截斷grep -l '…(truncated)' ~/.amem/wiki/*.md
raw 層~/.amem/raw/ 保留原件clipper 路徑不寫 raw。raw/ 只有 14 個 arxiv PDF 與 metabackground.js 只 POST /save_wiki_node,無 raw 寫入路徑
重剪語意去重、防 re-clip spam保留 first_seen_at 與 captured_at 清單,然後整檔覆蓋clipper_bridge.rs:359-397,std::fs::write 覆寫;captured_at 上限 20 筆
recall 檢索amem_recall / amem_groundtoken 計數 grep,非 BM25、非向量src/query/grep.rs:1 檔頭自述「MVP search… Token-count scoring」;src/mcp/mod.rs:167,175 走 helpers → grep
BM25 + 向量已建已建,但只服務 Zotero 參考文獻語料,不服務 wikisrc/refindex.rs(948 行)、src/embed.rs(519 行),入口是 refindex build / refcheck
Connectionsv0.3 填 wikilink31/35 個節點仍是字面 placeholder 註解<!-- amem-clipper v0.2 leaves this empty -->
index.md / log.mdSPEC 承諾不存在~/.amem/ 與 ~/.amem/wiki/ 均無
AI 對話自動存manifest 宣傳「Rich extraction on Claude / ChatGPT / Gemini」三個 autosave content script 都是 8 行 stub,只有 console.debugextension/content-scripts/{claude,chatgpt,gemini}-autosave.js 檔頭自述「Day 1-2: stub only」
AI 對話手動擷取同上真的可用。三個 extractor 各 200-230 行,含 artifact 與 binary 判別extension/extractors/claude.js 等;background.js:77-83 依 host 派送
離線容錯—無佇列、無重試。daemon 掛掉則該次擷取遺失background.js 唯一 alarm 是 amem-bridge-reconnect,只服務 bridge 輪詢
權限面—<all_urls> + 12 個 permission,含 tabCapture、downloads、scriptingextension/manifest.json
CWS 上架前次掃描判定「卡在 review」已公開上線。v0.3.0,更新 2026-08-13,3 users,0 則評論CWS 公開頁實測;chrome-extension skill 的 status 判讀為誤報

兩件要更正前次掃描的事。

第一,amem-clipper 已經上架了。前次掃描說「卡在 review 就等於沒有」, 那是 CWS API 的 uploadState: NOT_FOUND 造成的誤判。公開頁面查得到 v0.3.0。

第二,「verbatim archive」比前次掃描的描述更糟。前次說是「未整理的 DOM 傾印」, 問題是髒。實際上更嚴重:47% 的節點被截斷,而且沒有 raw 備份可以還原。 截斷後的 markdown 就是唯一一份副本。


2. Part A — 痛點排序

排序依據是「頻率 × amem 目前答得多差」。答得好的排在後面。

#痛點頻率amem 現況落差
1靜默的部分擷取與保真度崩壞7 repo,~90 comment,~70 reaction失敗最大
2存了找不回來4 repo + HN dominant失敗大
3權限過寬引發的供應鏈恐懼安裝決策點,有卸載實例失敗,且無謂大
4AI memory 錯誤、過期、無法驗證2026 最熱,HN 180 pts部分中
5離線時擷取不進去~92 reaction忽略中
6讀了不回頭的墳場HN 十年 dominant,tracker 僅 41 reaction忽略中(但被高估)
7登入頁與付費牆擷取全調查最高 reaction mass結構性解決無
8重複擷取6 repo,兩位維護者宣告做不到解決,但用資料遺失換反轉風險
9本機 LLM 連線摩擦5 repo,~80 comment部分小
10上架被下架的存亡風險有死亡案例已上架無

痛點 1 — 靜默的部分擷取與保真度崩壞(amem 答得最差)

陳述:擷取失敗最常見的形態不是報錯,是安靜地存了一個不完整或錯誤的東西。

一手證據:

  • kepano/defuddle#352(0 comment,2026-07-28,OPEN)— getElementSelector() 未 escape 冒號 id,React streaming SSR 觸發後 Defuddle 自己拋錯,靜默退回抓整個 <body>。這是「靜默部分擷取」的機制本體。 https://github.com/kepano/defuddle/issues/352
  • obsidianmd/obsidian-clipper#196(4 comment,2024-11-21)— jwhitley 回報存到了 頁面上從不可見的隱私聲明,可見內文一個字都沒存。他自己說「只發生在某些頁面, 我找不出規律」。https://github.com/obsidianmd/obsidian-clipper/issues/196
  • obsidianmd/obsidian-clipper#37(19 comment,32 reaction,CLOSED)— 「Download pictures to local」是該 repo 史上最高 reaction 的 closed issue。 遠端圖片連結會隨原站一起腐爛。https://github.com/obsidianmd/obsidian-clipper/issues/37
  • karakeep-app/karakeep#1652(16 reaction,OPEN)+ #1522(11 reaction)+ #594/#999/#1306 — 同一個 archive.org 中介需求被獨立提出五次,合計 27+ reaction。https://github.com/karakeep-app/karakeep/issues/1652
  • deathau/markdownload#366(7 comment,OPEN)— 「Only half of an article is extracted」,多位獨立回報者。一位指出根因是 Readability.js,另一位量化 「砍掉大約 16 段」。https://github.com/deathau/markdownload/issues/366
  • gildas-lormeau/SingleFile#1744(2 comment,OPEN)— Perplexity 與 ChatGPT 這類 對話視窗擷取被截斷。維護者:「It’s not an easy problem to deal with in a generic way.」https://github.com/gildas-lormeau/SingleFile/issues/1744
  • HN 44095939(bayindirh,2025-05-26)— 把需求講到最準:「Because I want the version I have seen. Not the edited/updated one.」 https://news.ycombinator.com/item?id=44095939

AMO 逐字評論獨立佐證同一件事。SingleFile 1033 則評分中 52 則低星,前三主題是 capture 不完整(4)、掉圖片與互動元素(3)、慢(2)。MarkDownload 有一則 3 星 (2023-06-17, alex):「it only converts what is visable on the screen rather than the whole page」。https://addons.mozilla.org/en-US/firefox/addon/single-file/reviews/

amem 現況:失敗,而且是最壞的一種失敗。

三個獨立缺陷疊在一起:

  1. 6000 字截斷,14/30 節點已中彈(clipper_bridge.rs:588)。
  2. 沒有 raw 備份,截斷後無法還原。
  3. 重剪整檔覆蓋,舊內容直接消失,沒有 diff、沒有版本、沒有 hash 可以偵測變化。

第三點值得單獨看。amem 用 URL 正規化把重剪折疊回同一個檔案,這解決了痛點 8。 但實作方式是覆寫,所以它同時製造了痛點 1。使用者兩週後重剪一個被改過的頁面, 第一版就永久消失了,而且系統不會告訴他。

對照 basicmachines-co/basic-memory#124(6 comment,4 reaction,OPEN)裡那句話: 「5 minutes into using it I realized I HAVE to have a mechanism for checkpoints… the LLM could easily mess up the memory in one prompt, and you want to be able to go back.」https://github.com/basicmachines-co/basic-memory/issues/124

這條落差最傷,因為 amem 的定位句押的就是這一格。全業界沒人做 chunk 級 hash, 這是真空象限。但 amem 目前只在 2/35 的面積上鋪了地基,而那 2 個還是不走 clipper 的 legacy PDF 節點。


痛點 2 — 存了找不回來(amem 答得差)

陳述:使用者的抱怨分兩種,檢索失敗與忘記自己存過。第二種更常被講。

一手證據:

  • HN 22105561(279 pts / 268 comment,2020-01-21)— hooande 一句話定義了它: 「My problem with bookmarks isn’t managing them, but remembering that they exist.」同 thread 的 JohnFen 補上為什麼「再 Google 一次」不算替代方案: 「Most of the things I bookmark are things that were really hard to find in the first place.」https://news.ycombinator.com/item?id=22105561
  • HN 44927588(jtqq,2025-08-16)— 自建 Emacs + org-roam + Elfeed + Wallabag 全套, 結論仍是「Retrievability is the biggest pain point right now」。下一步計畫是加 向量資料庫。https://news.ycombinator.com/item?id=44927588
  • karakeep-app/karakeep#2489(0 comment,2026-02-16,OPEN)— 「Full Text search STILL unavaliable… this seems as the most obvious use case」。六個月零回覆。 https://github.com/karakeep-app/karakeep/issues/2489
  • karakeep-app/karakeep#1955(4 comment,OPEN)— 部分關鍵字搜尋回傳不完整結果。 https://github.com/karakeep-app/karakeep/issues/1955
  • basicmachines-co/basic-memory#951(5 comment,OPEN,維護者自撰)— 有硬數字。 LoCoMo 1,986 個 query,single-hop R@5 basic-memory 0.486 vs mem0 0.592。 281 個獨有 miss 裡 146 是真 retrieval miss,主因是跨對話實體混淆。 https://github.com/basicmachines-co/basic-memory/issues/951
  • nashsu/llm_wiki#5(15 comment,OPEN)— 50GB+ 語料庫每次開啟 loading 極久, 且作者給不出建議上限。#397(3 reaction)— 6566 頁重建索引要 2 小時以上。 https://github.com/nashsu/llm_wiki/issues/397

amem 現況:失敗,而且失敗的方式很諷刺。

amem_recall 與 amem_ground 走 src/query/grep.rs,檔頭自述是「MVP search… Token-count scoring」。它把每個 wiki 檔整份讀進記憶體再數 token 命中次數。 35 個檔案可以,3,500 個不行。

諷刺的地方在於:BM25 加密集向量檢索已經寫好了。src/refindex.rs 948 行、 src/embed.rs 519 行,含 RRF 融合、bibliography chunk 過濾、向量指紋校驗。 但它服務的是 Zotero 的 78 篇 PDF 語料,不是 wiki。

也就是說,好的檢索建在對的技術上、錯的語料上。而 wiki 才是 clipper 的產物, 是產品入口。

~/.amem/ 沒有 index.md 也沒有 log.md,## Connections 在 31/35 個節點裡是空的。 所以「忘記自己存過」這半邊的痛點,amem 連最便宜的解法都沒有。


痛點 3 — 權限過寬引發的供應鏈恐懼(amem 答得差,且是無謂的差)

陳述:反對的不是 telemetry,是 <all_urls> 加 scripting 這個後門面積。使用者 會因此卸載,並且會寫下來。

一手證據:

  • karakeep-app/karakeep#2782(9 comment,2026-05-10,CLOSED)— 全調查最好的 單一物證。使用者 psla:「This permission is a no-go from me.」接著:「As much as I trust you and the project, supply chain attacks are a thing, and I generally don’t allow extensions that request full control over the page. Have you considered using optional permissions?」維護者 24 小時內改成 optional_permissions 並發 1.2.11。https://github.com/karakeep-app/karakeep/issues/2782
  • karakeep-app/karakeep#2790(2026-05-12,CLOSED)— 另一位獨立提出,附威脅模型: 「its a great backdoor if repo ever gets compromised」,並標 all_urls 🔴 HIGH、 scripting 🔴 HIGH。https://github.com/karakeep-app/karakeep/issues/2790
  • gildas-lormeau/SingleFile#1030(7 comment,OPEN 四年)— 「if I install this extension on Firefox, it will inject a script into every page I load even if I don’t intend to save the web page」。維護者答:必須永遠注入,無法關閉。 https://github.com/gildas-lormeau/SingleFile/issues/1030
  • obsidianmd/obsidian-clipper#165(2024-11-14)— brian-burton 因 0.9.6 新增下載權限 而卸載,並在 template 欄位寫「Sorry, I can’t provide a template because I’ve uninstalled Web Clipper」。https://github.com/obsidianmd/obsidian-clipper/issues/165
  • AMO 逐字,Instapaper 1 星(2025-08-30, dev421):「The permissions required by this add-on are ridiculous. “Access your data for all websites”!? You only need the current tab URL!」https://addons.mozilla.org/en-US/firefox/addon/instapaper-official/reviews/

amem 現況:失敗,而且有一半的權限沒有換到任何功能。

manifest 有 <all_urls> 加 12 個 permission,含 tabCapture、downloads、 scripting、tabGroups。這正是讓 psla 拒裝 karakeep 的權限輪廓。

更糟的是四個 always-injected content script 裡有三個是 8 行 stub。 claude-autosave.js、chatgpt-autosave.js、gemini-autosave.js 都只做一件事: 設一個 flag 然後 console.debug。它們對 claude.ai、chatgpt.com、gemini.google.com 永遠注入,換到零功能。

這是純負債。使用者付出「這個 extension 在我的 AI 對話頁上跑程式」的信任成本, 產品沒有拿到任何東西。真正的 AI 對話擷取在 extractors/ 裡,是按需 executeScript 派送的,不需要宣告 content script。

補一個對 amem 有利的觀察。維護者本人有時是隱私鷹派,而使用者站在他那邊。 kepano 在 #112 以安全理由永久拒絕雙向整合:「I don’t think most Obsidian users would trust giving their browser full access to their Obsidian vault.」同 thread 的 writtenfool 附議:「I would not want, at anytime, for a web app to have access to my local app.」amem 的 daemon 是本機的,這是可以講的故事,但前提是權限面要乾淨。


痛點 4 — AI memory 錯誤、過期、無法驗證(amem 部分答)

陳述:2026 年最熱的一條。抱怨不是記不住,是自動記下了錯的東西再拿去誤導使用者。

一手證據:

  • HN 48776232「Memorizing session transcripts isn’t useful」(180 pts / 159 comment, 2026-07-03)。Fabricio20 最完整:「I specifically disabled claude memory in a project because it kept writing down thigns to memory that didn’t need to be in memory, including severly wrong statements that then would confuse it later.」 他還指出自動記憶被自動重新啟用,要同時關 autoMemoryEnabled 與 autoDreamEnabled。https://news.ycombinator.com/item?id=48776232
  • 同 thread,mastax 指出污染的正回饋:「If you allow any low value things into memory, Claude will notice that established pattern and start trying to add low value memories」。semiquaver:「Claude’s own memories severely mislead it」。
  • HN 47531882(ravikirany22,2026-03-26)— 唯一有量化的一筆:「We’ve been auditing TypeScript repos and finding 10-84% of symbol references in AI config files are stale… it’s getting a confident lie.」 https://news.ycombinator.com/item?id=47531882
  • nashsu/llm_wiki#458(2 reaction,OPEN)— 「Concerns About LLM-Based Knowledge Bases: Hallucinations and the Need for Faithful Original Text Retrieval」: 「When I query specific regulations or raw tables, the output is often hallucinated… frequently diverges significantly from the original content.」 https://github.com/nashsu/llm_wiki/issues/458
  • nashsu/llm_wiki#229(5 reaction,OPEN)— 要一個修正入口。使用者指出唯一的修復 路徑是手改 markdown,而這與「Wiki 頁面全部由 LLM 維護」的核心設計理念相悖。 https://github.com/nashsu/llm_wiki/issues/229

這裡有一個對 amem 極重要的相關性,是 GitHub 挖掘得到的最有價值推論之一。 幻覺抱怨只出現在 llm_wiki 一家。 其他八個 repo 的 LLM 只寫 tag 與 summary, 不寫知識本體,所以沒有幻覺抱怨。llm_wiki 是唯一讓 LLM 撰寫 wiki 的,也是唯一 被抱怨幻覺的。

amem 的 clipper 節點把 LLM 產物(tldr)與原文(Content)並置。這個結構比 llm_wiki 安全。但這個安全性依賴原文真的在旁邊,而 47% 的節點原文被截斷了。

amem 現況:部分答,且答案正在漏氣。

有的部分:frontmatter 帶 url、first_seen_at、captured_at 清單。這比 Claude Code 內建 memory 多(後者只記 modified 時間戳,不記來源)。tldr 與 Content 分離 也對。

漏的部分:沒有內容 hash,所以無法偵測來源頁面變了。amem_ground 的 tool description 自述是「search the local wiki for the topic and return JSON hits」 (src/mcp/mod.rs:172),這是對自有 wiki 的檢索加引用,不是驗證。 SPEC 定義的 amem_factcheck 仍在 RFC-003 規劃中。

refverify(CrossRef + DataCite)與 refcheck(BM25 + 向量)是真的驗證, 而且做得好。它們只服務 PDF 語料。


痛點 5 — 離線時擷取不進去(amem 忽略)

陳述:擷取的那一刻經常是離線的,沒有工具會排隊。

這是本次調查最意外的一條,前次掃描與 design memo 都沒有列。

一手證據:

  • karakeep-app/karakeep#1077(9 comment,52 reaction,CLOSED)— 「Offline cache on Mobile app」。 https://github.com/karakeep-app/karakeep/issues/1077
  • karakeep-app/karakeep#2401(17 comment,20 reaction,2026-01-14)— 自架者離開 LAN 之後,分享選單一直轉。 https://github.com/karakeep-app/karakeep/issues/2401
  • karakeep-app/karakeep#274(3 comment,20 reaction)— 「Cache Hoards while away from server」。合計約 92 reaction。 https://github.com/karakeep-app/karakeep/issues/274
  • obsidianmd/obsidian-clipper#828(0 comment,2026-05-05,OPEN 三個月無人回)— 同一個形狀的靜默資料遺失:「In my case I lost ~10 job-prospect clips this morning before realizing none of them had hit disk. There is no recovery path other than redoing the captures from browser history.」 https://github.com/obsidianmd/obsidian-clipper/issues/828

amem 現況:忽略,而且暴露面比 obsidian-clipper 更大。

background.js 唯一的 alarm 是 amem-bridge-reconnect,只服務 bridge 輪詢, 不是擷取佇列。沒有 retry、沒有 navigator.onLine 判斷、沒有 pending capture 儲存。

amem 的架構讓這件事比競品更常發生。obsidian-clipper 只需要 Obsidian 這個 app 存在。amem 需要 amem mcp serve 這個 daemon 正在跑。daemon 沒跑、剛重開機、 或 crash 了,使用者按下擷取就是失敗。而擷取失敗的當下,使用者已經離開那個頁面了。


痛點 6 — 讀了不回頭的墳場(amem 忽略,但這條被高估了)

陳述:痛感真實且橫跨十年,但它不是使用者會為之投票的痛點。

HN 證據極厚,判定 dominant:

  • HN 44066646(rossant,2025-05-22,母 thread 1222 pts / 761 comment)— 「I just exported my data and found 13,000 unread articles out of a total of 34,000.」https://news.ycombinator.com/item?id=44066646
  • HN 46880996(al_borland,2026-02-04)— 「They are where my good intentions go to die.」https://news.ycombinator.com/item?id=46880996
  • HN 16306040 與 HN 44925963(pixelmonkey,相隔七年半講同一句)— 「I sometimes describe Instapaper as ‘/dev/null for web content’.」 https://news.ycombinator.com/item?id=44925963
  • HN 36146108(bsnnkv,2023-05-31)— 「both note apps and read it later queues/apps are where ideas go to die.」https://news.ycombinator.com/item?id=36146108

但 tracker 的 reaction 質量說了另一件事。

resurfacing 相關的三個 issue 合計 41 reaction(karakeep #435 10、#705 16、 #863 15)。同一個 repo 裡,單一個登入擷取 issue(#172)就有 57 reaction, 離線佇列合計 92 reaction。

而且十個 repo 裡沒有任何一個有「我的存檔是墳場」這種 issue。它只以間接形式 出現,例如 karakeep#863:「we keep bookmarking the pages or videos, but we do not have time to fully read it.」

這對 design memo 是一個修正。 memo 把 #3「write-only graveyard」與 #6「friction」 當成「真正殺死 clipper 的兩個痛」。HN 支持墳場的存在,但 tracker 說使用者不為它 投票,他們為「存不進去」和「找不到」投票。

amem 現況:忽略。 沒有 resurfacing、沒有 index.md、沒有隨機回顧, ## Connections 是空的。但依據上面的證據,這應該排在補完痛點 1 與 2 之後, 不該當頭條。


痛點 7 — 登入頁與付費牆擷取(amem 結構性解決)

陳述:server 端 crawler 永遠看不到登入使用者看到的東西。每個專案都獨立重新發現 唯一解是「從使用者自己的瀏覽器擷取」。這是全調查最高 reaction mass 的主題。

一手證據:

  • karakeep-app/karakeep#172(52 comment,57 reaction,CLOSED)— karakeep 史上最高 reaction 的 issue,「Local Scraper (Use browser auth)」。 結局是改成相容 SingleFile extension 的 REST endpoint。server 端 crawler 輸給了 瀏覽器 extension。 https://github.com/karakeep-app/karakeep/issues/172
  • karakeep-app/karakeep#414(66 comment,40 reaction,CLOSED)— 存到 cookie 同意視窗而非文章。有使用者的本機 Llama3.2 去摘要了付費牆公告。 https://github.com/karakeep-app/karakeep/issues/414
  • karakeep-app/karakeep#2814(4 comment,2026-05-16,OPEN)— 最關鍵的一筆。 karakeep 已經出貨 client-side crawling,付費牆仍然失敗:「Expected: Full article text is archived. Actual: Only the article preview/teaser is saved.」 他的 workaround 是用 Obsidian Web Clipper 抓,再寫 Python 腳本同步進 karakeep。 https://github.com/karakeep-app/karakeep/issues/2814
  • karakeep-app/karakeep#2885(2026-06-13,維護者本人開)— 「As part of a recent reddit crackdown on crawlers, the endpoint that we were using… is now completely blocked.」https://github.com/karakeep-app/karakeep/issues/2885
  • 2026 年的新退化不只 Reddit:#2952 archive.is 又要 captcha(2026-07)、 #2423 Cloudflare 封鎖(2026-01)、#2381 JS/cookie 攔截頁被存下來(2026-01)。

amem 現況:結構性解決,這是最該押的一格。

content script 跑在使用者已登入的分頁裡,沒有 cookie 轉移問題,沒有 server 再抓一次 的問題。~/.amem/wiki/ 裡真的躺著 Gmail 與 openreview 登入後才看得到的 capture。

趨勢對 amem 有利。bot wall 在 2026 年系統性收緊,server-side crawler 這條路正在 失效,而「已登入的真實瀏覽器」正在變成唯一可行路徑。

倫理阻力方面,一手來源查到的是零。使用者一律把它當純能力缺口。 最常見的自我正當化是「我付錢看的」(#2236,bronikowski:「when I read something I paid for」)。

但這一格已經有人在賣了。 見 §3 的 Web2MD。


痛點 8 — 重複擷取(amem 解決,但用資料遺失換)

陳述:沒有工具會在你重存之前告訴你「你已經存過了」。而且有兩位維護者公開宣告 這件事做不到。

  • obsidianmd/obsidian-clipper#112(8 comment,16 reaction,OPEN)— kepano 是架構性拒絕:「For now there are no plans to create a two-way integration… This is intentional.… I think it poses a security risk.」 #323 與 #521 兩個同樣的請求都被 closed as duplicate。三次請求,一次永久拒絕。 使用者 CarcajadaArtificial:「I just want to not end up with duplicate web clippings. That’s all.」https://github.com/obsidianmd/obsidian-clipper/issues/112
  • gildas-lormeau/SingleFile#1642(OPEN)— 維護者:「This is not really possible from a technical point of view, for privacy reasons. Extensions cannot scan folders on the filesytem.」#956 同樣答「A database is required… it seems complicated to implement reliably」。 https://github.com/gildas-lormeau/SingleFile/issues/1642
  • karakeep-app/karakeep#486(10 comment,7 reaction,CLOSED)— 十家裡唯一出貨 「已存過」指示器的,花了 13 個月。 https://github.com/karakeep-app/karakeep/issues/486
  • karakeep-app/karakeep#864(3 reaction,OPEN)— 去重不處理結尾斜線, example.com/a 與 example.com/a/ 算兩筆。 https://github.com/karakeep-app/karakeep/issues/864

amem 現況:解決了,而且解得比業界好,但代價是痛點 1。

node_id_from_url(clipper_bridge.rs:425-471)做的事比 karakeep 多: 剝除 11 個追蹤參數、小寫 host、去尾斜線、arxiv/github/HN/x.com 各有專屬 id 規則。karakeep #864 抱怨的尾斜線問題,amem 在 url_canonical:521 已經處理。 #633 抱怨的追蹤參數,amem 在 :500-503 已經處理。

這是一個乾淨的勝場。兩位維護者公開說做不到的事,amem 因為有本機 daemon 而做得到。

但實作用覆寫達成去重。 舊內容消失,沒有版本、沒有 hash。所以 amem 把「重複 spam」換成了「靜默資料遺失」。這兩個痛點在 amem 身上是同一行程式碼的兩面 (clipper_bridge.rs:381 的 std::fs::write)。


痛點 9 — 本機 LLM 連線摩擦(amem 部分答)

陳述:「接你自己的模型」是 AI memory 工具的第一大死路。錯誤訊息很籠統, timeout 不可設定。

  • karakeep-app/karakeep#185(20 comment,OPEN 27 個月)— 標題就是 「How to verify hoarder app is working with the local ollama」。 https://github.com/karakeep-app/karakeep/issues/185
  • karakeep-app/karakeep#424(19 comment,OPEN)— 使用者拿到的 log 是 inference job failed: TypeError: fetch failed,沒有 host、沒有 status。 維護者:「Unfortunately the logging does not show where the issue happens.」 https://github.com/karakeep-app/karakeep/issues/424
  • MODSetter/SurfSense 有十個 open 的本機 LLM 連線 issue(#1379、#1616、 #1550、#587、#517、#1464、#1405、#1394、#1493、#1518)。 https://github.com/MODSetter/SurfSense/issues/1379
  • obsidianmd/obsidian-clipper#515(2 reaction,OPEN)— 「Custom Ollama provider requires non-existent API key」,UI 要一個該 provider 根本沒有的憑證。 https://github.com/obsidianmd/obsidian-clipper/issues/515
  • karakeep 有五個獨立 open issue 都是同一件事:本機模型比硬寫的 timeout 慢 (#2770、#2679、#2994、#1129、#1806)。

amem 現況:部分答,而且方向對。

clipper_bridge.rs 的 summarize 有三條路徑:summarize_via_claude(:256)、 summarize_via_codex(:286)、summarize_via_ollama(:314)。有 CLI fallback 是 對的設計,使用者不必先裝 ollama 才能用。這比 SurfSense 與 karakeep 好。

未查證的部分:三條路徑全掛時的錯誤是否對使用者可讀。依據痛點 5 的分析, 擷取失敗沒有佇列,所以錯誤處理路徑值得單獨測。


痛點 10 — 上架被下架的存亡風險(amem 已上架)

陳述:clipper 的死亡證明是 Google 簽的,不是 bug 數量。

  • deathau/markdownload#378(9 reaction,OPEN)— 該 repo 最高 reaction 的 open issue,標題是「This extension is no longer available because it doesn’t follow best practices for Chrome extensions」。同一件事被獨立提報三次 (#357 3 reaction、#393 1 reaction),合計 13 reaction。3,997 star 的專案 死於下架。https://github.com/deathau/markdownload/issues/378
  • #393 裡有陌生人貼未審核的 fork,下一位留言者回「now the deployments website is unable, the zip is not found」。使用者被推向已經死掉的未簽名 fork。
  • HN 46880866(2026-02-04)— 「most of the things I come across are dead and gone, or seem abandoned somehow」。2026 年有三個 Show HN 把死亡寫進標題: 「because the others keep dying」(48745735)、「because Mozilla killed it」 (46956985)、「built after Pocket shut down」(48449568)。 https://news.ycombinator.com/item?id=46880866

amem 現況:已上架,前次掃描這一格判錯了。

CWS 公開頁查得到 amem Clipper v0.3.0,更新 2026-08-13,3 users,0 則評論。 llm_wiki 的 extension 至今仍要 chrome://extensions load unpacked。這一格 amem 贏。

順帶一個 null result:「load unpacked / Developer Mode」在十個 repo 裡是零筆 issue。 這不是真實使用者痛點。design memo 用 Developer Mode 摩擦當作不做 chrome.userScripts 的理由,那個結論仍然對(安全面與品牌面成立),但摩擦論據 本身沒有一手支持。


3. Part B — 前次掃描漏掉的競品

只列前次掃描的搜尋角度會漏掉的。星數與日期我親自用 gh api 於 2026-08-17 複驗。

3.1 Web2MD — 架構逐項複製,且已在賣

https://web2md.org/

今天出貨: Chrome extension、16 個站台專用 extractor(含 Reddit、YouTube、 GitHub、arXiv)、npx web2md CLI、npx web2md-mcp-server MCP server、 GPT-4 與 Claude token 計數、watch mode。

Agent Bridge 是核心賣點,官網原文:「Agent Bridge uses your actual Chrome with your cookies and login state — Reddit can’t tell the difference from normal browsing」。

這是 amem 痛點 7 那格優勢的逐字複製,而且已經包裝成產品在賣。 定價 Pro $4.17/月(原價標 $15)。免費層每日 3 次轉換。非開源。

單一勝出軸:CLI 批次同步。 amem 是一頁一頁點,Web2MD 有 web2md sync。

需注意:前次掃描漏掉它,是因為它不在 GitHub 上,搜 repo 搜不到。它也是 web2md.org SEO 內容農場的擁有者,該站產出偽裝成「honest review」的競品評測。 評論挖掘的子代理明確標記了這一點並拒絕採用其內容。

3.2 Letta — SKILL.md 自動生成,全自動且已出貨

https://github.com/letta-ai/letta-code — 3,017 star,v0.30.25 發佈於 2026-08-17(今天)。

我親自讀了 src/agent/subagents/builtin/reflection.md,逐字引用:

name: reflection description: Background agent that reflects on recent conversations to update memory and maintain skills

Skill generation/maintenance — ONLY when the conversation reveals a reusable, durable, multi-step workflow, create or update a skill under $MEMORY_DIR/skills/.

Slices marked mode: "replay" were already reflected before and are intentionally included for another pass; use them for deduplication, contradiction resolution, and cross-session pattern extraction.

最後一段特別重要。它殺掉的不只是「自動生成 SKILL.md」,還包括「跨 session 縱向 recurrence detection」這個備援縫。Letta 明文在做跨 session 模式抽取。

它還有 amem 沒設計的東西:skill 生命週期回收。操作集是 update / extend / deprecate / split / create,寫進 MemFS git repo,有版本。

單一勝出軸:同一個 artifact,更完整的生命週期,已在有資金的產品內出貨。

它缺的是 capture surface。它的語料是對話,不是網頁。

3.3 OpenAI Computer History — 廠商層的 sensor → recurrence → skill

https://learn.chatgpt.com/docs/customization/computer-history

官方文件原文:「Computer History turns your activity across apps and websites into memories and a timeline that ChatGPT and Codex can reference.」以及 「When Computer History notices repeatable work, a timeline entry can suggest a skill or automation.」

這是「跨網站活動 → 偵測重複 → 建議 skill」,由模型廠商出貨。

限制就是 amem 剩下的縫:僅 macOS 桌面版、Pro/Business/Enterprise、預設關閉、 EEA/瑞士/英國不可用。而且它讀 interaction event(點擊、輸入、app 切換), 明確不讀頁面內容本身。

單一勝出軸:通路加作業系統級 sensor。

3.4 其他架構撞車的

名稱出貨成熟度(gh api 複驗)單一勝出軸
RowboatMac/Win/Linux 桌面 app,Apache-2.017,290 star,pushed 2026-08-17論述撞車。README 寫「living Obsidian-style backlinked knowledge graph」「All data is stored locally as plain Markdown」
claude-memCLI,v13.15.290,938 star,pushed 2026-08-17「捕獲即記憶」的心智位置佔有量。無 skill 生成、無網頁擷取
openhumandmg/exe/deb/AppImage36,323 star,created 2026-02-18發版速度近乎每日
open-knowledge桌面 app + CLI,GPL-3.03,480 star,created 2026-06-03自述「AI-native markdown IDE and LLM wiki」,前端成熟度
screenpipe桌面20,980 star,YC S26捕獲面總量最大,隨時可加瀏覽器 lane

Rowboat 有一點要更正子代理的判讀。它的內建瀏覽器隔離於使用者主瀏覽器, README 明寫使用者必須在裡面重新登入。這是 amem 的優勢,不是 Rowboat 的。

3.5 格式層風險

GoogleCloudPlatform/knowledge-catalog — 8,667 star,created 2026-05-04,pushed 2026-08-15。Google 把「LLM 編譯的 markdown wiki」變成規格。子代理報告 v0.2 加入 trust signals,這直接踩進 amem 的 citation-grounding 差異化。該 v0.2 細節我未親自複驗。建議獨立評估 OKF 相容性。

3.6 負面結果(掃過、確認沒有)

這些空白是 amem 剩餘的可辯護空間。

  • MCP 生態沒有 clipper。 punkpeye/awesome-mcp-servers(92,462 star)全文搜 clipper / web clip / clip page 零筆。modelcontextprotocol/servers README 無任何 clipper server。
  • 大型 memory 玩家都沒有 capture surface。 graphiti、letta、memU、cognee、 txtai、memoripy、memento-mcp 的 repo 樹內都沒有 browser extension 目錄。 mem0ai/mem0-chrome-extension 已 archived,最後 push 2026-03-23。 最大玩家退出了瀏覽器擷取賽道。
  • browser-control MCP 沒有一個寫 KB。 chrome-devtools-mcp(49,285)、 browser-use(109,480)、playwright-mcp(36,195)、Skyvern、stagehand 全部純自動化、 零持久化。
  • 「從我剪過的東西裡找跨篇反覆出現的模式」沒有人做。 搜過 chrome extension SKILL.md、web clipper skill agent、capture to skill、 bookmarks to agent skill、recurring patterns into skills 全部零相關結果。
  • YC 四個 2026 batch 共 638 家,零家做「瀏覽器 extension + 本機 markdown wiki」。 (子代理用 YC 自家 API 分頁,我未複驗)

4. RFC-007 的 moat 判決:GONE

RFC-007 寫「Every clipper competitor stops at storage; nobody closes the loop into agent capability. Step 3 is the moat.」

前次掃描已經把它修正成「moat 不是 step 3 整段,而是 step 3 裡的自動生成那半段」。 那個修正現在也不成立了。 三份證據,全部我親自複驗。

證據一:Anthropic 官方說這件事不需要工具。

https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices 逐字:

Claude models understand the Skill format and structure natively. You don’t need special system prompts or a “writing skills” skill to get Claude to help create Skills. Simply ask Claude to create a Skill and it generates properly structured SKILL.md content with appropriate frontmatter and body content.

護城河不能建在供應商文件標註「此處不需要工具」的動作上。

證據二:Claude Code 二進位檔本身就會自動寫 SKILL.md。

我在本機 /Users/ydwu/.local/share/claude/versions/2.1.233 裡抓到字串, 逐字(去掉 UTF-16 間隔):

If this repo has no project verify skill (.claude/skills/verify/SKILL.md), that is a reason to run /verify, not to skip it: the run creates that file, saving the working build-and-drive recipe for future sessions.

同一份二進位檔裡 run-skill-generator 出現 7 次,註冊為 bundled skill, menuDescription 是「Create a skill that knows how to run this project’s app」。

/verify 寫檔是跑 verify 的副作用,使用者沒有要求產生 skill。 這就是無人在迴路的自動蒸餾,而且它是內建的、免費的。

證據三:Letta 已全自動出貨同型機制。 見 §3.2 的逐字引用。

證據四:「X → SKILL.md」已是 commoditized 類別。 星數我逐一用 gh api 複驗:

專案Star做什麼
microsoft/SkillOpt16,074把 SKILL.md 當可訓練參數優化,含 nightly self-evolution
yusufkaraaslan/Skill_Seekers14,77318 種來源(docs/repo/PDF/EPUB/YouTube)→ SKILL.md
bergside/design-md-chrome2,663Chrome extension 剪一頁 → 產 SKILL.md

最後一列要特別看。它已經佔住「Chrome extension 剪頁 → 吐 SKILL.md」這個一模一樣 的手勢。

還剩什麼

一條窄縫,是 feature 不是 moat:沒有人從「跨站累積數週的 clipped web pages」這個 語料出發做蒸餾。

三家的語料都不是網頁。Anthropic 的 capture surface 是螢幕錄影、當下對話、本機 repo。 Letta 的是對話 transcript。OpenAI Computer History 讀 interaction event,明確不讀頁面 內容。design-md-chrome 是一頁換一檔,沒有語料庫概念。

這條縫真實存在,但它撐不起「moat」這個詞。它撐得起一個功能。

建議:不要再把「產生 SKILL.md」寫進對外定位。


5. 建議:第一個該改的東西

先讓 clipper 節點停止破壞證據,再談任何新功能。

理由是這一個改動同時關掉排名第 1、第 2、第 4、第 8 四個痛點的落差, 而且它是 amem 唯一真空象限(chunk 級 hash provenance)的地基。

具體是三件小事,都在 clipper_bridge.rs 一個檔案裡:

  1. 移除 6000 字截斷,或把完整原文寫進 ~/.amem/raw/。 目前截斷後無副本可還原 (:588)。
  2. frontmatter 加 content_sha256,並對 chunk 逐段 hash。 cite.rs 已有 text_sha256,PDF 路徑在用,clipper 路徑沒接上。
  3. 重剪時若 hash 改變,另存版本而不是覆寫。 目前 :381 直接 std::fs::write。

做完這三件,「amem 記錄你真正讀過的東西,而且每一句話都能追回原文」才從願景變成 產品描述。在那之前,這句話的後半是不實陳述。

第二順位是把 refindex 的 BM25 加向量檢索接到 wiki 語料上。程式碼已經寫好了 (refindex.rs 948 行、embed.rs 519 行),只是指向錯的語料。這關掉痛點 2。

第三順位是刪掉三個 stub content script。這是零成本降低權限面(痛點 3), 它們目前換到零功能。


6. 取樣缺口與未查證項目

誠實揭露,這些不要當成已經查過。

  1. Reddit 完全未取樣。 www.reddit.com、old.reddit.com、api.reddit.com 三個 domain 都被 harness 阻擋。r/ObsidianMD、r/PKMS、r/selfhosted、r/DataHoarder、 r/LocalLLaMA、r/ClaudeAI 全部沒有覆蓋。自架與 DataHoarder 族群的觀點缺口很大。
  2. Chrome Web Store 的 1-3 星評論全部拿不到。 CWS 從瀏覽器核心層禁止 content script 注入 chromewebstore.google.com(bridge 回 "The extensions gallery cannot be scripted."),這條路永久不可行,不是 bug。 WebFetch 只回傳預設「最相關」排序,幾乎全 5 星。受影響最大的是 Recall (getrecall.ai),它是 amem 最直接的 AI-summary 競品,負評完全沒有資料。
  3. AI 摘要品質主題嚴重取樣不足。 AMO 唯一有 AI 的樣本是 Raindrop 的 2 則。 不要從這 2 則推論任何結論。
  4. AMO 樣本有 Firefox 結構偏誤。 31 則登入/cookie 抱怨有相當比例是 Firefox 跨站 cookie 預設造成的,不能直接外推到 Chrome。
  5. 「AI 摘要在筆記庫裡幻覺」找不到第一手案例。 forum.obsidian.md 搜 「AI summary hallucinate」零筆。obsidian-clipper 的 Interpreter 相關 issue 全部是 整合層故障,沒有一筆抱怨輸出內容錯誤。含意:使用者對 agent memory 的幻覺很敏感, 對 clip 時的摘要幻覺還沒有痛感,可能因為原文還在旁邊。
  6. 付費牆擷取的倫理/ToS 抱怨:零筆。 這是 null result,不是沒查。
  7. notes/wiki → 自動生成 skills:完全沒有人要求。 使用者的行為是從 session 事後蒸餾 skill(HN 47543139,627 pts / 265 comment,ccosky:「Anytime I do something as a one-off that I know I’ll do in the future, at the end of the session I’ll ask Claude to write a new skill based on what it did」)。 方向與 RFC-007 的假設相反。
  8. 未親自複驗: OKF v0.2 的 trust signals 細節、YC batch 歸屬與家數、 Minibase 與 LLMnesia 的使用者數、Gemini CLI 那個「require recurrence evidence before extracting skills」PR 的編號。
  9. amem 三條 summarize 路徑全掛時的錯誤可讀性未測。 依痛點 5 的分析值得單獨測。