formace-00 — Zotero reference corpus (ops)
Set up 2026-08-16. Backend for issue #24 (Zotero data dir as the librarian’s reference corpus, fine-grained ref check).
Where things live
| What | Path | Disk |
|---|---|---|
| Zotero 7 program (140.12.0esr) | /mnt/storage1/users/ydwu/opt/zotero | big disk (19T) |
Zotero data dir (zotero.sqlite, storage/*.pdf) | /mnt/storage1/users/ydwu/zotero-data | big disk |
| Zotero profile (prefs) | ~/.zotero/zotero/5sun4qjy.default | system disk |
| systemd user unit | ~/.config/systemd/user/zotero.service | system disk |
| amem corpus (source of the import) | ~/.amem/raw (157 files, 78 with metadata) | system disk |
Nothing of size sits on / — it runs ~82% full (319G free of 1.8T).
Service
systemctl --user status zotero # active/enabled
systemctl --user restart zotero
journalctl --user -u zotero -n 50
loginctl enable-linger ydwu is set, so the unit starts at boot without a
login session. Zotero is a GUI app, so the unit runs it under xvfb-run.
Library contents
78 papers + 78 PDF attachments, imported from ~/.amem/raw/*.meta.json by
generating BibTeX with file = {…} fields and POSTing to
http://127.0.0.1:23119/connector/import. Cite keys match amem’s
(belrose2023eliciting etc.), so Zotero items and amem wiki nodes line up.
Read the library:
curl -s "http://127.0.0.1:23119/api/users/0/items?limit=200"
curl -s "http://127.0.0.1:23119/api/users/0/collections"
The local read API is off by default; it is enabled via
extensions.zotero.httpServer.localAPI.enabled in the profile’s prefs.js.
Security — this is a shared machine
formace-00 has 25+ accounts (a class of students). Three exposures were found and closed on setup; keep them closed:
- Zotero’s HTTP server on
127.0.0.1:23119has no authentication. That is fine on a personal desktop and a data leak here — on a multi-user box localhost is not a security boundary, so any local account could read the whole library. The unit gates the port by uid via an iptables owner match, added inExecStartPreand removed inExecStopPost(so it needs noiptables-persistentand cannot outlive the service). zotero-datawas created world-readable (drwxrwxr-x) → nowdrwx------.~/.amemwas also world-readable (drwxrwsr-x) → nowdrwx------. The 157 raw PDFs and every wiki node were readable by every account.
New directories under ~ inherit a group-writable umask on this host — check
permissions after creating anything that holds corpus or credentials.
Gotchas hit during setup
pkill -f "opt/zotero/zotero"kills your own SSH command, because the pattern matches the command line running it. The script dies silently mid-way. Usepkill -x zotero-bininstead.- Zotero’s Linux download is served as
.tar.bz2but is actually XZ;tar -xfauto-detects, an explicit-jfails. - First launch must be given a display (
xvfb-run) to createzotero.sqlite; it never exits on its own, so run it undertimeoutwhen you only want initialisation. - Do not read
zotero.sqlitewhile Zotero holds it. Prefer a better-bibtex.bibauto-export (plain text, no lock) as the metadata source, or copy the sqlite first.
Chunked reference layer (built 2026-08-16)
src/refindex.rs in amem-librarian. Paragraph-level chunks, each hashed, with
BM25 retrieval — no model, no GPU, so the pipeline is verifiable end to end
before embeddings go on top.
export PATH="$HOME/.cargo/bin:$PATH" # cargo is NOT on the default ssh PATH
AM=/mnt/storage1/users/ydwu/cargo-target/release/amem
$AM refindex build --bib /mnt/storage1/users/ydwu/zotero-data/amem-corpus.bib
$AM refcheck "<a sentence from your draft>" [--cite-key rafailov2023direct] [--limit 5]
State on formace-00: 78 refs → 12,751 chunks, 13 s, index at
~/.amem/refindex/chunks.jsonl.
What it is for: a hit carries cite_key#p<page>c<n>, the page, a SHA-256 and
the query terms that actually matched, so “this paper does not support that
sentence” becomes visible rather than a hunch. Measured on the Mac corpus, a
DPO claim scores 32.2 against rafailov2023direct and 5.9 when misattributed
to vaswani2017attention.
Content addressing immediately paid for itself — identical hashes across different keys exposed duplicate papers in the corpus:
| identical chunks | same paper filed twice as |
|---|---|
| 197 | wei2022chain = wei2022cot |
| 156 | dpo_2023 = rafailov2023direct |
| 97 | merrill2023expressive = merrill2024expressive |
Worth deduping before the corpus grows; 12,751 chunks carry only 12,129 unique hashes.
Cross-machine ref check (built 2026-08-16)
amem refserve — a separate service from clipper_bridge on purpose. That
one drives the user’s real logged-in Chrome, so exposing it off-host would hand
a remote caller the browser. This one is read-only over the chunk corpus, so it
is the only piece with any business on a non-loopback address.
| service | amem-refserve.service (user unit, linger already on) |
| bind | 100.86.146.58:7602 — the tailnet address, not 0.0.0.0, so the lab LAN cannot see it |
| auth | bearer token, constant-time compare; token in vault as AMEM_REF_TOKEN, on the box at ~/.config/amem/refserve.env (mode 600) |
| guardrail | the server refuses to start on a non-loopback bind with no token, or a token under 16 chars |
From the Mac:
export AMEM_REF_TOKEN=$(bash ~/.claude/skills/vault/run.sh get AMEM_REF_TOKEN)
amem refcheck "<sentence from your draft>" --remote http://100.86.146.58:7602 \
[--cite-key rafailov2023direct] [--limit 5]
Verified from the Mac: no token → 401, wrong token → 401, valid token →
12,751 chunks searched. --cite-key vaswani2017attention on a compositionality
claim correctly returns NO SUPPORT, which is the wrong-citation signal working
across the network.
GET /status is deliberately unauthenticated and deliberately boring
({ok, chunks, refs}) so a health probe reveals nothing about the corpus.
Bibliography chunks are excluded
Reference lists mention every topic in the field while asserting none of them,
and they score well — a claim phrased like a paper’s title matches that title
in every bibliography citing it. Observed live: querying the CoT paper’s own
title returned three other papers’ reference lists above any real text. Chunks
are now flagged by signal-counting (arxiv:/doi:/bare URLs weighted double,
et al., In Proceedings, and 2022./2022b. year stamps) and skipped at
search time. 16.5% of the formace corpus flags; after the fix the same query
returns the actual paper’s title page first.
Dense retrieval (built 2026-08-16)
BM25 alone only finds passages sharing words with the claim. Measured: “transformers
struggle to compose multiple reasoning steps” never surfaced dziri2023faith, the
paper that argues exactly that, because it says “compositionality” and “error
propagation”. With vectors, all three top hits are that paper — including its
“Error Propagations: The Theoretical Limits” section.
ollama pull nomic-embed-text # 768-d, ~274 MB
$AM refindex embed # 12,994 vectors in 4m11s on the RTX 6000
Vectors live at ~/.amem/refindex/vectors.bin (raw LE f32, fixed stride, ~38 MB)
with vectors.meta.json carrying a fingerprint of the corpus they were built
from. A mismatch is refused, not warned about — vectors out of step with the
chunks would attribute quotes to the wrong papers. Re-run refindex embed after
any refindex build.
Fusion is reciprocal-rank, not a weighted score sum: BM25 is unbounded and corpus-dependent while cosine is in [-1,1], so weighting needs calibration that drifts as the corpus grows.
Read the components, not the fused rank. RRF orders well but says nothing about whether a passage is support, so hits carry both:
| bm25 | cosine | |
|---|---|---|
genuine support (dziri2023faith#p9c5) | 16.6 | 0.82 |
unrelated paper (liu2024spinquant#p5c4) | 1.8 | 0.52 |
Metadata verification — CrossRef and DataCite (built 2026-08-16)
amem refverify answers the question refcheck cannot: does this citation
exist, and is the metadata right? (refcheck answers what it says.)
Which registry matters. arXiv registers preprints through DataCite
(10.48550/arXiv.*); CrossRef 404s them. Verified by hand: CrossRef returns 404
for 10.48550/arXiv.1706.03762 while DataCite returns “Attention Is All You
Need”, 2017, publisher arXiv. Routing everything to CrossRef made all 7 local
preprints report ABSENT — indistinguishable from fabricated references. With
prefix routing: 7/7, and 74/75 on formace.
An arXiv id is turned into an exact DOI, so most lookups are direct resolution rather than fuzzy search. Short titles additionally need the first author’s surname to corroborate: “Attention Is All You Need” is five tokens and matched a 2025 book chapter at 71% Jaccard, reported as a year error when it was the wrong paper entirely.
Result on the 78-ref corpus — every duplicate content hashing suspected, confirmed by DOI:
10.48550/arxiv.2201.11903 ← wei2022chain = wei2022cot
10.48550/arxiv.2305.18290 ← dpo_2023 = rafailov2023direct
10.48550/arxiv.2310.07923 ← merrill2023expressive = merrill2024expressive
Note on the last one: both records carry year 2023 and DataCite agrees, so only
the cite key string merrill2024expressive is wrong — the metadata never was.
Not done yet
- better-bibtex plugin — not installed; stable cite keys currently come
from amem’s own metadata, not from Zotero. The
.bibcurrently indexed is generated from~/.amem/raw/*.meta.json, not exported by Zotero. - Dedupe the three duplicate cite-key groups, now confirmed by DOI.
- Re-embedding is manual —
refindex buildinvalidates the vectors by fingerprint (correctly), but nothing re-runsrefindex embedfor you. - The clipper bridge is still loopback-only, deliberately — see the
cross-machine section. Only
refserveis exposed. - GUI curation — no VNC on the box. Adding/annotating items by hand needs either a VNC server or doing curation on the Mac and syncing.