Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 20 additions & 8 deletions BACKLOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,14 +66,6 @@ quantized int8 vectors + RRF beats their JSON BM25 index), on multi-machine repl
(they have none), on enforcement, and on shipping a real MCP server. What follows are the
five places their design is genuinely better and the idea transfers._

- [ ] **Contextual-prefix indexed chunks, synthetic tier only** ([#193](https://github.com/CryptoJones/omind/issues/193)) — _enhancement (retrieval)_ —
they implement Anthropic's Contextual Retrieval pattern: each chunk gets a 1–2
sentence prefix situating it in its page, and the *prefixed* text is what gets
indexed and embedded. omind indexes raw heading-split chunks, so a mid-note chunk
competes on its own words alone. Take only their zero-egress synthetic tier
(`<title> — <summary>. Section: <heading path>.`), not the LLM-written tiers —
omind's embeddings stay local. Gate it behind `omind bench --quality` and keep it
only if recall moves.
- [ ] **Journaled plan→apply→recover transactions for multi-note operations** ([#194](https://github.com/CryptoJones/omind/issues/194)) — _enhancement (durability)_ —
omind has atomic per-file replace, an advisory write lock, and version
preconditions, but no journal and no rollback, so an interrupted multi-note
Expand Down Expand Up @@ -129,6 +121,26 @@ _No open items._

## Not planned

- [ ] **Contextual-prefix indexed chunks (Anthropic Contextual Retrieval)** ([#193](https://github.com/CryptoJones/omind/issues/193), closed not-planned) — _rejected on measurement_ —
built behind `OMI_CONTEXTUAL_CHUNKS` and evaluated on the live 784-note vault, both
ways, with the semantic leg on. **recall@1 60% → 60%, recall@5 60% → 60%, MRR 0.640 →
0.640**, for +31% index size (13,908 → 18,248 KiB) and +20% rebuild time. Of the five
labelled cases, three were already rank 1 in both; the two misses went 7→8 and 13→11.
Noise, not signal.

The reason is the part worth remembering: **the issue's premise did not hold for omind.**
It assumed a mid-note chunk "competes on its own words alone", which is true of
claude-obsidian but not here — `_ingest` has always embedded `title + heading + tags +
chunk.text`, and `chunks_fts` has always given BM25 separate title/heading/tags columns.
omind already had contextual retrieval without the name. The prefix's only genuine
addition is the note's `Summary`, which on descriptive titles is largely a restatement
of a signal already indexed. The upstream 35–49% figure is measured against a baseline
that indexes bare chunks; omind is not that baseline.

Caveat kept honest: the labelled set is only 5 cases, so each is worth 20pp and an
effect under ~15% could hide. A 25–40 case set is worth building for retrieval work in
general — and would be the thing that could reopen this.

- [ ] **Long game: fine-tune a model on the accumulated violation corpus** ([#91](https://github.com/CryptoJones/omind/issues/91), closed not-planned) — _roadmap (Phase 4)_ — deferred: the blocker is data, not compute. The live `compliance.jsonl` corpus is ~91% relevance-noise, ~6% real denies, and 100% DENY (zero ALLOW), so training on it as-is yields an always-deny model. Revisit only after `export-corpus` is reworked to synthesize balanced ALLOW examples (from the deterministic `guard.decide()`) and split the relevance corpus from the action corpus. The mechanical guard remains the backstop.
- [ ] **claude-obsidian's source capture, Canvas/Bases emitters, and methodology filing modes** — _rejected_ —
evaluated during the 2026-08-02 comparison. Their `capture` (immutable content-addressed
Expand Down
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,14 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
_The last of the 2026-08-01 review and 2026-08-02 comparison backlogs, released
together rather than as another run of point versions._

### Changed
- The FTS5 `snippet()` call names its column through `_FTS_TEXT_COLUMN` instead
of a bare `3`. `snippet()` takes a column *number*, so inserting any column
before `text` silently re-points every excerpt at the wrong column rather than
erroring — which is exactly what happened while evaluating
[#193](https://github.com/CryptoJones/omind/issues/193) (declined; see
`BACKLOG.md`). The experiment was reverted; this guardrail is worth keeping.

### Fixed
- **A `SCHEMA_VERSION` bump that adds a column no longer wedges an existing
search index** ([#210](https://github.com/CryptoJones/omind/issues/210)).
Expand Down
10 changes: 9 additions & 1 deletion src/omind/searchindex.py
Original file line number Diff line number Diff line change
Expand Up @@ -81,6 +81,13 @@
#: just should not outrank a comparable one that was actually verified. Much
#: gentler than the superseded penalty — low confidence is not obsolescence.
_LOW_CONFIDENCE_WEIGHT = 0.8
#: 0-based position of `text` in ``chunks_fts`` — FTS5's snippet() takes a
#: column NUMBER, not a name, so inserting any column before `text` silently
#: re-points every excerpt at the wrong column instead of erroring. That is not
#: hypothetical: the #193 experiment added a column ahead of it and excerpts
#: started returning the note's title line. Keep in step with the CREATE VIRTUAL
#: TABLE below.
_FTS_TEXT_COLUMN = 3
#: Only the fused head pays the extra local embedding pass.
_RERANK_DEPTH = 20
#: How many candidates each leg contributes before fusion.
Expand Down Expand Up @@ -1274,7 +1281,8 @@ def _excerpt(
return _collapse(str(row["text"]), max_chars)
if expr:
snippet = db.execute(
"SELECT snippet(chunks_fts, 3, '', '', '…', 24) AS s FROM chunks_fts"
f"SELECT snippet(chunks_fts, {_FTS_TEXT_COLUMN}, '', '', '…', 24) AS s"
" FROM chunks_fts"
" WHERE rowid = ? AND chunks_fts MATCH ?",
(chunk_id, expr),
).fetchone()
Expand Down