feat(hands): add wiki hand for LLM-maintained knowledge bases (#44)

Squashed replay of the original 9-commit branch onto current main.  The
original branch was 30+ commits behind, forked from before the skills
refactor (PR #42) and workflow template expansion (PR #36), so a
standard rebase hit heavy add/add conflicts on workflows/*.toml that
are unrelated to the wiki hand.

This replay keeps only the final hands/wiki/ tree state, which is the
actual intent of the PR (the author iterated several times on the same
files; squashing matches that).

Implements the "LLM Wiki" pattern (Andrej Karpathy) for building a
personal, Obsidian-compatible knowledge base.  Instead of on-the-fly
RAG, the wiki hand incrementally maintains a Markdown vault:

  hands/wiki/
  ├── HAND.toml           # hand manifest + [agents.*] sections
  ├── README.md           # user-facing docs
  ├── SKILL-main.md       # Librarian (coordinator) routing + FS ops
  ├── SKILL-ingestor.md   # Source extraction + [[wikilink]] writing
  ├── SKILL-analyst.md    # Synthesis with provenance citations
  └── SKILL-linter.md     # Broken link / orphan / contradiction audit

Closes librefang/librefang-registry#44 (via replay, not merge).
This commit is contained in:
Adrian Rogala authored and GitHub committed 2026-04-10 22:04:18 +08:00
1 parent 8d3b49192d
commit f9c7456900
6 files changed
+1734

No files matched your search

+1027
View File
File diff suppressed because it is too large. Load diff
+36
View File
@@ -0,0 +1,36 @@
# Wiki Hand
Autonomous personal knowledge base agent -- builds and maintains an Obsidian-compatible wiki from raw sources. It extracts entities, concepts, and claims with strict provenance tracking, cross-references everything, and performs periodic health checks.
## Configuration
| Field | Value |
|-------|-------|
| Category | `knowledge` |
| Agent | `librarian` |
| Routing | `wiki`, `knowledge base`, `ingest source`, `ingest deep`, `deep ingest`, `crawl`, `wiki query`, `wiki lint` |
## Integrations
- **Git** -- Required for version control. Every operation (ingest, query, lint, maintain) is automatically committed to your vault history.
- **qmd** -- (Optional) A local hybrid search engine. Recommended for wikis scaling beyond 200 pages.
## Settings
- **Wiki Vault Path** -- Directory where your wiki files will be stored (default: `./wiki`)
- **File-Back Mode** -- How to handle saving generated syntheses (`auto`, `ask`, `never`)
- **Search Backend** -- The search engine to use for queries (`index`, `qmd`)
- **Wiki Content Language** -- Target language for all generated content in the wiki (default: `en`)
## Features
- **Ingestion:** Give the agent raw sources (markdown, text, HTML, PDF via terminal) and it will extract factual claims, identify entities, and build connected concept pages.
- **Deep Ingestion (`ingest-deep`):** Provide a starting URL, and the agent will intelligently crawl the main page along with up to 5 of its most substantive sub-pages to build a comprehensive knowledge tree in one go.
- **Synthesized Queries:** Ask questions across your entire knowledge base. The agent provides answers backed by `[[wikilink]]` citations, effectively compiling its response from your collected literature.
- **Linting & Maintenance:** The agent can periodically scan the wiki to merge duplicate entities, fix broken links, identify contradictions, and point out orphan pages.
## Usage
```bash
librefang hand run wiki
```
+115
View File
@@ -0,0 +1,115 @@
---
name: wiki-analyst
version: "1.0.0"
description: Wiki Analyst skills — synthesis methodology, gap identification, contradiction handling, and query resolution.
author: Leszek3737
tags: [wiki, knowledge-base, analyst, synthesis, reasoning]
runtime: prompt_only
---
# Wiki Hand — SKILL-analyst.md
## 1. Obsidian Conventions
### Wikilinks
- Internal cross-references: `[[page-name]]` (no path, no extension)
- Display text override: `[[page-name|Display Text]]`
- Never use markdown `[text](path)` for internal links
- External URLs: standard markdown `[text](https://...)`
- Image embeds: `![[filename.png]]`
- Images live in `raw/assets/`
### File Naming
- Format: `kebab-case.md`
- Max 40 characters (excluding `.md`)
- ASCII only — transliterate non-ASCII characters (ü → ue, ñ → n, ś → s, etc.)
- Entities: canonical name (`john-doe.md`, not `dr-john-doe-phd.md`)
- Concepts: noun phrase (`knowledge-base.md`, not `about-knowledge-bases.md`)
- Sources: derived from title (`the-future-of-ai.md`)
- Syntheses: derived from query (`tradeoffs-x-vs-y.md`)
- No version suffixes (_v2, _revised, _updated)
### Frontmatter
- Valid YAML between `---` delimiters at the top of every page
- All fields lowercase with underscores
- Dates: ISO 8601 (`YYYY-MM-DD`)
- Lists: YAML sequences (not comma-separated strings)
- Dataview-compatible: all fields are queryable
### Dataview Compatibility
Useful queries the user can run in Obsidian:
```dataview
TABLE confidence, source_count, last_updated
FROM "pages/entities"
SORT source_count DESC
```
```dataview
LIST
FROM "pages/concepts"
WHERE confidence = "disputed"
```
```dataview
TABLE claim_count, date_ingested
FROM "pages/sources"
SORT date_ingested DESC
```
---
## 4. Provenance and Confidence
### Inline Provenance Syntax
Every key claim (bullet point in Key Claims, Key Points, Key Facts sections) requires a provenance tag at the end of the line:
```markdown
- The company reported $10M ARR in Q4 2025 ([[annual-report-2025]], extracted)
- This suggests a 40% year-over-year growth rate (inferred)
- However, a later filing revised this to $8.5M ([[sec-filing-q1-2026]], extracted)
- The actual growth rate remains disputed ([[annual-report-2025]], [[sec-filing-q1-2026]], disputed)
```
| Tag | Meaning | When to use |
|-----|---------|-------------|
| `([[source]], extracted)` | Directly stated in the source | Verbatim facts, statistics, dates, names |
| `([[source-a]], [[source-b]], extracted)` | Corroborated across sources | Same fact confirmed independently |
| `(inferred)` | Derived by the LLM | Connections, implications, patterns not explicitly stated |
| `([[source-a]], [[source-b]], disputed)` | Sources contradict | Present both claims, let the user judge |
### Page-Level Confidence
Set in frontmatter `confidence` field:
| Level | Rule |
|-------|------|
| `high` | All key claims are `extracted` from 2+ corroborating sources |
| `medium` | Mix of extracted and inferred, OR all extracted from a single source |
| `low` | Primarily inferred, OR based on one unverified source |
| `disputed` | Contains at least one claim tagged `disputed` |
**Propagation in syntheses:** A synthesis page's confidence = the MINIMUM confidence of its consulted pages. If any consulted page is `disputed`, the synthesis must flag this.
### Confidence Update Triggers
- New source corroborates an existing claim → consider upgrading to `high`
- New source contradicts an existing claim → downgrade to `disputed`
- Source retracted or superseded → re-assess claims dependent on it
---
## 7. Synthesis Methodology
### Combining Multiple Sources
1. **Identify overlapping claims** — same fact from different sources
2. **Note corroboration** — claims supported by 2+ sources get higher confidence
3. **Surface contradictions** — present both sides with citations, do not silently pick one
4. **Fill complementary gaps** — where sources complement each other, weave together
5. **Maintain provenance chain** — every claim in the synthesis traces back to source pages
### Handling Contradictions
```markdown
## 2. Page Templates
## 12. Worked Examples
+260
View File
@@ -0,0 +1,260 @@
---
name: wiki-ingestor
version: "1.0.0"
description: Wiki Ingestor skills — extraction heuristics, source/entity/concept templates, provenance, and cross-referencing.
author: Leszek3737
tags: [wiki, knowledge-base, ingestor, extraction, provenance]
runtime: prompt_only
---
# Wiki Hand — SKILL-ingestor.md
## 1. Obsidian Conventions
### Wikilinks
- Internal cross-references: `[[page-name]]` (no path, no extension)
- Display text override: `[[page-name|Display Text]]`
- Never use markdown `[text](path)` for internal links
- External URLs: standard markdown `[text](https://...)`
- Image embeds: `![[filename.png]]`
- Images live in `raw/assets/`
### File Naming
- Format: `kebab-case.md`
- Max 40 characters (excluding `.md`)
- ASCII only — transliterate non-ASCII characters (ü → ue, ñ → n, ś → s, etc.)
- Entities: canonical name (`john-doe.md`, not `dr-john-doe-phd.md`)
- Concepts: noun phrase (`knowledge-base.md`, not `about-knowledge-bases.md`)
- Sources: derived from title (`the-future-of-ai.md`)
- Syntheses: derived from query (`tradeoffs-x-vs-y.md`)
- No version suffixes (_v2, _revised, _updated)
### Frontmatter
- Valid YAML between `---` delimiters at the top of every page
- All fields lowercase with underscores
- Dates: ISO 8601 (`YYYY-MM-DD`)
- Lists: YAML sequences (not comma-separated strings)
- Dataview-compatible: all fields are queryable
### Dataview Compatibility
Useful queries the user can run in Obsidian:
```dataview
TABLE confidence, source_count, last_updated
FROM "pages/entities"
SORT source_count DESC
```
```dataview
LIST
FROM "pages/concepts"
WHERE confidence = "disputed"
```
```dataview
TABLE claim_count, date_ingested
FROM "pages/sources"
SORT date_ingested DESC
```
---
## 4. Provenance and Confidence
### Inline Provenance Syntax
Every key claim (bullet point in Key Claims, Key Points, Key Facts sections) requires a provenance tag at the end of the line:
```markdown
- The company reported $10M ARR in Q4 2025 ([[annual-report-2025]], extracted)
- This suggests a 40% year-over-year growth rate (inferred)
- However, a later filing revised this to $8.5M ([[sec-filing-q1-2026]], extracted)
- The actual growth rate remains disputed ([[annual-report-2025]], [[sec-filing-q1-2026]], disputed)
```
| Tag | Meaning | When to use |
|-----|---------|-------------|
| `([[source]], extracted)` | Directly stated in the source | Verbatim facts, statistics, dates, names |
| `([[source-a]], [[source-b]], extracted)` | Corroborated across sources | Same fact confirmed independently |
| `(inferred)` | Derived by the LLM | Connections, implications, patterns not explicitly stated |
| `([[source-a]], [[source-b]], disputed)` | Sources contradict | Present both claims, let the user judge |
### Page-Level Confidence
Set in frontmatter `confidence` field:
| Level | Rule |
|-------|------|
| `high` | All key claims are `extracted` from 2+ corroborating sources |
| `medium` | Mix of extracted and inferred, OR all extracted from a single source |
| `low` | Primarily inferred, OR based on one unverified source |
| `disputed` | Contains at least one claim tagged `disputed` |
**Propagation in syntheses:** A synthesis page's confidence = the MINIMUM confidence of its consulted pages. If any consulted page is `disputed`, the synthesis must flag this.
### Confidence Update Triggers
- New source corroborates an existing claim → consider upgrading to `high`
- New source contradicts an existing claim → downgrade to `disputed`
- Source retracted or superseded → re-assess claims dependent on it
---
## 5. Cross-Referencing Patterns
### When to Create a Dedicated Page
| Condition | Action |
|-----------|--------|
| Entity/concept is the main subject of a source | Create page |
| Entity/concept appears substantively in 3+ sources | Create page |
| Entity is the author and contributes beyond a byline | Create page |
| Passing mention in < 3 sources, not the main subject | Plain text only, NO page, NO wikilink |
| Librarian or user explicitly requests a page | Create regardless of threshold |
"Substantively" means discussed in at least one paragraph, not just name-dropped.
### Wikilink Rules
- ONLY link to pages that exist. Never create a [[wikilink]] to a non-existent page.
- When creating a new page during ingest, you may wikilink to it from other pages you're touching in the same operation.
- First mention per section: use wikilink `[[page-name]]`
- Subsequent mentions in the same section: plain text is fine
- On creating a new page: check if other existing pages mention this entity/concept in plain text and could now be wikilinked. Flag this in the manifest for Librarian to handle.
### Backlink Maintenance
- On page creation: add the new page to relevant existing pages' "See Also" sections
- On page deletion: replace all inbound [[wikilinks]] with plain text
- On merge: redirect all inbound [[wikilinks]] from deleted page to surviving page
---
## 6. Extraction Heuristics
### Entity Recognition
| Entity Kind | Signals | Examples |
|-------------|---------|----------|
| `person` | Proper name, pronouns, job titles, biographical context | "John Doe, CEO of Acme" |
| `organization` | Company names, institutions, teams, brands | "Google", "MIT", "the W3C" |
| `tool` | Software, libraries, frameworks, products, protocols | "PostgreSQL", "React", "gRPC" |
| `place` | Geographic names, facilities, regions | "Silicon Valley", "CERN", "the EU" |
**Disambiguation:** When the same name could refer to different entities (e.g., "Mercury" — planet, element, car brand), use context. Create separate pages with qualifiers: `mercury-planet.md`, `mercury-element.md`. Add all variants to `aliases` in frontmatter.
### Concept Identification
Concepts are abstract ideas, patterns, theories, techniques, or methodologies — they don't refer to a specific named thing.
| Signal | Example |
|--------|---------|
| Defined or explained in the source | "Microservices architecture is a design approach where..." |
| Compared or contrasted with alternatives | "Unlike monolithic systems, microservices..." |
| Listed as a technique, methodology, or pattern | "Key patterns include: CQRS, event sourcing, and saga" |
| Forms the basis of an argument | "The efficient market hypothesis suggests..." |
### Claim Extraction
A claim is a factual assertion that can be verified, challenged, or updated. Aim for 5-15 per source.
**IS a claim:**
- "Revenue grew 40% in 2024" — verifiable quantitative data
- "The system uses a Rust backend" — architectural fact
- "The study found no significant correlation" — research finding
- "Founded in 2019" — dated fact
- "The team grew from 12 to 85 engineers" — organizational data
**NOT a claim:**
- "This is an interesting approach" — opinion without substance
- "The report continues with..." — structural description
- "See section 3 for details" — internal reference
- "It's important to consider..." — filler
### Relationship Mapping
When extracting entities, note relationships for the Relationships/Connections sections:
| Type | Example | Notation |
|------|---------|----------|
| Organizational | "Jane is CTO of Acme" | [[jane-doe]] ↔ [[acme-corp]], "CTO of" |
| Collaborative | "Developed by MIT and Google" | [[mit]] ↔ [[google]], "co-developed" |
| Competitive | "Competing with PostgreSQL" | [[product]] ↔ [[postgresql]], "competitor" |
| Dependency | "Built on React" | [[product]] ↔ [[react]], "depends on" |
| Temporal | "Founded 2019, acquired 2023" | Dates in frontmatter + key facts |
| Causal | "Migration caused 2.3x cost increase" | Noted as claim with provenance |
### Source Format Handling
| Format | Preprocessing | Notes |
|--------|--------------|-------|
| `.md` | None needed | Read directly |
| `.txt` | None needed | Read directly |
| `.html` | Strip tags or use web_fetch | Preserve headings structure if possible |
| `.pdf` | Extract text via pdftotext | May lose formatting; note if tables are garbled |
| Images in source | Note `![[image.png]]` references | If LLM vision available, describe key images |
When a source contains images that convey information (charts, diagrams, screenshots): if vision capability is available, describe the image content in the Assessment section. If not, note "Source contains visual content not processed: {description of what images appear to show}."
---
## 2. Page Templates
### 2.1 Source Summary
```markdown
---
type: source
title: "{Original Title}"
author: "{Author Name}"
date_published: YYYY-MM-DD
date_ingested: YYYY-MM-DD
raw_path: "raw/{filename.ext}"
source_url: "{url or null}"
format: md
claim_count: 0
confidence: medium
tags:
- {tag}
---
# {Title}
{2-3 sentence summary in the configured language.}
## 12. Worked Examples
### Example A: Ingesting a Technical Article
**Source:** `raw/microservices-at-scale-2025.md` — a blog post by Jane Chen about Acme Corp's migration from monolith to microservices.
**Step 1 — Analyze:** Ingestor reads the source and identifies:
- Entities: Jane Chen (person), Acme Corp (organization), Kubernetes (tool), AWS (organization)
- Concepts: microservices architecture, service mesh, circuit breaker pattern
- Key claims: 8 factual assertions
**Step 2 — Classify thresholds:**
- Jane Chen → CREATE (author, contributes substantively)
- Acme Corp → CREATE (main subject)
- Kubernetes → MENTION-ONLY (passing mention, 1 source, below threshold)
- AWS → MENTION-ONLY (passing mention, 1 source)
- microservices architecture → CREATE (main topic)
- service mesh → CREATE (discussed substantively)
- circuit breaker pattern → CREATE (discussed with specific data)
**Step 3 — Source summary → `pages/sources/microservices-at-scale-2025.md`:**
```markdown
---
type: source
title: "Microservices at Scale: Lessons from Acme Corp"
author: "Jane Chen"
date_published: 2025-09-15
date_ingested: 2026-04-09
raw_path: "raw/microservices-at-scale-2025.md"
source_url: "https://blog.acme.com/microservices-at-scale"
format: md
claim_count: 8
confidence: medium
tags:
- architecture
- cloud
---
# Microservices at Scale: Lessons from Acme Corp
Jane Chen describes Acme Corp's two-year migration from a monolithic Rails application to a microservices architecture on Kubernetes. The post covers technical decisions, organizational challenges, and performance outcomes.
+95
View File
@@ -0,0 +1,95 @@
---
name: wiki-linter
version: "1.0.0"
description: Wiki Linter skills — structural checks, provenance validation, contradiction detection, and vault health.
author: Leszek3737
tags: [wiki, knowledge-base, linter, validation, quality]
runtime: prompt_only
---
# Wiki Hand — SKILL-linter.md
## 1. Obsidian Conventions
### Wikilinks
- Internal cross-references: `[[page-name]]` (no path, no extension)
- Display text override: `[[page-name|Display Text]]`
- Never use markdown `[text](path)` for internal links
- External URLs: standard markdown `[text](https://...)`
- Image embeds: `![[filename.png]]`
- Images live in `raw/assets/`
### File Naming
- Format: `kebab-case.md`
- Max 40 characters (excluding `.md`)
- ASCII only — transliterate non-ASCII characters (ü → ue, ñ → n, ś → s, etc.)
- Entities: canonical name (`john-doe.md`, not `dr-john-doe-phd.md`)
- Concepts: noun phrase (`knowledge-base.md`, not `about-knowledge-bases.md`)
- Sources: derived from title (`the-future-of-ai.md`)
- Syntheses: derived from query (`tradeoffs-x-vs-y.md`)
- No version suffixes (_v2, _revised, _updated)
### Frontmatter
- Valid YAML between `---` delimiters at the top of every page
- All fields lowercase with underscores
- Dates: ISO 8601 (`YYYY-MM-DD`)
- Lists: YAML sequences (not comma-separated strings)
- Dataview-compatible: all fields are queryable
### Dataview Compatibility
Useful queries the user can run in Obsidian:
```dataview
TABLE confidence, source_count, last_updated
FROM "pages/entities"
SORT source_count DESC
```
```dataview
LIST
FROM "pages/concepts"
WHERE confidence = "disputed"
```
```dataview
TABLE claim_count, date_ingested
FROM "pages/sources"
SORT date_ingested DESC
```
---
## 8. Lint Checks Reference
### Check Definitions
| Check | Severity | Method |
|-------|----------|--------|
| Missing frontmatter | critical | Parse YAML — absent or malformed |
| Missing required field | critical | Compare against schema.md |
| Dead wikilink | critical | `find pages/ -name "{name}.md"` for each [[link]] |
| No provenance tag | critical | Scan Key Claims/Points/Facts bullets for missing tag |
| Index/file mismatch | critical | Compare index.md entries with actual files |
| Orphan page | warning | `grep -rl` across ALL pages/ for inbound links |
| Contradiction | warning | Compare extracted claims across pages sharing tags |
| Stale content | warning | `last_updated` vs newest source's `date_ingested` |
| Weak provenance | warning | `(inferred)` with source_count=1, or `high` confidence with source_count<2 |
| Duplicate filenames | warning | Fuzzy match on filenames (e.g., acme-corp + acme-corporation) |
| File not in index | warning | File exists but no index.md entry |
| Missing page | info | Plain-text entity/concept in 3+ source summaries |
| Large page | info | Word count > 3000 |
### Auto-Fixable vs Requires Confirmation
**Auto-fixable** (Librarian applies directly):
- Missing `last_updated` → set to today
- File missing from index.md → add entry with description from page's first paragraph
- Index entry without corresponding file → remove entry
**Requires user confirmation:**
- Contradictions (user decides which claim is correct)
- Merge recommendations (user reviews combined content)
- Delete recommendations (user accepts information loss)
- Confidence changes (user validates reasoning)
- New page creation (user approves topic importance)
---
## 12. Worked Examples
+201
View File
@@ -0,0 +1,201 @@
---
name: wiki-librarian
version: "1.0.0"
description: Wiki Librarian skills — schema management, indexing, linting resolution, and overall vault health.
author: Leszek3737
tags: [wiki, knowledge-base, librarian, schema, maintenance]
runtime: prompt_only
---
# Wiki Hand — SKILL-main.md
## 1. Obsidian Conventions
### Wikilinks
- Internal cross-references: `[[page-name]]` (no path, no extension)
- Display text override: `[[page-name|Display Text]]`
- Never use markdown `[text](path)` for internal links
- External URLs: standard markdown `[text](https://...)`
- Image embeds: `![[filename.png]]`
- Images live in `raw/assets/`
### File Naming
- Format: `kebab-case.md`
- Max 40 characters (excluding `.md`)
- ASCII only — transliterate non-ASCII characters (ü → ue, ñ → n, ś → s, etc.)
- Entities: canonical name (`john-doe.md`, not `dr-john-doe-phd.md`)
- Concepts: noun phrase (`knowledge-base.md`, not `about-knowledge-bases.md`)
- Sources: derived from title (`the-future-of-ai.md`)
- Syntheses: derived from query (`tradeoffs-x-vs-y.md`)
- No version suffixes (_v2, _revised, _updated)
### Frontmatter
- Valid YAML between `---` delimiters at the top of every page
- All fields lowercase with underscores
- Dates: ISO 8601 (`YYYY-MM-DD`)
- Lists: YAML sequences (not comma-separated strings)
- Dataview-compatible: all fields are queryable
### Dataview Compatibility
Useful queries the user can run in Obsidian:
```dataview
TABLE confidence, source_count, last_updated
FROM "pages/entities"
SORT source_count DESC
```
```dataview
LIST
FROM "pages/concepts"
WHERE confidence = "disputed"
```
```dataview
TABLE claim_count, date_ingested
FROM "pages/sources"
SORT date_ingested DESC
```
---
## 3. Index and Log Formats
### index.md Structure
```markdown
# Wiki Index
## 5. Cross-Referencing Patterns
### When to Create a Dedicated Page
| Condition | Action |
|-----------|--------|
| Entity/concept is the main subject of a source | Create page |
| Entity/concept appears substantively in 3+ sources | Create page |
| Entity is the author and contributes beyond a byline | Create page |
| Passing mention in < 3 sources, not the main subject | Plain text only, NO page, NO wikilink |
| Librarian or user explicitly requests a page | Create regardless of threshold |
"Substantively" means discussed in at least one paragraph, not just name-dropped.
### Wikilink Rules
- ONLY link to pages that exist. Never create a [[wikilink]] to a non-existent page.
- When creating a new page during ingest, you may wikilink to it from other pages you're touching in the same operation.
- First mention per section: use wikilink `[[page-name]]`
- Subsequent mentions in the same section: plain text is fine
- On creating a new page: check if other existing pages mention this entity/concept in plain text and could now be wikilinked. Flag this in the manifest for Librarian to handle.
### Backlink Maintenance
- On page creation: add the new page to relevant existing pages' "See Also" sections
- On page deletion: replace all inbound [[wikilinks]] with plain text
- On merge: redirect all inbound [[wikilinks]] from deleted page to surviving page
---
## 9. Schema Evolution Patterns
### How to Propose Changes
When Librarian identifies a recurring pattern:
1. State the observation: "I've noticed several entities are research papers. Currently we classify these as tools."
2. Propose a minimal change: "Add `entity_kind: paper` to the entity page schema."
3. Show the before/after diff
4. Wait for user confirmation
5. Optionally suggest a lint pass to retroactively update existing pages
### Backward Compatibility
- New optional fields: add with a default value. Existing pages remain valid.
- New required fields: add as optional first, run lint to find pages missing the field, then promote to required.
- New page types: add template to schema.md, create directory if needed.
- Changed field names: rename in all pages, update schema, lint to verify.
### Common Schema Additions
| Change | When | Example |
|--------|------|---------|
| New `entity_kind` | 3+ entities don't fit existing kinds | `paper`, `event`, `dataset` |
| New tag namespace | Domain-specific taxonomy emerging | `ai/`, `bio/`, `finance/` |
| New frontmatter field | Recurring metadata across pages | `relevance_score`, `review_status` |
| New page type | Distinct content pattern | `timeline`, `glossary`, `comparison` |
---
## 10. Configuration & Settings Reference
The Wiki Hand behavior is governed by settings in `HAND.toml`. The agents should adjust their operations based on these configurations:
### `vault_path`
- **What it is:** The root directory for the Obsidian-compatible wiki.
- **Agent Behavior:** All `shell_exec`, `file_read`, `file_write`, and `file_list` operations must be relative to or prefixed with this path.
- **Best Practice:** Never assume the vault is in the current working directory. Always use the parameterized `{vault_path}`.
### `file_back_mode`
- **What it is:** Determines whether new syntheses generated during user queries are saved back to the wiki.
- **Agent Behavior:**
- `auto`: Automatically save to `pages/syntheses/`, update index, log, and commit.
- `ask`: Ask the user for permission before saving.
- `never`: Only provide the answer in chat. Do not persist the synthesis page.
### `search_backend`
- **What it is:** Specifies the mechanism for querying the wiki.
- **Agent Behavior:**
- `index`: Read `index.md` and visually scan entries. Best for smaller wikis (<200 pages).
- `qmd`: Use the `qmd search` CLI command. Required for larger wikis to avoid context limits.
- **Best Practice:** Proactively suggest switching to `qmd` when the wiki grows beyond 150-200 pages.
### `language`
- **What it is:** The configured language for the wiki content.
- **Agent Behavior:** All generated content, summaries, page titles, and synthesis responses must match this language, even if the user queries in another language or provides foreign-language source materials.
---
## 11. Common Pitfalls & Best Practices
### Tag Proliferation
- **Pitfall:** Creating dozens of highly specific, single-use tags (e.g., `#startup-founded-in-2023`).
- **Best Practice:** Stick to broad, thematic tags (e.g., `#startup`, `#technology`). If more specificity is needed, use `entity_kind` or create a synthesis page grouping them.
### Premature Page Creation (Stubs)
- **Pitfall:** Creating a dedicated entity page for something mentioned only once in passing.
- **Best Practice:** Wait until an entity has 2+ sources or significant context before promoting it to a dedicated page. Otherwise, leave it as a plain-text mention or a generic wikilink in the source summary.
### Broken YAML Frontmatter
- **Pitfall:** Generating invalid YAML (e.g., unescaped quotes in titles, incorrect list formatting for aliases).
- **Best Practice:** Always validate frontmatter syntax mentally before writing. Use arrays correctly: `aliases: ["Name 1", "Name 2"]` or list format. Quote string values if they contain colons.
### Losing Provenance
- **Pitfall:** Extracting a bold claim into an entity page without adding the inline `[[Source]]` tag.
- **Best Practice:** Every key claim or fact MUST have an inline source reference. Information without provenance decays the reliability of the entire wiki.
### Unnecessary Duplication
- **Pitfall:** Creating `Acme Corp.md` when `Acme Corporation.md` already exists, resulting in fragmented knowledge.
- **Best Practice:** Always use cross-category search or `grep` before creating new entity pages. Merge duplicate entries during routine linting.
---
## 2. Page Templates
### 2.1 Source Summary
```markdown
---
type: source
title: "{Original Title}"
author: "{Author Name}"
date_published: YYYY-MM-DD
date_ingested: YYYY-MM-DD
raw_path: "raw/{filename.ext}"
source_url: "{url or null}"
format: md
claim_count: 0
confidence: medium
tags:
- {tag}
---
# {Title}
{2-3 sentence summary in the configured language.}
## 12. Worked Examples