Files
Evan 102b506b0b fix(agents,hands): per-agent/per-hand mcp_servers / skills allowlists (#87) (#92)
All 32 agent manifests and 17 hands shipped with empty mcp_servers /
skills lists, which the kernel interprets as "no filter" — every
globally-configured MCP server's tools and every installed skill get
injected into the prompt on every LLM call. On a typical instance (9
MCP servers, ~85 MCP tools + ~82 built-in tools) that's ~50k input
tokens per turn spent on definitions the agent never uses.

Changes
-------

32 agents/*/agent.toml:
  - mcp_servers: 1-4 per agent. memory wherever state persists across
    turns; fetch / exa-search / brave-search only where the prompt
    actually calls for web; git / github / filesystem on engineering
    agents; gmail / google-calendar / linear / jira on productivity
    agents whose prompts mention them.
  - skills: per-role allowlist driven by what the system_prompt names
    (e.g. coder → rust/python/typescript/git/shell-scripting; devops-
    lead → docker/kubernetes/terraform/ansible/ci-cd/helm/prometheus/
    sysadmin). Generalists (assistant) keep skills = [] (see "Open
    items" below).
  - skills_disabled = true on the four short-conversational agents
    (hello-world, recipe-assistant, health-tracker, home-automation).
    Their system prompts never instruct the LLM to consult any skill,
    so loading all 60 was pure waste. They also drop the explicit
    max_history_messages override and inherit the kernel default (60).
  - max_history_messages tiered by workload shape:
      60  short conversational (hello-world, recipe, health-tracker,
          home-automation) — inherits the rising kernel default
          (`DEFAULT_MAX_HISTORY_MESSAGES = 60`); no override needed.
      60  single-turn task agents (writer, translator, doc-writer,
          email-assistant, customer-support, sales-assistant, recruit-
          er, social-media, personal-finance, tutor, travel-planner,
          meeting-assistant, ops, devops-lead, planner) — explicit
          override at the same value to lock the cap if the kernel
          default moves again.
      80  multi-step / tool-heavy (coder, debugger, architect, code-
          reviewer, test-engineer, security-auditor, analyst, data-
          scientist, academic-researcher, researcher, legal-assistant)
      120 coordinators (assistant, orchestrator) — long multi-agent
          sessions where prompt-cache continuity is critical
    All values sit at or above the kernel default. Pinning lower
    would thrash the prompt cache (the failure mode #91 fixed for
    the creator hand by *raising* the cap, not lowering it).

17 hands/*/HAND.toml:
  - hand-level mcp_servers / skills now declared on every hand, so
    every [agents.*] inside inherits a sensible allowlist.
  - skills_disabled = true placed on each [agents.*] inside clip and
    creator (pure media pipelines that don't benefit from any skill).
    HandDefinitionRaw in librefang-hands does NOT have a top-level
    skills_disabled field — declaring it at the hand top level would
    be silently dropped by serde, so the setting must live on the
    AgentManifest of each sub-agent role.
  - devteam: expand existing mcp_servers = ["github"] to include
    memory / git / filesystem; populate skills with the expected
    dev-team expertise (replacing the placeholder skills = []).
  - wiki: replace placeholder mcp_servers = [] with [memory, fetch,
    filesystem]. Hand-level skills stays [].
  - lead: hand-level skills was originally [email-writer, writing-
    coach, interview-prep]; interview-prep is for job-interview
    preparation, not lead generation. Replaced with data-analyst
    (used by the qualification-scoring step in the prompt).

schema.toml: register mcp_servers / skills / max_history_messages on
the agent field schema so machine consumers (RegistrySchema in
librefang-types) see the new top-level fields. The
max_history_messages description now points at
librefang_runtime::agent_loop::DEFAULT_MAX_HISTORY_MESSAGES (60
today) by name, so the schema doesn't go stale when the constant
moves again.

agents/README.md: example block + "Adding a New Agent" checklist
mention the allowlists; max_history_messages example is shown
commented out with a prompt-cache caveat.

Open items
----------

`assistant` (the default user-facing agent) keeps `skills = []`
deliberately. It is the generalist entry point — capping its skill
surface at a small allowlist would defeat its "delegate to any
specialist" job. The trade-off is that this single agent still pays
the full skill-definition load on every turn; operators who want a
strict allowlist for `assistant` can override it after install.

Why not adopt PR #89's approach
-------------------------------

#89 covers similar ground but with three issues this PR avoids:

1. mcp_servers = ["_none"] sentinel. #89's body explicitly notes
   it's pending upstream librefang#4808 (mcp_disabled). Shipping a
   magic-string today means coming back later to clean it up. This
   PR uses real allowlists.
2. max_history_messages = 8 / 12 / 15 / 20. Far below today's
   kernel default (60) and #91's direction for long-workflow hands
   (80–120). Every turn that hits the cap invalidates the cached
   prompt prefix; the cost of cache misses exceeds the saving from
   shorter history. This PR uses 60–120.
3. Doubling max_llm_tokens_per_hour (coder 200k→500k, assistant
   300k→500k) widens the per-agent budget — the opposite direction
   from #87's "reduce per-call cost" goal. Left to the operator's
   instance-specific tuning.

Refs librefang/librefang-registry#87, librefang/librefang-registry#89
2026-05-12 09:30:21 +09:00

133 lines
5.0 KiB
TOML

name = "academic-researcher"
version = "0.4.3-beta4-20260314"
description = "Academic research agent. Searches scholarly papers, summarizes findings, and generates literature reviews."
author = "librefang"
module = "builtin:chat"
tags = ["research", "academic", "papers", "literature-review", "science"]
# Per-agent resource allowlists (refs librefang/librefang-registry#87).
# Empty list = all available; explicit list filters the prompt surface
# so the LLM only sees what this agent actually uses.
mcp_servers = ["memory", "fetch", "exa-search", "brave-search"]
skills = ["technical-writer", "writing-coach", "pdf-reader"]
max_history_messages = 80
[metadata.routing]
aliases = [
"academic research",
"paper search",
"scholarly research",
"find papers",
]
weak_aliases = ["papers", "citations", "bibliography", "meta-analysis"]
[model]
provider = "default"
model = "default"
api_key_env = "GEMINI_API_KEY"
max_tokens = 8192
temperature = 0.3
system_prompt = """You are Academic Researcher, a scholarly research agent running inside the LibreFang Agent OS. You specialize in searching academic papers, summarizing research findings, and generating structured literature reviews.
RESEARCH METHODOLOGY:
1. SCOPE — Clarify the research question. Identify key concepts, synonyms, and related terms. Define inclusion/exclusion criteria (date range, field, study type).
2. SEARCH — Use web_search with academic queries (site:arxiv.org, site:scholar.google.com, site:pubmed.ncbi.nlm.nih.gov, site:semanticscholar.org). Use multiple query phrasings combining key terms with Boolean logic.
3. RETRIEVE — Use web_fetch to read full paper abstracts, methods, and conclusions. Don't rely on search snippets alone.
4. EVALUATE — Assess each source for relevance, methodology rigor, sample size, peer-review status, citation count, and journal impact. Prefer peer-reviewed publications over preprints.
5. SYNTHESIZE — Organize findings thematically. Identify consensus, contradictions, and gaps in the literature. Compare methodologies and results across studies.
6. CITE — Maintain proper academic citations throughout. Use a consistent citation format (APA-style by default).
SOURCE HIERARCHY (strongest to weakest):
- Systematic reviews and meta-analyses
- Randomized controlled trials / large-scale empirical studies
- Cohort and case-control studies
- Cross-sectional studies and surveys
- Case reports and expert opinions
- Preprints (flag as not yet peer-reviewed)
OUTPUT FORMATS:
For Paper Search:
- Title, Authors, Year, Journal/Venue
- Abstract summary (2-3 sentences)
- Key findings and methodology
- Relevance to the research question (high/medium/low)
For Literature Review:
- Introduction (research question and scope)
- Methodology (search strategy, databases, inclusion criteria)
- Thematic Analysis (organized by theme, not chronologically)
- Discussion (consensus, contradictions, gaps)
- Conclusion (state of knowledge, future directions)
- References (full citation list)
For Paper Summary:
- Citation (authors, year, title, journal)
- Objective / Research Question
- Methodology (design, sample, measures)
- Key Findings (with effect sizes and confidence intervals when available)
- Limitations
- Implications
GUIDELINES:
- Always distinguish between correlation and causation.
- Report effect sizes, confidence intervals, and p-values when available.
- Flag potential biases (funding, sample selection, publication bias).
- Note the recency of findings — flag if the field has evolved since publication.
- When sources conflict, present both sides with evidence strength assessment.
- Never overstate conclusions beyond what the evidence supports.
- Acknowledge limitations and gaps in the available literature."""
[[fallback_models]]
provider = "default"
model = "default"
api_key_env = "GROQ_API_KEY"
[resources]
max_llm_tokens_per_hour = 200000
max_concurrent_tools = 5
[capabilities]
tools = [
"web_search",
"web_fetch",
"file_read",
"file_write",
"file_list",
"memory_store",
"memory_recall",
]
network = ["*"]
memory_read = ["*"]
memory_write = ["self.*", "shared.*"]
[i18n.zh]
name = "学术研究员"
description = "学术研究 Agent:检索学术论文、归纳结论、撰写文献综述。"
[i18n.zh-TW]
name = "學術研究員"
description = "學術研究 Agent:檢索學術論文、歸納結論、撰寫文獻綜述。"
[i18n.ja]
name = "学術リサーチャー"
description = "学術論文の検索、要点整理、文献レビュー生成を行う研究 Agent。"
[i18n.ko]
name = "학술 연구원"
description = "학술 논문 검색, 요약, 문헌 리뷰 작성을 수행하는 연구 Agent."
[i18n.de]
name = "Akademischer Forscher"
description = "Forschungs-Agent für wissenschaftliche Recherche, Ergebniszusammenfassung und Literaturübersichten."
[i18n.es]
name = "Investigador académico"
description = "Agente de investigación académica: busca artículos, resume hallazgos y genera revisiones de literatura."
[i18n.fr]
name = "Chercheur académique"
description = "Agent de recherche académique : recherche d'articles, synthèse et revues de littérature."