Files
librefang-registry/hands/collector/HAND.toml
T
Evan 7881d327a5 refactor: migrate icon fields from emoji to lucide:<name> tokens (#63)
* refactor: migrate icon fields from emoji to lucide:<name> tokens

Every TOML manifest's `icon = "<emoji>"` line is replaced with
`icon = "lucide:<kebab-name>"` — a reference to a lucide-react icon,
which the librefang.ai site and dashboard render as crisp SVG. Reasons
for the switch:

- Emoji render very differently across OS/browser/font stacks; the
  registry catalog looked inconsistent from one row to the next.
- Five manifests (clip / creator / linkedin / reddit / twitter) had
  their icons stored as literal Python-style escape strings
  ("\\U0001F3AC") because the TOML parser upstream never decoded
  them. Switching away from emoji drops that class of bug entirely.
- As a drive-by, also decode the \\uXXXX accent escapes in the
  [i18n.fr] block of hands/creator/HAND.toml so "Créateur" shows
  up correctly.

87 files touched. example manifests left untouched (still "TODO").

* fix: backfill i18n name + drop the single-member email category

- Every existing [i18n.<lang>] block now has a `name` field. 60 files
  previously translated description but kept the English name
  implicitly — which rendered as "some English some Chinese" in the
  registry UI. Fill in the missing name from the English brand (or a
  known localized equivalent: DingTalk→钉钉, Feishu→飞书, Email→
  电子邮件 / メール / E-Mail / Correo / Courriel, and a handful of
  hands that have Chinese product names like 视频剪辑 Hand).
- channels/email.toml was the only item under category="email";
  reclassify it as "messaging" so the sub-category filter chip list
  on the category page isn't littered with singletons.

* feat(i18n): localize 76 agents/integrations/plugins into 7 languages

Adds full [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks with name + description to every manifest
that previously shipped English-only.

Coverage:
- 32 agents (academic-researcher, analyst, architect, assistant,
  code-reviewer, coder, customer-support, data-scientist, debugger,
  devops-lead, doc-writer, email-assistant, health-tracker,
  hello-world, home-automation, legal-assistant, meeting-assistant,
  ops, orchestrator, personal-finance, planner, recipe-assistant,
  recruiter, researcher, sales-assistant, security-auditor,
  social-media, test-engineer, translator, travel-planner, tutor,
  writer)
- 33 integrations (AWS, Azure, Bitbucket, Brave Search, Discord,
  Dropbox, Elasticsearch, Exa Search, Fetch, Filesystem, GCP, Git,
  GitHub, GitLab, Gmail, Google Calendar, Google Drive, Google Maps,
  Jira, Linear, Memory, MongoDB, Notion, PostgreSQL, Puppeteer, Redis,
  Sentry, Sequential Thinking, Slack, SQLite, Teams, Time, Todoist) —
  brand names kept as-is across all locales, only descriptions
  translated.
- 11 plugins (auto-summarizer, context-decay, conversation-logger,
  episodic-memory, guardrails, keyword-memory, mempalace-indexer,
  sentiment-tracker, todo-tracker, topic-memory, user-profile)

The descriptions are one-line summaries — hand-translated rather than
machine-generated, so technical terms (MCP, PR, CI/CD, etc.) stay
consistent across locales.

* feat(i18n): close remaining per-lang gaps for channels, workflows, devteam

Third pass on i18n coverage. Every non-example manifest now carries a
full set of [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks.

- 44 channel adapters: added French descriptions (zh/zh-TW/ja/ko/de/es
  were already present). Brand names kept as-is in all locales so users
  recognize Discord / Slack / LINE / etc. consistently.
- 22 workflows: filled zh-TW / ja / ko / de / es / fr blocks. Each
  translation mirrors the existing zh one in structure and tone so the
  catalog reads consistently across locales.
- hands/devteam/HAND.toml: added the four langs that were missing
  (zh-TW, de, es, fr).

Only the 6 templates under examples/ are left without i18n blocks on
purpose — they still contain "TODO:" placeholders.
2026-04-17 22:04:26 +09:00

974 lines
42 KiB
TOML

id = "collector"
version = "1.1.0"
name = "Collector Hand"
description = "Autonomous intelligence collector — monitors any target continuously with change detection and knowledge graphs"
category = "data"
tags = ["popular"]
icon = "lucide:search"
tools = [
"shell_exec",
"file_read",
"file_write",
"file_list",
"web_fetch",
"web_search",
"memory_store",
"memory_recall",
"schedule_create",
"schedule_list",
"schedule_delete",
"knowledge_add_entity",
"knowledge_add_relation",
"knowledge_query",
"event_publish",
]
[routing]
aliases = [
"monitor changes",
"track updates",
"collect intelligence",
"osint",
"change detection",
"gather info",
"market intelligence",
]
weak_aliases = [
"watch",
"signals",
"continuous monitoring",
"surveillance",
"intel",
"news monitoring",
]
# ─── Configurable settings ───────────────────────────────────────────────────
[[settings]]
key = "target_subject"
label = "Target Subject"
description = "What to monitor (company name, person, technology, market, topic)"
setting_type = "text"
default = ""
[[settings]]
key = "collection_depth"
label = "Collection Depth"
description = "How deep to dig on each cycle"
setting_type = "select"
default = "deep"
[[settings.options]]
value = "surface"
label = "Surface (headlines only)"
[[settings.options]]
value = "deep"
label = "Deep (full articles + sources)"
[[settings.options]]
value = "exhaustive"
label = "Exhaustive (multi-hop research)"
[[settings]]
key = "update_frequency"
label = "Update Frequency"
description = "How often to run collection sweeps"
setting_type = "select"
default = "daily"
[[settings.options]]
value = "hourly"
label = "Every hour"
[[settings.options]]
value = "every_6h"
label = "Every 6 hours"
[[settings.options]]
value = "daily"
label = "Daily"
[[settings.options]]
value = "weekly"
label = "Weekly"
[[settings]]
key = "focus_area"
label = "Focus Area"
description = "Lens through which to analyze collected intelligence"
setting_type = "select"
default = "general"
[[settings.options]]
value = "market"
label = "Market Intelligence"
[[settings.options]]
value = "business"
label = "Business Intelligence"
[[settings.options]]
value = "competitor"
label = "Competitor Analysis"
[[settings.options]]
value = "person"
label = "Person Tracking"
[[settings.options]]
value = "technology"
label = "Technology Monitoring"
[[settings.options]]
value = "general"
label = "General Intelligence"
[[settings]]
key = "alert_on_changes"
label = "Alert on Changes"
description = "Publish an event when significant changes are detected"
setting_type = "toggle"
default = "true"
[[settings]]
key = "report_format"
label = "Report Format"
description = "Output format for intelligence reports"
setting_type = "select"
default = "markdown"
[[settings.options]]
value = "markdown"
label = "Markdown"
[[settings.options]]
value = "json"
label = "JSON"
[[settings.options]]
value = "html"
label = "HTML"
[[settings]]
key = "max_sources_per_cycle"
label = "Max Sources Per Cycle"
description = "Maximum number of sources to process per collection sweep"
setting_type = "select"
default = "30"
[[settings.options]]
value = "10"
label = "10 sources"
[[settings.options]]
value = "30"
label = "30 sources"
[[settings.options]]
value = "50"
label = "50 sources"
[[settings.options]]
value = "100"
label = "100 sources"
[[settings]]
key = "track_sentiment"
label = "Track Sentiment"
description = "Analyze and track sentiment trends over time"
setting_type = "toggle"
default = "false"
[[settings]]
key = "source_reliability_threshold"
label = "Source Reliability Threshold"
description = "Minimum source tier required to include a data point (lower tiers are discarded unless they are the sole source for a structural change)"
setting_type = "select"
default = "tier_3"
[[settings.options]]
value = "tier_1"
label = "Tier 1 only (official/primary sources)"
[[settings.options]]
value = "tier_2"
label = "Tier 2+ (institutional and above)"
[[settings.options]]
value = "tier_3"
label = "Tier 3+ (professional and above)"
[[settings.options]]
value = "tier_4"
label = "Tier 4+ (community and above)"
[[settings.options]]
value = "tier_5"
label = "All sources (no filtering)"
[[settings]]
key = "change_significance_threshold"
label = "Change Significance Threshold"
description = "Minimum significance score (0-100) for a change to be classified as IMPORTANT. Changes below this threshold are classified as MINOR."
setting_type = "select"
default = "60"
[[settings.options]]
value = "40"
label = "40 (more sensitive — more alerts)"
[[settings.options]]
value = "50"
label = "50 (balanced)"
[[settings.options]]
value = "60"
label = "60 (default)"
[[settings.options]]
value = "70"
label = "70 (stricter — fewer alerts)"
[[settings.options]]
value = "80"
label = "80 (very strict — only critical-level)"
# ─── Agent configuration ─────────────────────────────────────────────────────
[agents.main]
coordinator = true
name = "collector-hand"
description = "AI intelligence collector — monitors any target continuously with OSINT techniques, knowledge graphs, and change detection"
module = "builtin:chat"
provider = "default"
model = "default"
max_tokens = 16384
temperature = 0.3
max_iterations = 60
system_prompt = """You are Collector Hand — an autonomous intelligence collector that monitors any target 24/7, building a living knowledge graph and detecting changes over time.
## IMPORTANT: Shell Execution Rules
- Execute ONE command per shell_exec call. NEVER chain commands with `;`, `&&`, `||`, or pipes `|`.
- NEVER use backticks, `$()`, `${}`, or I/O redirection (`>`, `<`, `>>`).
- If you need multiple commands, make separate shell_exec calls for each.
- If a file_read fails, check the path exists first with shell_exec before retrying.
## Phase 0 — Platform Detection & State Recovery (ALWAYS DO THIS FIRST)
Detect the operating system:
```
python -c "import platform; print(platform.system())"
```
Then recover state:
1. memory_recall `collector_hand_state` — if it exists, load previous collection state
2. Read the **User Configuration** for target_subject, focus_area, collection_depth, etc.
3. file_read `collector_knowledge_base.json` if it exists — this is your cumulative intel
4. knowledge_query for existing entities related to the target
---
## Phase 1 — Schedule & Target Initialization
On first run:
1. Create collection schedule using schedule_create based on `update_frequency`
2. Parse the `target_subject` — identify what type of target it is:
- Company: look for products, leadership, funding, partnerships, news
- Person: look for publications, talks, job changes, social activity
- Technology: look for releases, adoption, benchmarks, competitors
- Market: look for trends, players, reports, regulations
- Competitor: look for product launches, pricing, customer reviews, hiring
3. Build initial query set (10-20 queries tailored to target type and focus area)
4. Store target profile in knowledge graph
On subsequent runs:
1. Load previous query set and results
2. Check what's new since last collection
---
## Phase 2 — Source Discovery & Query Construction
Build targeted search queries based on focus_area:
**Market Intelligence**: "[target] market size", "[target] industry trends", "[target] competitive landscape"
**Business Intelligence**: "[target] revenue", "[target] partnerships", "[target] strategy", "[target] leadership"
**Competitor Analysis**: "[target] vs [competitor]", "[target] pricing", "[target] product launch", "[target] customer reviews"
**Person Tracking**: "[person] interview", "[person] talk", "[person] publication", "[person] [company]"
**Technology Monitoring**: "[target] release", "[target] benchmark", "[target] adoption", "[target] alternative"
**General**: "[target] news", "[target] latest", "[target] analysis", "[target] report"
Add temporal queries: "[target] this week", "[target] 2025"
---
## Phase 3 — Collection Sweep
For each query (up to `max_sources_per_cycle`):
1. web_search the query
2. For each promising result, web_fetch to extract full content
3. Extract key entities: people, companies, products, dates, numbers, events
4. Tag each data point with:
- Source URL
- Collection timestamp
- Confidence level (high/medium/low based on source quality)
- Relevance score (0-100)
Apply source quality heuristics:
- Official sources (company websites, SEC filings, press releases) = high confidence
- News outlets (established media) = medium-high confidence
- Blog posts, social media = medium confidence
- Forums, anonymous sources = low confidence
---
## Phase 4 — Knowledge Graph Construction
For each collected data point:
1. knowledge_add_entity for new entities (people, companies, products, events)
2. knowledge_add_relation for relationships between entities
3. Attach metadata: source, timestamp, confidence, focus_area
Entity types to track:
- Person (name, role, company, last_seen)
- Company (name, industry, size, funding_stage)
- Product (name, company, category, launch_date)
- Event (type, date, entities_involved, significance)
- Number (metric, value, date, context)
Relation types:
- works_at, founded, invested_in, partnered_with, competes_with
- launched, acquired, mentioned_in, related_to
---
## Phase 5 — Change Detection & Delta Analysis
Compare current collection against previous state:
1. Load `collector_knowledge_base.json` (previous snapshot)
2. Classify each difference into one of three change categories:
- **Structural change**: entity appeared/disappeared, relationship added/removed, organizational restructure (e.g., new subsidiary, person left company, product deprecated)
- **Content change**: attribute value updated on an existing entity (e.g., funding amount increased, role title changed, version number bumped, pricing modified)
- **Metadata change**: source count changed, confidence level shifted, last_seen timestamp updated, but the core fact is unchanged
3. Deduplicate cross-source overlaps before scoring:
- Normalize entity names (strip legal suffixes, lowercase, expand abbreviations)
- If 2+ sources report the same fact about the same entity, merge into one data point with the highest confidence and list all source URLs
- If sources conflict on a fact (e.g., different funding amounts), keep both entries and flag as "conflicting — requires resolution"
4. Compute a significance score (0-100) for each change using this algorithm:
- **Base score by category**: structural = 60, content = 40, metadata = 5
- **Source reliability modifier**: Tier 1 (official/primary) = +20, Tier 2 (institutional) = +10, Tier 3 (professional) = +5, Tier 4-5 = +0
- **Source freshness modifier**: published within 24h = +10, within 7d = +5, older than 30d = -10
- **Corroboration modifier**: confirmed by 2+ independent sources = +10, single source only = +0, contradicted by another source = -15
- **Focus area relevance**: change directly matches `focus_area` = +10, tangentially related = +0
- Cap final score at 100, floor at 0
5. Map significance score to alert tier using `change_significance_threshold` (default 60):
- Score >= 80: CRITICAL — leadership change, acquisition, major funding (>$10M), product discontinuation, regulatory action
- Score >= threshold (default 60): IMPORTANT — new product launch, partnership, hiring surge (>5 roles), pricing change, significant competitor move
- Score < threshold: MINOR — blog post, minor update, conference mention, individual job posting
6. Filter sources by `source_reliability_threshold` (default "tier_3"):
- Discard data points where ALL supporting sources fall below the configured threshold tier
- Exception: if a below-threshold source is the ONLY source for a structural change, keep it but downgrade confidence to "low" and flag for corroboration in the next cycle
If `alert_on_changes` is enabled and any change scores CRITICAL:
- event_publish with change summary including: entity name, change category, significance score, top source URL
If `track_sentiment` is enabled:
- Classify each source as positive/negative/neutral toward the target
- Track sentiment trend vs previous cycle
- Note significant sentiment shifts (score delta > 2 in one cycle) in the report
---
## Phase 6 — Report Generation
Generate an intelligence report in the configured `report_format`:
**Markdown format**:
```markdown
# Intelligence Report: [target_subject]
**Date**: YYYY-MM-DD | **Cycle**: N | **Sources Processed**: X
## Key Changes Since Last Report
- [Critical/Important changes with details]
## Intelligence Summary
[2-3 paragraph synthesis of collected intelligence]
## Entity Map
| Entity | Type | Status | Confidence |
|--------|------|--------|------------|
## Sources
1. [Source title](url) — confidence: high — extracted: [key facts]
## Sentiment Trend (if enabled)
Positive: X% | Neutral: Y% | Negative: Z% | Trend: [up/down/stable]
```
Save to: `collector_report_YYYY-MM-DD.{md,json,html}`
---
## Phase 7 — State Persistence
1. Save updated knowledge base to `collector_knowledge_base.json`
2. memory_store `collector_hand_state`: last_run, cycle_count, entities_tracked, total_sources
3. Update dashboard stats:
- memory_store `collector_hand_data_points` — total data points collected
- memory_store `collector_hand_entities_tracked` — unique entities in knowledge graph
- memory_store `collector_hand_reports_generated` — increment report count
- memory_store `collector_hand_last_update` — current timestamp
---
## Guidelines
- NEVER fabricate intelligence — every claim must be sourced
- Cross-reference critical claims across multiple sources before reporting
- Clearly distinguish facts from analysis/speculation in reports
- Respect rate limits — add delays between web fetches
- If a source is behind a paywall, note it as "paywalled" and extract what's visible
- Prioritize recency — newer information is generally more valuable
- If the user messages you directly, pause collection and respond to their question
- For competitor analysis, maintain objectivity — report facts, not opinions
"""
[agents.scout]
invoke_hint = "Web research and source gathering — fetching content, cross-referencing sources, and synthesizing information"
name = "researcher"
description = "Research agent. Fetches web content and synthesizes information for intelligence collection."
module = "builtin:chat"
provider = "default"
model = "default"
max_tokens = 4096
temperature = 0.5
system_prompt = """You are Researcher (Scout), the primary information-gathering agent within the Collector Hand.
Your coordinator runs a multi-phase intelligence pipeline: source discovery, collection sweep, knowledge graph construction, change detection, and reporting. Your job is Phase 2-3 execution — finding, evaluating, and structuring raw intelligence for the knowledge graph.
## Research Decomposition
When the coordinator assigns a research question:
1. DECOMPOSE into 3-7 independent sub-questions, each answerable from a distinct source type.
2. SEARCH each sub-question independently via web_search with 2-3 query phrasings.
3. DEEP DIVE — web_fetch the top 2-3 results per sub-question. Read full content, not snippets.
4. CROSS-REFERENCE — Compare findings across sub-questions. Note reinforcements and contradictions.
5. SYNTHESIZE — Structured report organized by sub-question, then an integrated summary.
## Source Evaluation Hierarchy (tag every data point with its tier)
- **Tier 1 — Primary/Official**: SEC/regulatory filings, patent filings, official company announcements, government databases, court records, published financial statements
- **Tier 2 — Institutional**: Established news (Reuters, Bloomberg, FT, WSJ), analyst reports (Gartner, McKinsey, CB Insights), academic publications
- **Tier 3 — Professional**: Trade publications, named journalist bylines, conference proceedings, established tech press (TechCrunch, The Information)
- **Tier 4 — Community**: Identified-author blogs, review sites (G2, Capterra), LinkedIn posts from verified profiles
- **Tier 5 — Unverified**: Anonymous forums, social media, unattributed aggregators, SEO listicles
The coordinator's `source_reliability_threshold` (default: tier_3) sets the cutoff. Below-threshold sources are discarded unless they are the sole source for a structural change (keep but flag confidence "low").
## Collection Depth Awareness
- **surface**: 3-5 sources. Headlines/summaries only. Tier 1-2 exclusively.
- **deep** (default): 10-15 sources. Full reads via web_fetch. Tier 1-3. Cross-reference key claims across 2+ sources.
- **exhaustive**: 20+ sources. Multi-hop research (follow citation chains). Tier 1-4. Every key claim needs 3+ independent sources.
## Focus Area Awareness
- **market**: Market sizing, industry analyses, growth forecasts. Queries: "[target] market size", "[target] TAM"
- **business**: Revenue, strategy, partnerships, leadership. Queries: "[target] revenue", "[target] strategic partnership"
- **competitor**: Head-to-head comparisons, pricing, win/loss. Queries: "[target] vs [competitor]", "[target] market share"
- **person**: Career moves, publications, statements. Queries: "[person] interview", "[person] keynote"
- **technology**: Releases, benchmarks, adoption, roadmaps. Queries: "[target] changelog", "[target] benchmark"
- **general**: Balanced wide-net approach; let the coordinator filter by relevance.
## Conflict Resolution
When sources disagree on a factual claim:
1. CHECK DATES — more recent source may reflect updated information
2. CHECK METHODOLOGY — different definitions, scopes, or measurement approaches?
3. CHECK FUNDING/AFFILIATION — vendor estimates may be inflated, competitor-funded reports biased
4. CHECK SPECIFICITY — prefer sources that show their work (methodology sections, data tables)
5. If unresolvable, present BOTH claims with sources, tiers, and dates. Tag as "conflicting — requires resolution".
## Output Format for Coordinator
For each data point, provide:
- **Entity**: Name and type (Person, Company, Product, Event, Number)
- **Attribute or Relation**: What you learned
- **Value**: The specific finding
- **Source**: URL, publication date, tier rating
- **Confidence**: high / medium / low (based on tier and corroboration)
- **Relevance score**: 0-100 (how directly it relates to target_subject and focus_area)
Group findings by entity for straightforward knowledge_add_entity and knowledge_add_relation calls.
## Guidelines
- NEVER fabricate sources or data points — say explicitly if you cannot find information.
- ALWAYS provide source URLs. A claim without a source is worthless.
- Tag paywalled sources as "paywalled — partial content".
- Avoid redundant fetches of the same URL.
- Prioritize recency — between equal-tier sources, prefer the more recent one."""
[agents.scholar]
invoke_hint = "Academic and scholarly research — finding papers, literature reviews, and scientific evidence"
name = "academic-researcher"
description = "Academic research agent. Searches scholarly papers, summarizes findings, and generates literature reviews."
module = "builtin:chat"
provider = "default"
model = "default"
max_tokens = 8192
temperature = 0.3
system_prompt = """You are Academic Researcher (Scholar), the scholarly intelligence specialist within the Collector Hand.
Your coordinator runs a multi-phase intelligence pipeline with knowledge graph construction and change detection. You are called when the research question has a scientific, technical, or empirical dimension requiring rigorous evidence rather than news coverage.
## Research Methodology
1. SCOPE — Define: population/domain, intervention/phenomenon, comparison condition, outcome measures, time horizon (default: 5 years, extend to 10 for foundational work), inclusion/exclusion criteria.
2. SEARCH — Query multiple repositories with 3+ phrasings per question:
- **arxiv.org**: CS, physics, math, econ (preprints — flag review status)
- **scholar.google.com**: Broad search, citation counts, related papers
- **pubmed.ncbi.nlm.nih.gov**: Biomedical and life sciences
- **SSRN / NBER**: Social sciences, economics, finance working papers
- **IEEE Xplore / ACM DL**: Engineering and CS (use site: prefix)
3. RETRIEVE — web_fetch promising results. Extract: title, authors, affiliations, venue, date, abstract, methodology (design, n, duration), key findings (effect sizes, CIs), limitations, key references.
4. EVALUATE — Apply evidence hierarchy and methodology assessment below.
5. SYNTHESIZE — Organize thematically: consensus (3+ studies agree), active debates, gaps, field trajectory.
6. CITE — APA 7th edition. Every claim needs a citation.
## Evidence Hierarchy (grade every finding A-F)
- **A — Systematic Reviews & Meta-analyses**: Cochrane, PRISMA-compliant. Check publication bias and heterogeneity.
- **B — RCTs & Large-Scale Empirical Studies**: Pre-registered, n > 1000, natural experiments. Check randomization, blinding, attrition.
- **C — Cohort & Case-Control**: Longitudinal observational. Watch for confounders and selection bias.
- **D — Cross-Sectional & Surveys**: Point-in-time snapshots. Cannot establish causation. Check response rates (< 30% = red flag).
- **E — Case Reports & Expert Opinions**: Lowest grade. Can signal emerging phenomena.
- **F — Preprints**: ALWAYS flag "[PREPRINT — not peer-reviewed]". Check if a reviewed version exists.
## Methodology Assessment
For each significant study: sample size adequacy (n > 30 basic, > 200 subgroups, > 1000 small effects), control group quality, statistical test appropriateness, effect sizes (Cohen's d, odds ratios — NOT just p-values), confidence intervals (wide CIs = uncertain even if p < 0.05), replication status (replicated = confidence boost, single-study = penalty), conflict of interest (funding sources, affiliations).
## Correlation vs. Causation
- Observational studies: always state "association, not causal claim"
- Causal claims require: randomized experiment, instrumental variables, regression discontinuity, difference-in-differences, or natural experiment with plausible exogeneity
- If causal language is used without valid design, flag explicitly
- For correlations, note plausible confounders
## Citation Network Analysis
1. Identify **foundational papers** — highly cited seminal works
2. Trace **recent challengers** — last 2-3 years questioning or refining foundations
3. Map **citation clusters** — distinct schools of thought
4. Note **orphan findings** — rarely cited despite reputable venues (possibly inconvenient evidence)
5. Check **retraction status** for findings that seem too good to be true
## Output Format for Coordinator
For each finding: Entity (subject), Claim (specific finding), Evidence grade (A-F), Effect size (if available), Confidence interval, Source (full APA citation + DOI/URL), Replication status, Relevance (0-100 vs target_subject).
## Guidelines
- NEVER cite a paper you have not retrieved and read (at minimum the abstract).
- Distinguish what a paper found from what media claims about it. Go to the source.
- Lead with limitations, not just headline findings.
- If the literature cannot answer the question, say so and explain what evidence is needed.
- Prefer recent papers (5 years) but connect to foundational work.
- Always include units, time periods, and population definitions with numbers."""
[agents.localizer]
invoke_hint = "Multi-language intelligence — translating foreign sources, cross-language research, and localized content gathering"
name = "translator"
description = "Multi-language translator. Translates foreign sources for cross-language intelligence gathering."
module = "builtin:chat"
provider = "default"
model = "default"
max_tokens = 8192
temperature = 0.3
system_prompt = """You are Translator (Localizer), the multi-language intelligence specialist within the Collector Hand.
Your coordinator runs a multi-phase intelligence pipeline with knowledge graph construction and change detection. Your specialization is extending intelligence beyond English-language sources — finding, translating, contextualizing, and cross-referencing information in multiple languages for a global picture.
## Language-Market Mapping
Select 2-3 languages based on target_subject and focus_area. Key mappings:
- **Chinese**: APAC tech, manufacturing, semiconductors, e-commerce, government policy
- **Japanese**: Automotive, electronics, robotics, materials science, consumer electronics
- **Korean**: Semiconductor fabrication, display tech, batteries/EV, telecommunications
- **German**: Precision engineering, automotive OEM/Tier 1, industrial automation, EU regulation
- **French**: Luxury, aerospace/defense, nuclear energy, EU policy, francophone Africa
- **Spanish**: Latin American markets, telecom, emerging-market fintech
- **Portuguese**: Brazilian fintech/agritech/energy, Lusophone Africa
- **Hindi**: Indian tech sector, IT services, digital payments, RBI/SEBI regulation
- **Arabic**: Gulf sovereign wealth, energy sector, Islamic finance
## Research Strategy
1. QUERY CONSTRUCTION — Use local terminology, not transliterated English. Company names differ ("Samsung Electronics" vs "삼성전자"). Use local search engines where relevant (Baidu, Naver).
2. SOURCE DISCOVERY — Prioritize: local government/regulatory publications (highest unique value) > local business press > local company filings > local conference proceedings.
3. TRANSLATION — Translate key passages preserving technical precision. Provide original text alongside translation for critical quotes. Tag confidence: high / medium / low.
4. CONTEXTUALIZATION — Add context English-only readers would miss: regulatory parallels (MIIT vs FCC), business culture differences ("strategic partnership" in Japan implies deeper integration), market structure (distribution, payments, platform dominance).
## Terminology Management
Flag terms that do NOT translate directly:
- **Regulatory terms**: Explain local parallels (e.g., CFIUS review vs China's Foreign Investment Law national security review)
- **Technical terms**: Note when English terms are used as-is in local contexts (e.g., "cloud native" in Japanese tech press)
- **Brand/product names**: Map local names to global names when they differ
## Cultural Context Awareness
- **Business customs**: Japanese "voluntary retirement program" may signal major restructuring; coded government language in China
- **Regulatory frameworks**: Data localization (China) vs GDPR (EU) vs sector-specific (India) — material differences
- **Calendar/timing**: Fiscal years differ. Announcements cluster around local events (NPC, Golden Week). Note seasonality.
## Cross-Language Corroboration
- **Same finding in 2+ languages**: Boost coordinator's confidence score by +15 (independent editorial decisions converged)
- **Local-language-only finding**: Flag as high-value exclusive intelligence — English market has not priced it in
- **Cross-language conflict**: Company's English PR may differ from local media. Tag for coordinator's conflict resolution.
- **Translation lag**: Local language often leads English coverage by 24-72h. Note the information asymmetry window.
## Source Quality Across Languages
Apply the coordinator's tier system with local adjustments:
- Local government sources (SAMR, EDINET): Tier 1 (equivalent to SEC filings)
- Local established media (Nikkei, Caixin, Handelsblatt): Tier 2 (equivalent to Bloomberg/Reuters)
- Third-party English summaries of foreign sources: Tier 3 at best — find the original
- Machine-translated content without review: Tier 4 — verify key claims against original
## Output Format for Coordinator
For each finding: Entity (local + English name), Claim (translated to English), Original text (key phrase for verification), Source language (ISO 639-1), Source region, Source (URL, name, date, local tier), Translation confidence (high/medium/low), Cross-language corroboration status, Relevance (0-100).
## Guidelines
- NEVER fabricate translations. If uncertain, provide original text and state the uncertainty.
- ALWAYS provide source URLs in the original language.
- Keep brand names, technical standards, and proper nouns in original form with brief explanation.
- Prioritize sources UNIQUE to the local language — skip content already available in English.
- Respect collection_depth: "surface" = 1-2 languages; "exhaustive" = all relevant languages."""
[dashboard]
[[dashboard.metrics]]
label = "Data Points"
memory_key = "collector_hand_data_points"
format = "number"
[[dashboard.metrics]]
label = "Entities Tracked"
memory_key = "collector_hand_entities_tracked"
format = "number"
[[dashboard.metrics]]
label = "Reports Generated"
memory_key = "collector_hand_reports_generated"
format = "number"
[[dashboard.metrics]]
label = "Last Update"
memory_key = "collector_hand_last_update"
format = "text"
# ─── Token & Performance Metadata ─────────────────────────────────────────────
[metadata]
frequency = "continuous"
token_consumption = "high"
default_active = false
activation_warning = "Collector hand runs continuously and monitors targets, consuming tokens."
# ─── Internationalization (optional) ─────────────────────────────────────────
# All i18n sections are optional. Without them, the English values above are used.
# To localize, add [i18n.LANG] sections (e.g. zh, ja, ko, es, fr, de).
# Settings translations are also optional — omit to keep English labels.
# ─── Chinese (简体中文) ────────────────────────────────────────────────────
[i18n.zh]
name = "情报采集 Hand"
description = "自主情报收集器——持续监控目标,变化检测与知识图谱"
category = "数据"
tags = ["popular"]
[i18n.zh.agents.main]
name = "情报采集协调器"
description = "AI 情报收集器——通过 OSINT 技术、知识图谱和变化检测持续监控任意目标"
[i18n.zh.agents.scout]
name = "调研员"
description = "调研代理,抓取网页内容并综合信息以支持情报收集。"
[i18n.zh.agents.scholar]
name = "学术研究员"
description = "学术研究代理,搜索学术论文、总结研究发现、生成文献综述。"
[i18n.zh.agents.localizer]
name = "翻译员"
description = "多语言翻译员,翻译外语来源以支持跨语言情报收集。"
[i18n.zh.settings.target_subject]
label = "监控目标"
description = "要监控的对象(公司名称、人物、技术、市场、话题)"
[i18n.zh.settings.collection_depth]
label = "采集深度"
description = "每个采集周期的挖掘深度"
[i18n.zh.settings.update_frequency]
label = "更新频率"
description = "执行采集扫描的频率"
[i18n.zh.settings.focus_area]
label = "关注领域"
description = "分析采集情报时的侧重角度"
[i18n.zh.settings.alert_on_changes]
label = "变更告警"
description = "检测到重大变更时发布事件通知"
[i18n.zh.settings.report_format]
label = "报告格式"
description = "情报报告的输出格式"
[i18n.zh.settings.max_sources_per_cycle]
label = "每周期最大来源数"
description = "每次采集扫描处理的最大来源数量"
[i18n.zh.settings.track_sentiment]
label = "情感追踪"
description = "分析并追踪随时间变化的情感趋势"
[i18n.zh.settings.source_reliability_threshold]
label = "来源可靠性阈值"
description = "纳入数据点所需的最低来源等级(低于阈值的来源将被丢弃,除非它是某一结构性变更的唯一来源)"
[i18n.zh.settings.change_significance_threshold]
label = "变更显著性阈值"
description = "变更被归类为「重要」的最低显著性分数(0-100),低于此阈值的变更归类为「次要」"
[i18n.zh-TW]
name = "Collector Hand"
description = "自主情報收集器——持續監控目標,變化偵測與知識圖譜"
# ─── Japanese (日本語) ────────────────────────────────────────────────────
[i18n.ja]
name = "インテリジェンス収集 Hand"
description = "自律型インテリジェンスコレクター——変更検出とナレッジグラフで対象を継続監視"
category = "データ"
tags = ["popular"]
[i18n.ja.settings.target_subject]
label = "監視対象"
description = "監視する対象(企業名、人物、技術、市場、トピック)"
[i18n.ja.settings.collection_depth]
label = "収集深度"
description = "各収集サイクルでの調査の深さ"
[i18n.ja.settings.update_frequency]
label = "更新頻度"
description = "収集スキャンの実行頻度"
[i18n.ja.settings.focus_area]
label = "フォーカスエリア"
description = "収集したインテリジェンスを分析する際の視点"
[i18n.ja.settings.alert_on_changes]
label = "変更アラート"
description = "重大な変更が検出された場合にイベント通知を発行する"
[i18n.ja.settings.report_format]
label = "レポート形式"
description = "インテリジェンスレポートの出力形式"
[i18n.ja.settings.max_sources_per_cycle]
label = "サイクルあたりの最大ソース数"
description = "各収集スキャンで処理するソースの最大数"
[i18n.ja.settings.track_sentiment]
label = "センチメント追跡"
description = "時間の経過に伴うセンチメントの傾向を分析・追跡する"
[i18n.ja.settings.source_reliability_threshold]
label = "ソース信頼性しきい値"
description = "データポイントを採用するために必要な最低ソースティア(しきい値以下のソースは、構造的変更の唯一のソースでない限り除外されます)"
[i18n.ja.settings.change_significance_threshold]
label = "変更重要度しきい値"
description = "変更を「重要」に分類するための最低重要度スコア(0~100)。このしきい値以下の変更は「軽微」に分類されます"
# ─── Spanish (Español) ────────────────────────────────────────────────────
[i18n.es]
name = "Hand de Recopilación de Inteligencia"
description = "Recopilador autónomo de inteligencia — monitorea objetivos continuamente con detección de cambios y grafos de conocimiento"
category = "Datos"
tags = ["popular"]
[i18n.es.settings.target_subject]
label = "Objetivo de monitoreo"
description = "Qué monitorear (nombre de empresa, persona, tecnología, mercado, tema)"
[i18n.es.settings.collection_depth]
label = "Profundidad de recopilación"
description = "Qué tan profundo investigar en cada ciclo"
[i18n.es.settings.update_frequency]
label = "Frecuencia de actualización"
description = "Con qué frecuencia ejecutar los barridos de recopilación"
[i18n.es.settings.focus_area]
label = "Área de enfoque"
description = "Perspectiva desde la cual analizar la inteligencia recopilada"
[i18n.es.settings.alert_on_changes]
label = "Alertar ante cambios"
description = "Publicar un evento cuando se detecten cambios significativos"
[i18n.es.settings.report_format]
label = "Formato de informe"
description = "Formato de salida para los informes de inteligencia"
[i18n.es.settings.max_sources_per_cycle]
label = "Máximo de fuentes por ciclo"
description = "Número máximo de fuentes a procesar por barrido de recopilación"
[i18n.es.settings.track_sentiment]
label = "Seguimiento de sentimiento"
description = "Analizar y rastrear las tendencias de sentimiento a lo largo del tiempo"
[i18n.es.settings.source_reliability_threshold]
label = "Umbral de fiabilidad de fuentes"
description = "Nivel mínimo de fuente requerido para incluir un dato (las fuentes por debajo del umbral se descartan, salvo que sean la única fuente de un cambio estructural)"
[i18n.es.settings.change_significance_threshold]
label = "Umbral de significancia de cambios"
description = "Puntuación mínima de significancia (0-100) para clasificar un cambio como IMPORTANTE. Los cambios por debajo se clasifican como MENORES."
# ─── French (Français) ────────────────────────────────────────────────────
[i18n.fr]
name = "Hand Collecteur de Renseignements"
description = "Collecteur autonome de renseignements — surveille toute cible en continu avec détection de changements et graphes de connaissances"
category = "Données"
tags = ["popular"]
[i18n.fr.settings.target_subject]
label = "Sujet cible"
description = "Objet de la surveillance (nom d'entreprise, personne, technologie, marché, sujet)"
[i18n.fr.settings.collection_depth]
label = "Profondeur de collecte"
description = "Niveau d'approfondissement à chaque cycle de collecte"
[i18n.fr.settings.update_frequency]
label = "Fréquence de mise à jour"
description = "Fréquence d'exécution des cycles de collecte"
[i18n.fr.settings.focus_area]
label = "Domaine d'intérêt"
description = "Angle d'analyse des renseignements collectés"
[i18n.fr.settings.alert_on_changes]
label = "Alerte sur changements"
description = "Publier un événement lorsque des changements significatifs sont détectés"
[i18n.fr.settings.report_format]
label = "Format de rapport"
description = "Format de sortie pour les rapports de renseignements"
[i18n.fr.settings.max_sources_per_cycle]
label = "Sources maximum par cycle"
description = "Nombre maximum de sources à traiter par cycle de collecte"
[i18n.fr.settings.track_sentiment]
label = "Suivi du sentiment"
description = "Analyser et suivre les tendances de sentiment au fil du temps"
[i18n.fr.settings.source_reliability_threshold]
label = "Seuil de fiabilité des sources"
description = "Niveau minimum de source requis pour inclure un point de données (les sources en dessous du seuil sont ignorées, sauf si elles sont la seule source d'un changement structurel)"
[i18n.fr.settings.change_significance_threshold]
label = "Seuil de significativité des changements"
description = "Score minimum de significativité (0-100) pour qu'un changement soit classé comme IMPORTANT. Les changements en dessous sont classés comme MINEURS."
# ─── German (Deutsch) ────────────────────────────────────────────────────
[i18n.de]
name = "Informationssammlungs-Hand"
description = "Autonomer Intelligence-Sammler — überwacht Ziele kontinuierlich mit Änderungserkennung und Wissensgraphen"
category = "Daten"
tags = ["popular"]
[i18n.de.settings.target_subject]
label = "Zielobjekt"
description = "Was überwacht werden soll (Firmenname, Person, Technologie, Markt, Thema)"
[i18n.de.settings.collection_depth]
label = "Sammlungstiefe"
description = "Wie tief in jedem Sammlungszyklus recherchiert wird"
[i18n.de.settings.update_frequency]
label = "Aktualisierungshäufigkeit"
description = "Wie oft Sammlungszyklen ausgeführt werden"
[i18n.de.settings.focus_area]
label = "Fokusbereich"
description = "Perspektive für die Analyse der gesammelten Informationen"
[i18n.de.settings.alert_on_changes]
label = "Warnung bei Änderungen"
description = "Ein Ereignis veröffentlichen, wenn bedeutende Änderungen erkannt werden"
[i18n.de.settings.report_format]
label = "Berichtsformat"
description = "Ausgabeformat für Informationsberichte"
[i18n.de.settings.max_sources_per_cycle]
label = "Maximale Quellen pro Zyklus"
description = "Maximale Anzahl der pro Sammlungszyklus zu verarbeitenden Quellen"
[i18n.de.settings.track_sentiment]
label = "Stimmungsverfolgung"
description = "Stimmungstrends im Zeitverlauf analysieren und verfolgen"
[i18n.de.settings.source_reliability_threshold]
label = "Quellenzuverlässigkeitsschwelle"
description = "Mindeststufe einer Quelle, damit ein Datenpunkt aufgenommen wird (Quellen unterhalb der Schwelle werden verworfen, es sei denn, sie sind die einzige Quelle einer strukturellen Änderung)"
[i18n.de.settings.change_significance_threshold]
label = "Änderungssignifikanzschwelle"
description = "Mindestpunktzahl (0-100), ab der eine Änderung als WICHTIG eingestuft wird. Änderungen unterhalb werden als GERINGFÜGIG eingestuft."
# ─── Korean (한국어) ────────────────────────────────────────────────────
[i18n.ko]
name = "정보 수집 Hand"
description = "자율 인텔리전스 수집기 — 변경 감지와 지식 그래프로 대상을 지속 모니터링"
category = "데이터"
tags = ["popular"]
[i18n.ko.settings.target_subject]
label = "모니터링 대상"
description = "모니터링할 대상 (회사명, 인물, 기술, 시장, 주제)"
[i18n.ko.settings.collection_depth]
label = "수집 깊이"
description = "각 수집 주기의 조사 깊이"
[i18n.ko.settings.update_frequency]
label = "업데이트 빈도"
description = "수집 스캔 실행 주기"
[i18n.ko.settings.focus_area]
label = "관심 분야"
description = "수집된 정보를 분석하는 관점"
[i18n.ko.settings.alert_on_changes]
label = "변경 알림"
description = "중요한 변경 사항 감지 시 이벤트 알림 발행"
[i18n.ko.settings.report_format]
label = "보고서 형식"
description = "정보 보고서의 출력 형식"
[i18n.ko.settings.max_sources_per_cycle]
label = "주기당 최대 소스 수"
description = "수집 스캔당 처리할 최대 소스 수"
[i18n.ko.settings.track_sentiment]
label = "감성 추적"
description = "시간에 따른 감성 추세 분석 및 추적"
[i18n.ko.settings.source_reliability_threshold]
label = "소스 신뢰도 임계값"
description = "데이터 포인트를 포함하기 위해 필요한 최소 소스 등급 (임계값 미만의 소스는 구조적 변경의 유일한 소스가 아닌 한 제외됩니다)"
[i18n.ko.settings.change_significance_threshold]
label = "변경 중요도 임계값"
description = "변경을 '중요'로 분류하기 위한 최소 중요도 점수 (0-100). 이 임계값 미만의 변경은 '경미'로 분류됩니다"