Files
librefang-registry/workflows/data-pipeline.toml
T
Evan 7881d327a5 refactor: migrate icon fields from emoji to lucide:<name> tokens (#63)
* refactor: migrate icon fields from emoji to lucide:<name> tokens

Every TOML manifest's `icon = "<emoji>"` line is replaced with
`icon = "lucide:<kebab-name>"` — a reference to a lucide-react icon,
which the librefang.ai site and dashboard render as crisp SVG. Reasons
for the switch:

- Emoji render very differently across OS/browser/font stacks; the
  registry catalog looked inconsistent from one row to the next.
- Five manifests (clip / creator / linkedin / reddit / twitter) had
  their icons stored as literal Python-style escape strings
  ("\\U0001F3AC") because the TOML parser upstream never decoded
  them. Switching away from emoji drops that class of bug entirely.
- As a drive-by, also decode the \\uXXXX accent escapes in the
  [i18n.fr] block of hands/creator/HAND.toml so "Créateur" shows
  up correctly.

87 files touched. example manifests left untouched (still "TODO").

* fix: backfill i18n name + drop the single-member email category

- Every existing [i18n.<lang>] block now has a `name` field. 60 files
  previously translated description but kept the English name
  implicitly — which rendered as "some English some Chinese" in the
  registry UI. Fill in the missing name from the English brand (or a
  known localized equivalent: DingTalk→钉钉, Feishu→飞书, Email→
  电子邮件 / メール / E-Mail / Correo / Courriel, and a handful of
  hands that have Chinese product names like 视频剪辑 Hand).
- channels/email.toml was the only item under category="email";
  reclassify it as "messaging" so the sub-category filter chip list
  on the category page isn't littered with singletons.

* feat(i18n): localize 76 agents/integrations/plugins into 7 languages

Adds full [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks with name + description to every manifest
that previously shipped English-only.

Coverage:
- 32 agents (academic-researcher, analyst, architect, assistant,
  code-reviewer, coder, customer-support, data-scientist, debugger,
  devops-lead, doc-writer, email-assistant, health-tracker,
  hello-world, home-automation, legal-assistant, meeting-assistant,
  ops, orchestrator, personal-finance, planner, recipe-assistant,
  recruiter, researcher, sales-assistant, security-auditor,
  social-media, test-engineer, translator, travel-planner, tutor,
  writer)
- 33 integrations (AWS, Azure, Bitbucket, Brave Search, Discord,
  Dropbox, Elasticsearch, Exa Search, Fetch, Filesystem, GCP, Git,
  GitHub, GitLab, Gmail, Google Calendar, Google Drive, Google Maps,
  Jira, Linear, Memory, MongoDB, Notion, PostgreSQL, Puppeteer, Redis,
  Sentry, Sequential Thinking, Slack, SQLite, Teams, Time, Todoist) —
  brand names kept as-is across all locales, only descriptions
  translated.
- 11 plugins (auto-summarizer, context-decay, conversation-logger,
  episodic-memory, guardrails, keyword-memory, mempalace-indexer,
  sentiment-tracker, todo-tracker, topic-memory, user-profile)

The descriptions are one-line summaries — hand-translated rather than
machine-generated, so technical terms (MCP, PR, CI/CD, etc.) stay
consistent across locales.

* feat(i18n): close remaining per-lang gaps for channels, workflows, devteam

Third pass on i18n coverage. Every non-example manifest now carries a
full set of [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks.

- 44 channel adapters: added French descriptions (zh/zh-TW/ja/ko/de/es
  were already present). Brand names kept as-is in all locales so users
  recognize Discord / Slack / LINE / etc. consistently.
- 22 workflows: filled zh-TW / ja / ko / de / es / fr blocks. Each
  translation mirrors the existing zh one in structure and tone so the
  catalog reads consistently across locales.
- hands/devteam/HAND.toml: added the four langs that were missing
  (zh-TW, de, es, fr).

Only the 6 templates under examples/ are left without i18n blocks on
purpose — they still contain "TODO:" placeholders.
2026-04-17 22:04:26 +09:00

107 lines
3.5 KiB
TOML

id = "data-pipeline"
name = "Data Analysis & Transform"
description = "Clean, transform, analyse, and summarise raw data. Handles messy inputs: CSV, JSON, tables, or free-form text."
category = "engineering"
tags = ["data", "analysis", "transform", "etl"]
[i18n.zh]
name = "数据分析与转换"
description = "清洗、转换、分析并摘要原始数据,支持 CSV、JSON、表格或自由文本等杂乱输入。"
[[parameters]]
name = "raw_data"
description = "The raw data to process (paste CSV, JSON, table, or any structured text)"
param_type = "string"
required = true
[[parameters]]
name = "output_format"
description = "Desired output format for the transformed data (json, csv, markdown-table)"
param_type = "string"
required = false
default = "json"
[[parameters]]
name = "analysis_goal"
description = "What you want to understand or extract from the data"
param_type = "string"
required = false
default = "general summary and key insights"
[[steps]]
name = "profile"
prompt_template = """
You are a data engineer. Profile the following raw data:
1. Detected format and structure
2. Number of rows and columns/fields
3. Data types per field
4. Missing value count per field
5. Obvious data quality issues (duplicates, inconsistent formatting, out-of-range values)
6. Sample of first 5 rows for reference
Raw data:
{{raw_data}}
"""
[[steps]]
name = "clean_transform"
prompt_template = """
Clean and transform the raw data based on the profile below. Apply:
- Trim whitespace from string fields
- Normalise dates to ISO-8601
- Remove exact duplicate rows
- Convert empty strings to null
- Standardise inconsistent categorical values (e.g. 'Y'/'Yes'/'yes' → true)
- Flag but preserve rows with suspicious values (do not silently drop them)
Output the cleaned data in {{output_format}} format, followed by a change log listing every transformation applied.
Data profile:
{{profile}}
Raw data:
{{raw_data}}
"""
depends_on = ["profile"]
[[steps]]
name = "analyse"
prompt_template = """
Analyse the cleaned data to address the following goal: {{analysis_goal}}
Provide:
1. Key statistics (counts, totals, averages, distributions as appropriate)
2. Notable patterns or trends
3. Outliers or anomalies worth investigating
4. Correlations between fields (if applicable)
5. Top 3 actionable insights from the data
Cleaned data:
{{clean_transform}}
"""
depends_on = ["clean_transform"]
[i18n.zh-TW]
name = "資料分析與轉換"
description = "清洗、轉換、分析並摘要原始資料,支援 CSV、JSON、表格或自由文字等雜亂輸入。"
[i18n.ja]
name = "データ分析と変換"
description = "生データをクレンジング・変換・分析・要約。CSV、JSON、表、自由記述など乱雑な入力に対応。"
[i18n.ko]
name = "데이터 분석 및 변환"
description = "원본 데이터를 정제·변환·분석·요약합니다. CSV, JSON, 표, 자유 텍스트 등 지저분한 입력을 처리."
[i18n.de]
name = "Datenanalyse & Transformation"
description = "Rohdaten bereinigen, transformieren, analysieren und zusammenfassen. Verarbeitet unsaubere Eingaben wie CSV, JSON, Tabellen oder Freitext."
[i18n.es]
name = "Análisis y transformación de datos"
description = "Limpia, transforma, analiza y resume datos crudos. Admite entradas desordenadas: CSV, JSON, tablas o texto libre."
[i18n.fr]
name = "Analyse et transformation de données"
description = "Nettoie, transforme, analyse et résume les données brutes. Gère les entrées désordonnées : CSV, JSON, tableaux ou texte libre."