Files
librefang-registry/workflows/data-pipeline.toml
T
Evan 05bdf02169 feat(workflows): expand template library from 9 to 22 + multiline string cleanup (#36)
* feat(workflows): add 13 workflow templates across engineering, business, and productivity

Engineering:
- bug-triage: reproduce path → root cause → fix plan
- api-design: resource model → endpoints → OpenAPI spec
- incident-postmortem: timeline → RCA → full postmortem report
- test-generation: code analysis → edge cases → full test suite
- refactor-plan: smell analysis → prioritised opportunities → migration plan

Business:
- competitor-analysis: profiles → SWOT → strategy report
- product-spec: problem definition → user stories → full PRD
- market-research: landscape → segments → research report

Productivity/Thinking:
- meeting-summary: raw notes → structured summary → follow-up email
- decision-matrix: criteria → weighted scoring → recommendation memo
- learning-plan: gap analysis → roadmap → week-1 day-by-day plan
- job-application: job analysis → tailored resume → cover letter → interview prep
- blog-post: research → outline → draft → SEO optimisation

Closes #1912 on librefang/librefang

* fix(workflows): overhaul existing 9 templates

- data-pipeline: redesigned — original 'extract from URL' step was
  broken (LLMs cannot fetch URLs); replaced with paste-data approach
  (profile → clean_transform → analyse) with analysis_goal parameter
- translate-polish: added target_language and register parameters;
  added back-translation step for accuracy verification
- weekly-report: added team/audience parameters, richer extraction
  step, added Metrics and Notes sections
- content-pipeline: added audience/tone parameters, added outline step
  between research and writing
- content-review: fix category 'content' → 'creation'
- customer-support: fix category 'support' → 'business'

* style(workflows): convert all prompt_template strings to TOML multiline syntax

Replace \n escape sequences with real newlines using triple-quote
multiline strings ("""...""") across all 22 workflow templates.
No content changes — formatting only.

* style(hands): replace \n escape in reddit writer format example with multiline code block
2026-04-01 18:21:09 +08:00

83 lines
2.3 KiB
TOML

id = "data-pipeline"
name = "Data Analysis & Transform"
description = "Clean, transform, analyse, and summarise raw data. Handles messy inputs: CSV, JSON, tables, or free-form text."
category = "engineering"
tags = ["data", "analysis", "transform", "etl"]
[i18n.zh]
name = "数据分析与转换"
description = "清洗、转换、分析并摘要原始数据,支持 CSV、JSON、表格或自由文本等杂乱输入。"
[[parameters]]
name = "raw_data"
description = "The raw data to process (paste CSV, JSON, table, or any structured text)"
param_type = "string"
required = true
[[parameters]]
name = "output_format"
description = "Desired output format for the transformed data (json, csv, markdown-table)"
param_type = "string"
required = false
default = "json"
[[parameters]]
name = "analysis_goal"
description = "What you want to understand or extract from the data"
param_type = "string"
required = false
default = "general summary and key insights"
[[steps]]
name = "profile"
prompt_template = """
You are a data engineer. Profile the following raw data:
1. Detected format and structure
2. Number of rows and columns/fields
3. Data types per field
4. Missing value count per field
5. Obvious data quality issues (duplicates, inconsistent formatting, out-of-range values)
6. Sample of first 5 rows for reference
Raw data:
{{raw_data}}
"""
[[steps]]
name = "clean_transform"
prompt_template = """
Clean and transform the raw data based on the profile below. Apply:
- Trim whitespace from string fields
- Normalise dates to ISO-8601
- Remove exact duplicate rows
- Convert empty strings to null
- Standardise inconsistent categorical values (e.g. 'Y'/'Yes'/'yes' → true)
- Flag but preserve rows with suspicious values (do not silently drop them)
Output the cleaned data in {{output_format}} format, followed by a change log listing every transformation applied.
Data profile:
{{profile}}
Raw data:
{{raw_data}}
"""
depends_on = ["profile"]
[[steps]]
name = "analyse"
prompt_template = """
Analyse the cleaned data to address the following goal: {{analysis_goal}}
Provide:
1. Key statistics (counts, totals, averages, distributions as appropriate)
2. Notable patterns or trends
3. Outliers or anomalies worth investigating
4. Correlations between fields (if applicable)
5. Top 3 actionable insights from the data
Cleaned data:
{{clean_transform}}
"""
depends_on = ["clean_transform"]