Files
librefang-registry/skills/data-pipeline/SKILL.md
T
Evan b567db71ba feat(skills): restore 60 bundled skills (#42)
* feat(skills): restore ansible skill

* feat(skills): restore api-tester skill

* feat(skills): restore aws skill

* feat(skills): restore azure skill

* feat(skills): restore ci-cd skill

* feat(skills): restore code-reviewer skill

* feat(skills): restore compliance skill

* feat(skills): restore confluence skill

* feat(skills): restore crypto-expert skill

* feat(skills): restore css-expert skill

* feat(skills): restore data-analyst skill

* feat(skills): restore data-pipeline skill

* feat(skills): restore docker skill

* feat(skills): restore elasticsearch skill

* feat(skills): restore email-writer skill

* feat(skills): restore figma-expert skill

* feat(skills): restore gcp skill

* feat(skills): restore git-expert skill

* feat(skills): restore github skill

* feat(skills): restore golang-expert skill

* feat(skills): restore graphql-expert skill

* feat(skills): restore helm skill

* feat(skills): restore interview-prep skill

* feat(skills): restore jira skill

* feat(skills): restore kubernetes skill

* feat(skills): restore linear-tools skill

* feat(skills): restore linux-networking skill

* feat(skills): restore llm-finetuning skill

* feat(skills): restore ml-engineer skill

* feat(skills): restore mongodb skill

* feat(skills): restore nextjs-expert skill

* feat(skills): restore nginx skill

* feat(skills): restore notion skill

* feat(skills): restore oauth-expert skill

* feat(skills): restore openapi-expert skill

* feat(skills): restore pdf-reader skill

* feat(skills): restore postgres-expert skill

* feat(skills): restore presentation skill

* feat(skills): restore project-manager skill

* feat(skills): restore prometheus skill

* feat(skills): restore prompt-engineer skill

* feat(skills): restore python-expert skill

* feat(skills): restore react-expert skill

* feat(skills): restore redis-expert skill

* feat(skills): restore regex-expert skill

* feat(skills): restore rust-expert skill

* feat(skills): restore security-audit skill

* feat(skills): restore sentry skill

* feat(skills): restore shell-scripting skill

* feat(skills): restore slack-tools skill

* feat(skills): restore sql-analyst skill

* feat(skills): restore sqlite-expert skill

* feat(skills): restore sysadmin skill

* feat(skills): restore technical-writer skill

* feat(skills): restore terraform skill

* feat(skills): restore typescript-expert skill

* feat(skills): restore vector-db skill

* feat(skills): restore wasm-expert skill

* feat(skills): restore web-search skill

* feat(skills): restore writing-coach skill
2026-04-08 15:18:53 +08:00

39 lines
3.3 KiB
Markdown

---
name: data-pipeline
description: "Data pipeline expert for ETL, Apache Spark, Airflow, dbt, and data quality"
---
# Data Pipeline Expert
A data engineering specialist with extensive experience designing and operating production ETL/ELT pipelines, orchestration frameworks, and data quality systems. This skill provides guidance for building reliable, observable, and scalable data pipelines using industry-standard tools like Apache Airflow, Spark, and dbt across batch and streaming architectures.
## Key Principles
- Prefer ELT over ETL when your target warehouse can handle transformations; load raw data first, then transform in place for reproducibility and auditability
- Design every pipeline step to be idempotent; re-running a task with the same inputs must produce the same outputs without side effects or duplicates
- Partition data by time or logical keys at every stage; partitioning enables incremental processing, efficient pruning, and manageable backfill operations
- Instrument pipelines with data quality checks between stages; catching bad data early prevents cascading corruption through downstream tables
- Separate orchestration (when and what order) from computation (how); the scheduler should not perform heavy data processing itself
## Techniques
- Build Airflow DAGs with task-level retries, timeouts, and SLAs; use sensors for external dependencies and XCom for lightweight inter-task communication
- Design Spark jobs with proper partitioning (repartition/coalesce), broadcast joins for small dimension tables, and caching for reused DataFrames
- Structure dbt projects with staging models (source cleaning), intermediate models (business logic), and mart models (final consumption tables)
- Write dbt tests at multiple levels: schema tests (not_null, unique, accepted_values), relationship tests, and custom data tests for business rules
- Implement data quality gates using frameworks like Great Expectations: define expectations on row counts, column distributions, and referential integrity
- Use Change Data Capture (CDC) patterns with tools like Debezium to stream database changes into event pipelines without polling
## Common Patterns
- **Incremental Load**: Process only new or changed records using high-watermark columns (updated_at) or CDC events, falling back to full reload on schema changes
- **Backfill Strategy**: Design DAGs with date-parameterized runs so historical reprocessing uses the same code path as daily runs, just with different date ranges
- **Dead Letter Queue**: Route failed records to a separate table or topic for investigation and reprocessing instead of halting the entire pipeline
- **Schema Evolution**: Use schema registries (Avro, Protobuf) or column-add-only policies to evolve data contracts without breaking downstream consumers
## Pitfalls to Avoid
- Do not perform heavy computation inside Airflow operators; delegate to Spark, dbt, or external compute and use Airflow only for orchestration
- Do not skip data validation after ingestion; silent schema changes from upstream sources are the most common cause of pipeline failures
- Do not hardcode connection strings or credentials in pipeline code; use secrets managers and environment-based configuration
- Do not run full table scans on every pipeline execution when incremental processing is feasible; it wastes compute and increases latency