diff --git a/README.md b/README.md index c595b2d..45aa487 100644 --- a/README.md +++ b/README.md @@ -1,46 +1,149 @@ # LibreFang Registry -Community-maintained content registry for [LibreFang](https://github.com/librefang/librefang) -- the open-source Agent Operating System. +Community-maintained content registry for [LibreFang](https://github.com/librefang/librefang) — the open-source Agent Operating System. -This repository is the single source of truth for all installable content definitions. Anyone can submit a PR to add new agents, hands, integrations, skills, or provider models -- no changes to the LibreFang binary required. +This repository is the **single source of truth** for all installable content definitions. Anyone can submit a PR to add new agents, hands, integrations, skills, or provider models — no changes to the LibreFang binary required. -## Structure +## Overview + +| Type | Count | Description | +|------|------:|-------------| +| [Hands](#hands) | 14 | User-facing "apps" — agent + tools + settings + dashboard | +| [Agents](#agents) | 32 | Autonomous agent definitions with model config and tools | +| [Integrations](#integrations) | 25 | MCP server connections (GitHub, Slack, DBs, etc.) | +| [Providers](#providers) | 48 | LLM provider & model metadata with pricing | +| [Models](#providers) | 223 | Individual model definitions across all providers | +| [Aliases](#aliases) | 70 | Short names mapped to canonical model IDs | +| [Plugins](#plugins) | 10 | Memory, guardrails, and conversation plugins | +| [Skills](#skills) | 2 | Reusable prompt templates and Python scripts | +| [Workflows](#workflows) | 9 | Pre-built multi-agent workflow definitions | +| [Templates](#templates) | 6 | Starter templates for each content type | + +## Repository Structure ``` librefang-registry/ -├── agents/ # Agent definitions (TOML manifests) -│ ├── hello-world/agent.toml -│ ├── researcher/agent.toml -│ └── ... (33 agents) -├── hands/ # Hand definitions (TOML + docs) -│ ├── browser/HAND.toml -│ ├── trader/HAND.toml -│ └── ... (14 hands) -├── integrations/ # MCP server integration templates +├── agents/ # Agent definitions (TOML manifests) +│ ├── hello-world/ +│ │ └── agent.toml +│ ├── researcher/ +│ │ └── agent.toml +│ └── ... (32 agents) +├── hands/ # Hand definitions (app bundles) +│ ├── browser/ +│ │ ├── HAND.toml # Metadata, tools, settings, i18n (6 languages) +│ │ └── SKILL.md # Domain expert knowledge injected at runtime +│ ├── trader/ +│ │ ├── HAND.toml +│ │ └── SKILL.md +│ └── ... (14 hands) +├── integrations/ # MCP server integration templates │ ├── github.toml │ ├── slack.toml -│ └── ... (25 integrations) -├── skills/ # Reusable skill definitions -│ ├── custom-skill-prompt/skill.toml -│ └── custom-skill-python/ -├── providers/ # LLM provider & model metadata +│ └── ... (25 integrations) +├── providers/ # LLM provider & model metadata │ ├── anthropic.toml │ ├── openai.toml -│ └── ... (46 providers, 190+ models) -├── plugins/ # Plugin packages (10 plugins) -├── aliases.toml # Global model alias mappings -├── schema.toml # Provider/model schema reference +│ └── ... (48 providers, 223 models) +├── plugins/ # Memory, guardrails, and utility plugins +│ ├── episodic-memory/ +│ ├── guardrails/ +│ └── ... (10 plugins) +├── skills/ # Reusable skill definitions +│ ├── custom-skill-prompt/skill.toml +│ └── custom-skill-python/ +├── workflows/ # Pre-built multi-agent workflow definitions +│ ├── code-review.toml +│ ├── research.toml +│ └── ... (9 workflows) +├── templates/ # Starter templates for each content type +│ ├── agent.toml +│ ├── HAND.toml +│ └── ... (6 templates) +├── docs/ # Additional documentation +│ └── content-guide.md # Content contribution guidelines +├── aliases.toml # Global model alias mappings (70 aliases) +├── schema.toml # Provider/model schema reference ├── scripts/ -│ └── validate.py # Validation script +│ └── validate.py # Content validation script ├── CONTRIBUTING.md -└── LICENSE # MIT +└── LICENSE # MIT ``` ## Content Types +### Hands + +Hands are the **user-facing "apps"** in LibreFang. Each hand bundles an agent, tools, user-configurable settings, dashboard metrics, dependency checks, and i18n translations into a single deployable unit. + +Every hand includes a `SKILL.md` — domain-specific expert knowledge that is injected into the agent's context at runtime, giving it deep expertise in its domain. + +| Icon | Hand | Category | Description | +|:----:|------|----------|-------------| +| 📈 | analytics | data | Data collection, analysis, visualization, dashboards, and automated reporting | +| 🔌 | apitester | development | Endpoint discovery, request validation, load testing, and regression detection | +| 🌐 | browser | productivity | Web navigation, form filling, and multi-step web tasks with user approval | +| 🎬 | clip | content | Turns long-form video into viral short clips with captions and thumbnails | +| 🔍 | collector | data | Intelligence collection, change detection, and knowledge graphs | +| 👷 | devops | development | CI/CD management, infrastructure monitoring, deployment, and incident response | +| 📊 | lead | data | Lead generation, enrichment, scoring, and scheduled delivery | +| 💼 | linkedin | communication | Profile optimization, content creation, networking, and engagement | +| 🔮 | predictor | data | Signal collection, calibrated predictions, and accuracy tracking | +| 📢 | reddit | communication | Subreddit monitoring, content posting, and engagement tracking | +| 🧪 | researcher | productivity | Deep research, cross-referencing, fact-checking, and structured reports | +| 🎯 | strategist | productivity | Market research, competitive analysis, and strategic planning | +| 📈 | trader | data | Multi-signal analysis, adversarial reasoning, and risk management | +| 𝕏 | twitter | communication | Content creation, scheduled posting, engagement, and analytics | + +**HAND.toml format:** + +```toml +id = "browser" +name = "Browser Hand" +description = "Autonomous web browser" +category = "productivity" +icon = "🌐" +tools = ["browser_navigate", "browser_click", "browser_type"] + +[routing] +aliases = ["browse", "open website"] +weak_aliases = ["web", "url"] + +[[requires]] +key = "chromium" +requirement_type = "binary" +check_value = "chromium" + +[[settings]] +key = "headless" +setting_type = "toggle" +default = "true" + +[agent] +name = "browser-hand" +module = "builtin:chat" +system_prompt = """You are an autonomous web browser agent...""" + +[dashboard] +[[dashboard.metrics]] +label = "Pages Visited" +memory_key = "pages_visited" +format = "number" + +# i18n — 6 languages supported: zh, ja, ko, es, fr, de +[i18n.zh] +name = "浏览器 Hand" +description = "自主网页浏览器" +category = "生产力" + +[i18n.zh.settings.headless] +label = "无头模式" +description = "在后台运行浏览器" +``` + ### Agents -Agent definitions in `agents//agent.toml` describe autonomous agents with their model config, tools, capabilities, and routing aliases. +Agent definitions describe autonomous agents with model configuration, tools, capabilities, and routing aliases. ```toml name = "hello-world" @@ -56,30 +159,11 @@ system_prompt = "You are a helpful assistant." tools = ["web_search", "file_read"] ``` -### Hands - -Hands in `hands//HAND.toml` are higher-level application bundles -- the user-facing "apps" in LibreFang. Each hand bundles an agent config, tools, settings, dashboard metrics, and dependency requirements. - -```toml -id = "browser" -name = "Browser Hand" -category = "productivity" -tools = ["browser_navigate", "browser_click", "browser_type"] - -[agent] -name = "browser-hand" -module = "builtin:chat" -system_prompt = "You are an autonomous web browser agent..." - -[[settings]] -key = "headless" -setting_type = "toggle" -default = "true" -``` +**32 built-in agents:** academic-researcher, analyst, architect, assistant, code-reviewer, coder, customer-support, data-scientist, debugger, devops-lead, doc-writer, email-assistant, health-tracker, hello-world, home-automation, legal-assistant, meeting-assistant, ops, orchestrator, personal-finance, planner, recipe-assistant, recruiter, researcher, sales-assistant, security-auditor, social-media, test-engineer, translator, travel-planner, tutor, writer ### Integrations -Integration templates in `integrations/.toml` define MCP server connections (GitHub, Slack, databases, etc.) with transport config, required env vars, and setup instructions. +Integration templates define [MCP](https://modelcontextprotocol.io/) server connections with transport configuration, required environment variables, and setup instructions. ```toml id = "github" @@ -96,9 +180,47 @@ name = "GITHUB_PERSONAL_ACCESS_TOKEN" is_secret = true ``` +**25 integrations across 6 categories:** + +| Category | Integrations | +|----------|-------------| +| DevTools | bitbucket, github, gitlab, jira, linear, sentry | +| Data | elasticsearch, mongodb, postgresql, redis, sqlite | +| Productivity | dropbox, gmail, google-calendar, google-drive, notion, todoist | +| Communication | discord, slack, teams | +| Cloud | aws, azure, gcp | +| AI Search | brave-search, exa-search | + +### Providers + +Provider files define LLM providers and their models with pricing, context windows, and capability flags. See [schema.toml](schema.toml) for the full field reference. + +**48 providers** including: Anthropic, OpenAI, Google Gemini, DeepSeek, Groq, Mistral, Cohere, xAI, Together, Fireworks, Ollama (local), LM Studio (local), vLLM (self-hosted), and many more. + +**223 models** with metadata for each: pricing (input/output per token), context window size, capability flags (vision, function calling, streaming), and tier classification. + +### Aliases + +Global model alias mappings in [aliases.toml](aliases.toml) let users reference models by short names: + +```toml +"sonnet" = "claude-sonnet-4-6" +"gpt4" = "gpt-4o" +"flash" = "gemini-2.5-flash" +"deepseek" = "deepseek-chat" +``` + +Models can also define aliases directly in their provider TOML files, which are auto-registered at load time. + +### Plugins + +Plugins extend agent capabilities with memory systems, safety guardrails, and conversation utilities. + +**10 plugins:** auto-summarizer, context-decay, conversation-logger, episodic-memory, guardrails, keyword-memory, sentiment-tracker, todo-tracker, topic-memory, user-profile + ### Skills -Skills in `skills//skill.toml` are reusable prompt templates or Python scripts that agents can invoke. +Reusable prompt templates or Python scripts that agents can invoke. ```toml [skill] @@ -112,13 +234,28 @@ type = "promptonly" template = "Create a meeting agenda for: {{topic}}" ``` -### Providers +### Workflows -Provider files in `providers/.toml` define LLM providers and their models with pricing, context windows, and capability flags. See [schema.toml](schema.toml) for the full field reference. +Pre-built multi-agent workflow definitions in `workflows/.toml` orchestrate multiple agents for complex tasks. -## How LibreFang Uses This Registry +**9 workflows:** brainstorm, code-review, content-pipeline, content-review, customer-support, data-pipeline, research, translate-polish, weekly-report -LibreFang ships with built-in content compiled into the binary. This repository serves as the upstream source for updates and community contributions. +### Templates + +Starter templates in `templates/` for creating new content. Copy a template to get started quickly: + +```bash +cp templates/agent.toml agents/my-agent/agent.toml +cp templates/HAND.toml hands/my-hand/HAND.toml +``` + +**6 templates:** agent.toml, HAND.toml, integration.toml, plugin.toml, provider.toml, skill.toml + +See also [docs/content-guide.md](docs/content-guide.md) for naming conventions and contribution guidelines. + +## Usage + +### Install from Registry ```bash # Update all registry content @@ -133,15 +270,15 @@ librefang integration install github ### Custom Local Content -You can also create custom content locally without submitting a PR: +Create custom content locally without submitting to this registry: ```bash -# Create a custom agent +# Custom agent mkdir -p ~/.librefang/agents/my-agent # Edit ~/.librefang/agents/my-agent/agent.toml -# Add custom models to your config -# ~/.librefang/model_catalog.toml +# Custom model aliases +# Add to ~/.librefang/model_catalog.toml ``` ## Validation @@ -150,9 +287,9 @@ mkdir -p ~/.librefang/agents/my-agent python scripts/validate.py ``` -This validates all provider TOML files for correctness: required fields, valid tiers, non-negative costs, no duplicate IDs. +Validates all content files for correctness: required fields, valid types, non-negative costs, no duplicate IDs. -## How to Contribute +## Contributing 1. Fork this repository 2. Add or edit content in the appropriate directory @@ -161,19 +298,6 @@ This validates all provider TOML files for correctness: required fields, valid t See [CONTRIBUTING.md](CONTRIBUTING.md) for detailed instructions for each content type. -## Current Stats - -| Type | Count | -|------|-------| -| Agents | 33 | -| Hands | 14 | -| Integrations | 25 | -| Skills | 2 | -| Plugins | 10 | -| Providers | 46 | -| Models | 220+ | -| Aliases | 80+ | - ## License MIT License. See [LICENSE](LICENSE). diff --git a/hands/README.md b/hands/README.md index 8802579..6ca99e0 100644 --- a/hands/README.md +++ b/hands/README.md @@ -10,7 +10,7 @@ Hand definitions for LibreFang. Hands are the user-facing "apps" -- higher-level hands/ ├── browser/ │ ├── HAND.toml # Hand definition -│ └── SKILL.md # Documentation +│ └── SKILL.md # Expert knowledge for the agent ├── trader/ │ ├── HAND.toml │ └── SKILL.md @@ -47,6 +47,17 @@ name = "hand-agent" module = "builtin:chat" system_prompt = """...""" +# Optional: i18n for name, description, category, and settings +# Supported languages: zh, ja, ko, es, fr, de +[i18n.zh] +name = "浏览器 Hand" +description = "自主网页浏览器" +category = "生产力" + +[i18n.zh.settings.headless] # Per-setting label/description translation +label = "无头模式" +description = "在后台运行浏览器" + [dashboard] # Dashboard metrics [[dashboard.metrics]] label = "Tasks Completed" @@ -58,15 +69,24 @@ format = "number" | Hand | Category | Description | |------|----------|-------------| -| browser | productivity | Autonomous web browser | -| trader | data | Crypto/stock trading assistant | -| researcher | productivity | Deep research automation | -| analytics | data | Data analysis and dashboards | -| ... | | See each directory for details | +| analytics | data | Data analytics, visualization, dashboards, and automated reporting | +| apitester | development | API testing, endpoint discovery, load testing, and regression detection | +| browser | productivity | Web navigation, form filling, and multi-step web tasks | +| clip | content | Turns long-form video into short clips with captions and thumbnails | +| collector | data | Intelligence collection, change detection, and knowledge graphs | +| devops | development | CI/CD management, infrastructure monitoring, and incident response | +| lead | data | Lead generation, enrichment, scoring, and scheduled delivery | +| linkedin | communication | LinkedIn content creation, networking, and engagement | +| predictor | data | Signal collection, calibrated predictions, and accuracy tracking | +| reddit | communication | Subreddit monitoring, content posting, and engagement tracking | +| researcher | productivity | Deep research, cross-referencing, fact-checking, and reports | +| strategist | productivity | Market research, competitive analysis, and strategic planning | +| trader | data | Market intelligence, multi-signal analysis, and risk management | +| twitter | communication | Twitter/X content creation, scheduling, and performance tracking | ## Adding a New Hand -1. Create `hands//HAND.toml` (and optionally `SKILL.md`) +1. Create `hands//HAND.toml` and `SKILL.md` (expert knowledge for the agent) 2. Ensure `id` matches the directory name 3. Run `python scripts/validate.py` 4. Submit a PR diff --git a/hands/analytics/HAND.toml b/hands/analytics/HAND.toml index 1d35f06..169c41e 100644 --- a/hands/analytics/HAND.toml +++ b/hands/analytics/HAND.toml @@ -484,6 +484,217 @@ token_consumption = "high" default_active = true # Note: High consumption when actively analyzing data, lower when idle +# ─── Internationalization (optional) ───────────────────────────────────────── +# All i18n sections are optional. Without them, the English values above are used. +# To localize, add [i18n.LANG] sections (e.g. zh, ja, ko, es, fr, de). +# Settings translations are also optional — omit to keep English labels. + +# ─── Chinese (简体中文) ──────────────────────────────────────────────────── + [i18n.zh] name = "数据分析 Hand" description = "自主数据分析智能体——数据采集、分析、可视化、仪表盘和自动化报告" +category = "数据" + +[i18n.zh.settings.data_source] +label = "数据源" +description = "主要数据源类型" + +[i18n.zh.settings.analysis_type] +label = "分析类型" +description = "默认分析方法" + +[i18n.zh.settings.output_format] +label = "输出格式" +description = "分析结果的呈现方式" + +[i18n.zh.settings.visualization] +label = "可视化" +description = "生成图表和可视化内容" + +[i18n.zh.settings.auto_schedule] +label = "定时报告" +description = "按计划自动生成报告" + +[i18n.zh.settings.report_frequency] +label = "报告频率" +description = "定时报告的生成频率" + +[i18n.zh.settings.confidence_threshold] +label = "置信度阈值" +description = "报告中纳入分析结论的最低置信度要求" + +# ─── Japanese (日本語) ──────────────────────────────────────────────────── + +[i18n.ja] +name = "データ分析 Hand" +description = "自律型データ分析エージェント——データ収集、分析、可視化、ダッシュボード、自動レポート生成" +category = "データ" + +[i18n.ja.settings.data_source] +label = "データソース" +description = "主要なデータソースの種類" + +[i18n.ja.settings.analysis_type] +label = "分析タイプ" +description = "デフォルトの分析アプローチ" + +[i18n.ja.settings.output_format] +label = "出力形式" +description = "分析結果の表示方法" + +[i18n.ja.settings.visualization] +label = "可視化" +description = "チャートやビジュアライゼーションを生成する" + +[i18n.ja.settings.auto_schedule] +label = "定期レポート" +description = "スケジュールに基づいてレポートを自動生成する" + +[i18n.ja.settings.report_frequency] +label = "レポート頻度" +description = "定期レポートの生成頻度" + +[i18n.ja.settings.confidence_threshold] +label = "信頼度しきい値" +description = "レポートに分析結果を含めるための最低信頼度" + +# ─── Spanish (Español) ──────────────────────────────────────────────────── + +[i18n.es] +name = "Hand de Analítica" +description = "Agente autónomo de analítica de datos — recopilación, análisis, visualización, paneles de control e informes automatizados" +category = "Datos" + +[i18n.es.settings.data_source] +label = "Fuente de datos" +description = "Tipo de fuente de datos principal" + +[i18n.es.settings.analysis_type] +label = "Tipo de análisis" +description = "Enfoque de análisis predeterminado" + +[i18n.es.settings.output_format] +label = "Formato de salida" +description = "Cómo presentar los resultados del análisis" + +[i18n.es.settings.visualization] +label = "Visualización" +description = "Generar gráficos y visualizaciones" + +[i18n.es.settings.auto_schedule] +label = "Informes programados" +description = "Generar informes automáticamente según un calendario" + +[i18n.es.settings.report_frequency] +label = "Frecuencia de informes" +description = "Con qué frecuencia generar los informes programados" + +[i18n.es.settings.confidence_threshold] +label = "Umbral de confianza" +description = "Nivel mínimo de confianza para incluir hallazgos en los informes" + +# ─── French (Français) ──────────────────────────────────────────────────── + +[i18n.fr] +name = "Hand Analytique" +description = "Agent autonome d'analyse de données — collecte, analyse, visualisation, tableaux de bord et rapports automatisés" +category = "Données" + +[i18n.fr.settings.data_source] +label = "Source de données" +description = "Type de source de données principal" + +[i18n.fr.settings.analysis_type] +label = "Type d'analyse" +description = "Approche d'analyse par défaut" + +[i18n.fr.settings.output_format] +label = "Format de sortie" +description = "Mode de présentation des résultats d'analyse" + +[i18n.fr.settings.visualization] +label = "Visualisation" +description = "Générer des graphiques et des visualisations" + +[i18n.fr.settings.auto_schedule] +label = "Rapports programmés" +description = "Générer automatiquement des rapports selon un calendrier" + +[i18n.fr.settings.report_frequency] +label = "Fréquence des rapports" +description = "Fréquence de génération des rapports programmés" + +[i18n.fr.settings.confidence_threshold] +label = "Seuil de confiance" +description = "Niveau de confiance minimum pour inclure les résultats dans les rapports" + +# ─── German (Deutsch) ──────────────────────────────────────────────────── + +[i18n.de] +name = "Analytik-Hand" +description = "Autonomer Datenanalyse-Agent — Datenerfassung, Analyse, Visualisierung, Dashboards und automatisierte Berichte" +category = "Daten" + +[i18n.de.settings.data_source] +label = "Datenquelle" +description = "Primärer Datenquellentyp" + +[i18n.de.settings.analysis_type] +label = "Analysetyp" +description = "Standard-Analyseansatz" + +[i18n.de.settings.output_format] +label = "Ausgabeformat" +description = "Darstellung der Analyseergebnisse" + +[i18n.de.settings.visualization] +label = "Visualisierung" +description = "Diagramme und Visualisierungen generieren" + +[i18n.de.settings.auto_schedule] +label = "Geplante Berichte" +description = "Berichte automatisch nach Zeitplan generieren" + +[i18n.de.settings.report_frequency] +label = "Berichtshäufigkeit" +description = "Häufigkeit der geplanten Berichtserstellung" + +[i18n.de.settings.confidence_threshold] +label = "Konfidenzschwelle" +description = "Mindest-Konfidenzniveau für die Aufnahme von Ergebnissen in Berichte" + +# ─── Korean (한국어) ──────────────────────────────────────────────────── + +[i18n.ko] +name = "데이터 분석 Hand" +description = "자율 데이터 분석 에이전트 — 데이터 수집, 분석, 시각화, 대시보드 및 자동화 보고서" +category = "데이터" + +[i18n.ko.settings.data_source] +label = "데이터 소스" +description = "주요 데이터 소스 유형" + +[i18n.ko.settings.analysis_type] +label = "분석 유형" +description = "기본 분석 방법" + +[i18n.ko.settings.output_format] +label = "출력 형식" +description = "분석 결과 표시 방식" + +[i18n.ko.settings.visualization] +label = "시각화" +description = "차트 및 시각화 콘텐츠 생성" + +[i18n.ko.settings.auto_schedule] +label = "정기 보고서" +description = "일정에 따라 자동으로 보고서 생성" + +[i18n.ko.settings.report_frequency] +label = "보고서 빈도" +description = "정기 보고서 생성 주기" + +[i18n.ko.settings.confidence_threshold] +label = "신뢰도 임계값" +description = "보고서에 분석 결과를 포함하기 위한 최소 신뢰도 수준" diff --git a/hands/analytics/SKILL.md b/hands/analytics/SKILL.md index 16a404d..264b35f 100644 --- a/hands/analytics/SKILL.md +++ b/hands/analytics/SKILL.md @@ -337,3 +337,702 @@ Level 4: What to do (prescriptive) | Timeliness | Current | Data refreshed daily | | Uniqueness | 99% | 1% duplicate records found | ``` + +--- + +## Worked Examples + +### Example 1: E-commerce Sales Analysis + +**Goal**: Analyze 12 months of order data to identify revenue drivers, customer segments, and growth trends. + +#### Step 1 — Load and clean +```python +import pandas as pd +import numpy as np + +df = pd.read_csv('orders.csv', parse_dates=['order_date']) + +# Quick audit +print(f"Rows: {len(df):,} Columns: {df.shape[1]}") +print(df.isnull().sum()[df.isnull().sum() > 0]) + +# Clean +df = df.dropna(subset=['customer_id', 'order_total']) +df['order_total'] = df['order_total'].clip(lower=0) # Remove negative values +df['order_month'] = df['order_date'].dt.to_period('M') +``` + +#### Step 2 — Revenue trend analysis +```python +monthly = ( + df.groupby('order_month') + .agg(revenue=('order_total', 'sum'), + orders=('order_id', 'nunique'), + customers=('customer_id', 'nunique')) + .reset_index() +) +monthly['aov'] = monthly['revenue'] / monthly['orders'] # Average order value +monthly['revenue_mom'] = monthly['revenue'].pct_change() # Month-over-month growth + +fig, axes = plt.subplots(2, 1, figsize=(12, 8), sharex=True) +axes[0].bar(monthly['order_month'].astype(str), monthly['revenue'], color='steelblue') +axes[0].set_title('Monthly Revenue', fontsize=14, fontweight='bold') +axes[0].set_ylabel('Revenue ($)') + +axes[1].plot(monthly['order_month'].astype(str), monthly['aov'], marker='o', color='coral') +axes[1].set_title('Average Order Value', fontsize=14, fontweight='bold') +axes[1].set_ylabel('AOV ($)') +plt.xticks(rotation=45, ha='right') +plt.tight_layout() +plt.savefig('revenue_trend.png', dpi=150, bbox_inches='tight') +plt.close() +``` + +#### Step 3 — Customer segmentation (RFM) +```python +snapshot_date = df['order_date'].max() + pd.Timedelta(days=1) + +rfm = df.groupby('customer_id').agg( + recency=('order_date', lambda x: (snapshot_date - x.max()).days), + frequency=('order_id', 'nunique'), + monetary=('order_total', 'sum') +) + +# Score each dimension 1-4 using quartiles +for col in ['recency', 'frequency', 'monetary']: + labels = [4, 3, 2, 1] if col == 'recency' else [1, 2, 3, 4] + rfm[f'{col}_score'] = pd.qcut(rfm[col], q=4, labels=labels, duplicates='drop') + +rfm['rfm_score'] = (rfm['recency_score'].astype(int) + + rfm['frequency_score'].astype(int) + + rfm['monetary_score'].astype(int)) + +# Segment mapping +def segment(row): + r, f, m = int(row['recency_score']), int(row['frequency_score']), int(row['monetary_score']) + if r >= 3 and f >= 3: + return 'Champions' + elif r >= 3 and f < 3: + return 'New / Promising' + elif r < 3 and f >= 3: + return 'At Risk' + else: + return 'Needs Attention' + +rfm['segment'] = rfm.apply(segment, axis=1) +print(rfm.groupby('segment').agg( + count=('monetary', 'size'), + avg_revenue=('monetary', 'mean'), + avg_frequency=('frequency', 'mean') +).sort_values('avg_revenue', ascending=False)) +``` + +#### Step 4 — Cohort retention analysis +```python +df['cohort'] = df.groupby('customer_id')['order_date'].transform('min').dt.to_period('M') +df['order_period'] = df['order_date'].dt.to_period('M') +df['cohort_index'] = (df['order_period'] - df['cohort']).apply(lambda x: x.n) + +cohort_table = ( + df.groupby(['cohort', 'cohort_index'])['customer_id'] + .nunique() + .reset_index() + .pivot(index='cohort', columns='cohort_index', values='customer_id') +) + +# Convert to retention percentages +retention = cohort_table.div(cohort_table[0], axis=0) * 100 + +fig, ax = plt.subplots(figsize=(14, 8)) +sns.heatmap(retention, annot=True, fmt='.0f', cmap='YlOrRd_r', ax=ax) +ax.set_title('Cohort Retention (% of original customers)', fontsize=14, fontweight='bold') +ax.set_xlabel('Months Since First Purchase') +ax.set_ylabel('Cohort') +plt.tight_layout() +plt.savefig('cohort_retention.png', dpi=150, bbox_inches='tight') +plt.close() +``` + +--- + +### Example 2: A/B Test Analysis + +**Goal**: Evaluate whether a new checkout flow (variant B) improves conversion rate over the existing flow (variant A). + +#### Step 1 — Sample size calculation (pre-test) +```python +from scipy import stats +import numpy as np + +baseline_rate = 0.12 # Current conversion rate: 12% +mde = 0.02 # Minimum detectable effect: 2 percentage points +alpha = 0.05 # Significance level +power = 0.80 # Statistical power + +# Using the normal approximation formula +p1 = baseline_rate +p2 = baseline_rate + mde +p_avg = (p1 + p2) / 2 + +z_alpha = stats.norm.ppf(1 - alpha / 2) # Two-tailed +z_beta = stats.norm.ppf(power) + +n_per_group = ((z_alpha * np.sqrt(2 * p_avg * (1 - p_avg)) + + z_beta * np.sqrt(p1 * (1 - p1) + p2 * (1 - p2))) ** 2 + / (p2 - p1) ** 2) + +print(f"Required sample size per group: {int(np.ceil(n_per_group)):,}") +print(f"Total required: {int(np.ceil(n_per_group)) * 2:,}") +``` + +#### Step 2 — Run the test and collect results +```python +ab = pd.read_csv('ab_test_results.csv') + +summary = ab.groupby('variant').agg( + visitors=('user_id', 'nunique'), + conversions=('converted', 'sum') +) +summary['conversion_rate'] = summary['conversions'] / summary['visitors'] +print(summary) +``` + +#### Step 3 — Statistical significance +```python +a = ab[ab['variant'] == 'A'] +b = ab[ab['variant'] == 'B'] + +# Chi-squared test for proportions +contingency = pd.crosstab(ab['variant'], ab['converted']) +chi2, p_value, dof, expected = stats.chi2_contingency(contingency) + +# Proportions z-test (more direct) +from statsmodels.stats.proportion import proportions_ztest +successes = [summary.loc['B', 'conversions'], summary.loc['A', 'conversions']] +trials = [summary.loc['B', 'visitors'], summary.loc['A', 'visitors']] +z_stat, p_val = proportions_ztest(successes, trials, alternative='larger') + +print(f"Z-statistic: {z_stat:.4f}") +print(f"P-value: {p_val:.4f}") +print(f"Significant: {'Yes' if p_val < 0.05 else 'No'} (at alpha=0.05)") +``` + +#### Step 4 — Effect size and confidence interval +```python +p_a = summary.loc['A', 'conversion_rate'] +p_b = summary.loc['B', 'conversion_rate'] +n_a = summary.loc['A', 'visitors'] +n_b = summary.loc['B', 'visitors'] + +lift = (p_b - p_a) / p_a +se_diff = np.sqrt(p_a * (1 - p_a) / n_a + p_b * (1 - p_b) / n_b) +ci_lower = (p_b - p_a) - 1.96 * se_diff +ci_upper = (p_b - p_a) + 1.96 * se_diff + +print(f"Control rate: {p_a:.4f}") +print(f"Variant rate: {p_b:.4f}") +print(f"Absolute lift: {p_b - p_a:+.4f}") +print(f"Relative lift: {lift:+.2%}") +print(f"95% CI for diff: [{ci_lower:+.4f}, {ci_upper:+.4f}]") +``` + +#### Step 5 — Recommendation template +``` +## A/B Test Report: New Checkout Flow + +| Metric | Control (A) | Variant (B) | +|---------------------|-------------|-------------| +| Visitors | 15,204 | 15,198 | +| Conversions | 1,824 | 2,127 | +| Conversion Rate | 12.00% | 13.99% | + +**Result**: Statistically significant (p = 0.0003, alpha = 0.05) +**Lift**: +1.99pp absolute / +16.6% relative +**95% CI**: [+0.90pp, +3.08pp] +**Recommendation**: Deploy variant B. The effect is both statistically +and practically significant with a lower bound above the +1pp threshold. +``` + +--- + +### Example 3: Customer Churn Analysis + +**Goal**: Identify which factors most strongly predict customer churn and quantify their relative importance. + +#### Step 1 — Feature engineering +```python +df = pd.read_csv('customers.csv') + +# Create behavioral features from raw data +features = df.copy() +features['tenure_months'] = (pd.Timestamp.now() - pd.to_datetime(df['signup_date'])).dt.days / 30 +features['support_tickets_per_month'] = df['total_tickets'] / features['tenure_months'].clip(lower=1) +features['avg_session_minutes'] = df['total_session_minutes'] / df['total_sessions'].clip(lower=1) +features['days_since_last_login'] = (pd.Timestamp.now() - pd.to_datetime(df['last_login'])).dt.days +features['has_premium'] = (df['plan'] == 'premium').astype(int) + +# Drop raw columns, keep engineered features +feature_cols = [ + 'tenure_months', 'support_tickets_per_month', 'avg_session_minutes', + 'days_since_last_login', 'has_premium', 'monthly_spend', 'num_features_used' +] +``` + +#### Step 2 — Correlation analysis +```python +churn_corr = features[feature_cols + ['churned']].corr()['churned'].drop('churned').sort_values() + +fig, ax = plt.subplots(figsize=(8, 5)) +churn_corr.plot(kind='barh', ax=ax, color=['coral' if x > 0 else 'steelblue' for x in churn_corr]) +ax.set_title('Feature Correlation with Churn', fontsize=14, fontweight='bold') +ax.set_xlabel('Pearson Correlation') +ax.axvline(x=0, color='black', linewidth=0.5) +plt.tight_layout() +plt.savefig('churn_correlations.png', dpi=150, bbox_inches='tight') +plt.close() +``` + +#### Step 3 — Key driver identification via group comparison +```python +churned = features[features['churned'] == 1] +retained = features[features['churned'] == 0] + +comparison = [] +for col in feature_cols: + t_stat, p_val = stats.ttest_ind(churned[col].dropna(), retained[col].dropna()) + d = cohens_d(churned[col].dropna(), retained[col].dropna()) # From earlier definition + comparison.append({ + 'feature': col, + 'churned_mean': churned[col].mean(), + 'retained_mean': retained[col].mean(), + 'diff_pct': (churned[col].mean() - retained[col].mean()) / retained[col].mean() * 100, + 'cohens_d': abs(d), + 'p_value': p_val, + 'significant': p_val < 0.05 + }) + +result = pd.DataFrame(comparison).sort_values('cohens_d', ascending=False) +print(result.to_string(index=False)) +``` + +#### Step 4 — Interpret and report +``` +## Churn Driver Analysis + +**Top 3 factors distinguishing churned vs. retained customers:** + +| Factor | Churned (avg) | Retained (avg) | Diff | Effect Size | +|----------------------------|---------------|----------------|----------|-------------| +| Days since last login | 34.2 | 8.7 | +293% | Large | +| Support tickets per month | 2.8 | 0.9 | +211% | Large | +| Number of features used | 3.1 | 7.4 | -58% | Medium | + +**Actionable insights:** +1. Customers inactive >14 days are 4x more likely to churn -- trigger re-engagement email at day 10 +2. High support ticket rate signals frustration -- escalate accounts with >2 tickets/month to success team +3. Low feature adoption correlates with churn -- implement onboarding flow targeting unused features +``` + +--- + +## Advanced pandas Patterns + +### Window Functions + +```python +# Expanding window (cumulative statistics) +df['cumulative_avg'] = df['value'].expanding().mean() +df['cumulative_max'] = df['value'].expanding().max() + +# Exponentially weighted moving average (EWMA) -- emphasizes recent values +df['ewma_7'] = df['value'].ewm(span=7).mean() # Span-based decay +df['ewma_a'] = df['value'].ewm(alpha=0.3).mean() # Explicit decay factor + +# Comparison: rolling vs. EWMA +# - rolling(7).mean() weights all 7 values equally +# - ewm(span=7).mean() weights recent values exponentially more +# Use EWMA when recent data matters more (stock prices, real-time metrics) + +# Rolling with min_periods (handles early rows with insufficient data) +df['rolling_avg'] = df['value'].rolling(window=30, min_periods=5).mean() + +# Rolling rank (percentile within window) +df['rolling_pctile'] = df['value'].rolling(90).rank(pct=True) +``` + +### Multi-Index Operations + +```python +# Create multi-index from groupby +multi = df.groupby(['region', 'product']).agg( + revenue=('amount', 'sum'), + units=('quantity', 'sum') +) + +# Access levels +multi.loc['North'] # All products in North region +multi.loc[('North', 'Widget')] # Specific region + product +multi.xs('Widget', level='product') # All regions for Widget + +# Swap and sort levels +multi = multi.swaplevel().sort_index() + +# Reset to flat columns +flat = multi.reset_index() + +# Stack / unstack (reshape between long and wide) +wide = multi['revenue'].unstack(level='product') # Products become columns +long = wide.stack() # Back to multi-index +``` + +### Merge and Join Patterns + +```python +# Inner join (only matching rows) +merged = orders.merge(customers, on='customer_id', how='inner') + +# Left join with indicator (see which rows matched) +merged = orders.merge(customers, on='customer_id', how='left', indicator=True) +unmatched = merged[merged['_merge'] == 'left_only'] + +# Join on multiple keys +merged = df1.merge(df2, on=['date', 'region'], how='left') + +# Join with different column names +merged = orders.merge(products, left_on='prod_id', right_on='product_id') + +# Anti-join (rows in A that have no match in B) +anti = df_a.merge(df_b, on='key', how='left', indicator=True) +anti = anti[anti['_merge'] == 'left_only'].drop(columns='_merge') + +# Self-join (compare rows within same table) +df_prev = df[['customer_id', 'order_date', 'amount']].rename( + columns={'order_date': 'prev_date', 'amount': 'prev_amount'} +) +df_with_prev = df.merge(df_prev, on='customer_id', how='left') +df_with_prev = df_with_prev[df_with_prev['prev_date'] < df_with_prev['order_date']] +``` + +### Apply and Transform + +```python +# transform() returns same-shaped output -- useful for group-level stats on each row +df['group_mean'] = df.groupby('category')['value'].transform('mean') +df['pct_of_group'] = df['value'] / df.groupby('category')['value'].transform('sum') +df['z_within_group'] = df.groupby('category')['value'].transform( + lambda x: (x - x.mean()) / x.std() +) + +# apply() for multi-column group operations +def top_n(group, n=3): + return group.nlargest(n, 'value') + +top3_per_category = df.groupby('category', group_keys=False).apply(top_n, n=3) + +# Vectorized operations (prefer these over apply when possible) +# Slow: +df['result'] = df.apply(lambda row: row['a'] * row['b'] + row['c'], axis=1) +# Fast: +df['result'] = df['a'] * df['b'] + df['c'] + +# np.where for conditional columns (vectorized if/else) +df['tier'] = np.where(df['revenue'] > 10000, 'high', 'low') + +# np.select for multiple conditions +conditions = [ + df['revenue'] > 10000, + df['revenue'] > 5000, + df['revenue'] > 0, +] +choices = ['high', 'medium', 'low'] +df['tier'] = np.select(conditions, choices, default='none') +``` + +### Memory Optimization for Large Datasets + +```python +# Check current memory usage +print(df.memory_usage(deep=True).sum() / 1024**2, "MB") + +# Downcast numeric types +df['int_col'] = pd.to_numeric(df['int_col'], downcast='integer') # int64 -> int8/16/32 +df['float_col'] = pd.to_numeric(df['float_col'], downcast='float') # float64 -> float32 + +# Use category type for low-cardinality strings +for col in df.select_dtypes(include='object'): + if df[col].nunique() / len(df) < 0.5: # Less than 50% unique values + df[col] = df[col].astype('category') + +# Read in chunks for files that exceed memory +chunks = pd.read_csv('huge_file.csv', chunksize=100_000) +results = [] +for chunk in chunks: + processed = chunk.groupby('category')['value'].sum() + results.append(processed) +final = pd.concat(results).groupby(level=0).sum() + +# Specify dtypes at load time (avoids loading as float64/object first) +dtypes = { + 'id': 'int32', + 'category': 'category', + 'value': 'float32', + 'flag': 'bool' +} +df = pd.read_csv('data.csv', dtype=dtypes) + +# Use pyarrow backend for better memory efficiency (pandas 2.0+) +df = pd.read_csv('data.csv', engine='pyarrow', dtype_backend='pyarrow') +``` + +--- + +## Dashboard and Reporting Patterns + +### Executive Dashboard Template + +```python +import matplotlib.pyplot as plt +import matplotlib.gridspec as gridspec +from matplotlib.patches import FancyBboxPatch + +def executive_dashboard(kpis, trend_df, comparison_df, output='dashboard.png'): + """ + kpis: dict with keys like {'Revenue': '$1.2M', 'Growth': '+15%', ...} + trend_df: DataFrame with 'date' and 'value' columns + comparison_df: DataFrame with 'category' and 'current'/'previous' columns + """ + fig = plt.figure(figsize=(16, 10)) + gs = gridspec.GridSpec(3, len(kpis), hspace=0.4, wspace=0.3) + + # Row 1: KPI cards + for i, (label, value) in enumerate(kpis.items()): + ax = fig.add_subplot(gs[0, i]) + ax.text(0.5, 0.6, value, ha='center', va='center', + fontsize=28, fontweight='bold', color='#2c3e50') + ax.text(0.5, 0.2, label, ha='center', va='center', + fontsize=12, color='#7f8c8d') + ax.set_xlim(0, 1) + ax.set_ylim(0, 1) + ax.axis('off') + # Card background + rect = FancyBboxPatch((0.05, 0.05), 0.9, 0.9, boxstyle="round,pad=0.05", + facecolor='#f8f9fa', edgecolor='#dee2e6') + ax.add_patch(rect) + + # Row 2: Trend line + ax_trend = fig.add_subplot(gs[1, :]) + ax_trend.plot(trend_df['date'], trend_df['value'], linewidth=2, color='steelblue') + ax_trend.fill_between(trend_df['date'], trend_df['value'], alpha=0.1, color='steelblue') + ax_trend.set_title('Trend Over Time', fontsize=13, fontweight='bold') + ax_trend.set_ylabel('Value') + + # Row 3: Period comparison (grouped bar) + ax_comp = fig.add_subplot(gs[2, :]) + x = range(len(comparison_df)) + width = 0.35 + ax_comp.bar([i - width/2 for i in x], comparison_df['previous'], width, + label='Previous', color='#bdc3c7') + ax_comp.bar([i + width/2 for i in x], comparison_df['current'], width, + label='Current', color='steelblue') + ax_comp.set_xticks(list(x)) + ax_comp.set_xticklabels(comparison_df['category'], rotation=45, ha='right') + ax_comp.set_title('Current vs. Previous Period', fontsize=13, fontweight='bold') + ax_comp.legend() + + plt.savefig(output, dpi=150, bbox_inches='tight', facecolor='white') + plt.close() +``` + +### Weekly Metrics Report Template + +```python +def weekly_report(df, date_col='date', metric_col='value', group_col=None): + """Generate a standard weekly metrics summary.""" + df[date_col] = pd.to_datetime(df[date_col]) + df['week'] = df[date_col].dt.isocalendar().week.astype(int) + df['year'] = df[date_col].dt.year + + current_week = df['week'].max() + prev_week = current_week - 1 + + curr = df[df['week'] == current_week] + prev = df[df['week'] == prev_week] + + report = { + 'period': f"Week {current_week}", + 'total': curr[metric_col].sum(), + 'mean': curr[metric_col].mean(), + 'median': curr[metric_col].median(), + 'wow_change': (curr[metric_col].sum() - prev[metric_col].sum()) + / prev[metric_col].sum() * 100 + if prev[metric_col].sum() != 0 else None, + } + + if group_col: + report['by_group'] = curr.groupby(group_col)[metric_col].agg(['sum', 'mean', 'count']) + + # Sparkline trend (last 8 weeks) + weekly_totals = ( + df.groupby('week')[metric_col].sum() + .tail(8) + .reset_index() + ) + + fig, ax = plt.subplots(figsize=(6, 2)) + ax.plot(weekly_totals['week'], weekly_totals[metric_col], marker='o', + linewidth=2, color='steelblue', markersize=4) + ax.fill_between(weekly_totals['week'], weekly_totals[metric_col], + alpha=0.1, color='steelblue') + ax.set_title(f'{metric_col.title()} — Last 8 Weeks', fontsize=10) + ax.tick_params(labelsize=8) + plt.tight_layout() + plt.savefig('weekly_sparkline.png', dpi=150, bbox_inches='tight') + plt.close() + + return report +``` + +### Anomaly Detection Patterns + +```python +def detect_anomalies(series, method='zscore', threshold=3.0, window=30): + """ + Detect anomalies in a numeric series. + + Methods: + - 'zscore': Flag values beyond `threshold` standard deviations from mean + - 'iqr': Flag values beyond 1.5x IQR from quartiles + - 'rolling': Flag values beyond `threshold` std devs from rolling mean + """ + anomalies = pd.Series(False, index=series.index) + + if method == 'zscore': + z = (series - series.mean()) / series.std() + anomalies = z.abs() > threshold + + elif method == 'iqr': + q1 = series.quantile(0.25) + q3 = series.quantile(0.75) + iqr = q3 - q1 + anomalies = (series < q1 - 1.5 * iqr) | (series > q3 + 1.5 * iqr) + + elif method == 'rolling': + rolling_mean = series.rolling(window, min_periods=5).mean() + rolling_std = series.rolling(window, min_periods=5).std() + anomalies = (series - rolling_mean).abs() > threshold * rolling_std + + return anomalies + + +# Usage: detect and visualize +anomalies = detect_anomalies(df['metric'], method='rolling', threshold=2.5, window=30) + +fig, ax = plt.subplots(figsize=(14, 5)) +ax.plot(df.index, df['metric'], linewidth=1, color='steelblue', label='Metric') +ax.scatter(df.index[anomalies], df['metric'][anomalies], + color='red', s=40, zorder=5, label='Anomaly') +ax.legend() +ax.set_title('Anomaly Detection (Rolling Z-Score)', fontsize=14, fontweight='bold') +plt.tight_layout() +plt.savefig('anomalies.png', dpi=150, bbox_inches='tight') +plt.close() + +print(f"Detected {anomalies.sum()} anomalies out of {len(series):,} data points") +``` + +**Method selection guide:** + +| Method | Best For | Assumptions | Sensitivity | +|--------|----------|-------------|-------------| +| Z-score | Stationary data with normal distribution | Constant mean and variance | Low (misses local anomalies) | +| IQR | Skewed distributions, outlier screening | None (non-parametric) | Medium | +| Rolling z-score | Time series with trends or seasonality | Local stationarity within window | High (adapts to drift) | + +--- + +## Common Analytics Pitfalls + +### Simpson's Paradox + +A trend that appears in grouped data reverses when the groups are combined. + +``` +Department A: Drug works better (80% vs 70%) +Department B: Drug works better (50% vs 40%) +Combined: Drug appears WORSE (55% vs 60%) <-- paradox +``` + +**Why it happens**: Unequal group sizes create a confounding effect. Department B (with lower overall rates) sent most patients to the drug group. + +**Prevention**: Always segment data by relevant confounders before drawing conclusions. If aggregate and segmented results disagree, trust the segmented analysis and report the confounding variable. + +### Survivorship Bias + +Analyzing only entities that "survived" a selection process, ignoring those that dropped out. + +**Classic examples:** +- Studying only successful companies to find success patterns (ignoring failed companies with the same patterns) +- Analyzing only current customers to understand satisfaction (ignoring those who already left) +- Looking at fund performance by examining only funds that still exist (dead funds were closed) + +**Prevention**: Always ask "what is missing from this dataset?" before drawing conclusions. If possible, include data from non-survivors. Explicitly note the selection criteria and what it excludes. + +### Correlation vs. Causation + +A statistically significant correlation between X and Y does not mean X causes Y. Possible explanations: + +| Explanation | Example | +|-------------|---------| +| X causes Y | Exercise reduces blood pressure | +| Y causes X | Depression reduces exercise (not exercise causes depression) | +| Z causes both | Income drives both education spending AND health outcomes | +| Coincidence | Ice cream sales correlate with drowning deaths (both driven by summer) | + +**Prevention**: Establish causation only with randomized controlled experiments (A/B tests). For observational data, state findings as "associated with" not "causes." Look for confounders and test whether the relationship holds when controlling for them. + +### Cherry-Picking Time Windows + +Selecting a start/end date that makes a metric look better or worse than the true trend. + +```python +# Example: same data, different conclusions +# "Revenue up 40%!" -- comparing Jan (seasonal low) to Dec (seasonal high) +# "Revenue flat." -- comparing Dec 2024 to Dec 2025 (year-over-year) + +# Prevention: always use year-over-year comparison for seasonal data +df['yoy_change'] = df.groupby(df['date'].dt.month)['revenue'].pct_change(periods=12) +``` + +**Prevention checklist:** +- Compare like-for-like periods (YoY for seasonal businesses) +- Show the full time range, not a selected subset +- Use multiple time windows (WoW, MoM, QoQ, YoY) and note if they disagree +- Include a moving average to show the underlying trend separate from noise + +### Small Sample Size Issues + +Small samples produce unstable statistics that can flip with just a few more observations. + +```python +# Illustrate instability: conversion rates with small vs. large samples +from scipy.stats import beta + +# Scenario: 3 conversions out of 10 visitors (30%) +a_small, b_small = 3 + 1, 10 - 3 + 1 # Beta posterior +ci_small = beta.interval(0.95, a_small, b_small) +print(f"n=10: 30% conversion, 95% CI: [{ci_small[0]:.1%}, {ci_small[1]:.1%}]") +# Output: 95% CI: [9.9%, 56.8%] -- extremely wide, almost useless + +# Scenario: 300 conversions out of 1000 visitors (30%) +a_large, b_large = 300 + 1, 1000 - 300 + 1 +ci_large = beta.interval(0.95, a_large, b_large) +print(f"n=1000: 30% conversion, 95% CI: [{ci_large[0]:.1%}, {ci_large[1]:.1%}]") +# Output: 95% CI: [27.2%, 32.9%] -- narrow and actionable +``` + +**Rules of thumb:** +- n < 30: Do not draw firm conclusions. Report as directional only. +- Conversion rates need hundreds (not dozens) of conversions to stabilize. +- Always report confidence intervals alongside point estimates. +- If sample size is fixed and small, use exact tests (Fisher's exact) rather than approximations (chi-squared). diff --git a/hands/apitester/HAND.toml b/hands/apitester/HAND.toml index f295e5d..34a7d0c 100644 --- a/hands/apitester/HAND.toml +++ b/hands/apitester/HAND.toml @@ -302,9 +302,26 @@ If `approval_mode` is ENABLED: If `approval_mode` is DISABLED: Execute load tests directly. +### Structured Load Test Profiles + +Run profiles in order. Each answers a different question. Stop a profile early if exit criteria are met. + +**Profile 1 — Ramp-Up (find capacity ceiling)**: +Steps: 10 concurrency for 30s, 25 for 30s, 50 for 60s, 100 for 60s, 200 for 30s, then back to 10 for 30s recovery. +Exit: stop stepping up when error rate >10% or p95 >2s. Record last healthy step as "max safe concurrency." + +**Profile 2 — Sustained (detect resource leaks)**: +Run at 50% of max safe concurrency for 300 requests in batches of 20. Compare average response time of first quarter vs last quarter. A >25% increase signals connection pool exhaustion or memory growth. + +**Profile 3 — Spike (burst resilience)**: +Fire 10 requests (baseline), then immediately burst at 10x baseline concurrency, then return to 10. Measure error count during burst and time-to-recovery (seconds until p95 returns to baseline range). + +**Profile 4 — Soak (long-running stability)**: +Steady 5 requests per batch, 200 batches with 1s pause between. Track response time trend. Flag if final-quarter average exceeds first-quarter average by >30%. + Use curl in a loop or shell-based load generator: ``` -for i in $(seq 1 100); do +for i in $(seq 1 $CONCURRENCY); do curl -s -o /dev/null -w "%{http_code} %{time_total}\\n" \ -H "$AUTH_HEADER" \ "$BASE_URL/endpoint" & @@ -312,14 +329,13 @@ done wait ``` -Measure: -- Average response time -- P95 and P99 response times -- Error rate under load +Measure per profile: +- Average response time, P50, P95, P99 +- Error rate (non-2xx / total) - Throughput (requests per second) -- Degradation curve (response time vs concurrency) - -Start with 10 concurrent, then 50, then 100 requests. +- Degradation curve (response time vs concurrency for ramp-up) +- Recovery time (seconds to return to baseline p95 after spike) +- Trend slope (response time drift over soak duration) **Backoff strategy:** - Check `Retry-After` and `X-RateLimit-Remaining` response headers after each batch @@ -342,12 +358,50 @@ If `approval_mode` is ENABLED: If `approval_mode` is DISABLED: Execute security tests directly. -1. **Authentication tests**: Missing auth, invalid auth, expired tokens -2. **Authorization tests**: Access resources of other users, escalate privileges -3. **Input injection**: SQL injection, XSS, command injection in parameters -4. **Headers**: Missing security headers (CORS, HSTS, X-Frame-Options) -5. **Rate limiting**: Verify rate limits are enforced -6. **Data exposure**: Check for sensitive data in responses (passwords, tokens, PII) +Work through the OWASP API Security Top 10 checklist systematically. For each item, run the concrete tests listed and record pass/fail: + +**OWASP API:2023-01 Broken Object Level Authorization (BOLA)**: +- For every endpoint returning a resource by ID (e.g. `/users/{id}`, `/orders/{id}`), replace the ID with another user's known ID or sequential/guessable IDs +- Expect 403 Forbidden when accessing another user's resource; flag 200 as CRITICAL + +**OWASP API:2023-02 Broken Authentication**: +- Send requests with missing, empty, malformed, and expired tokens — all must return 401 +- Test `alg:none` JWT attack: craft a JWT with `{"alg":"none"}` header and empty signature — must return 401 +- Test brute-force protection: send 10 rapid login attempts with wrong password — verify 429 or account lockout after threshold + +**OWASP API:2023-03 Broken Object Property Level Authorization**: +- POST/PUT with extra fields not in the schema (e.g. `"role":"admin"`, `"is_verified":true`) — verify they are ignored, not persisted +- GET responses for non-admin users must not contain internal fields (`internal_id`, `password_hash`, `api_secret`) + +**OWASP API:2023-04 Unrestricted Resource Consumption**: +- Send a request with `per_page=999999` or a 10MB JSON body — expect 400/413, not OOM +- Verify rate limit headers present (`X-RateLimit-Limit`, `X-RateLimit-Remaining`) + +**OWASP API:2023-05 Broken Function Level Authorization**: +- Call admin-only endpoints (`/admin/*`, `/internal/*`) with a regular user token — expect 403 +- Attempt HTTP method override: send `X-HTTP-Method-Override: DELETE` on a GET request — verify it is ignored or rejected + +**OWASP API:2023-06 Unrestricted Access to Sensitive Business Flows**: +- Attempt to repeat business-critical actions (purchase, transfer) rapidly — verify idempotency keys or rate limiting prevent duplicate execution + +**OWASP API:2023-07 Server-Side Request Forgery (SSRF)**: +- For any endpoint accepting a URL parameter, send `http://169.254.169.254/latest/meta-data/` (cloud metadata) and `http://localhost:6379/` — expect rejection or error, not a proxied response + +**OWASP API:2023-08 Security Misconfiguration**: +- Check response headers: `Strict-Transport-Security`, `X-Content-Type-Options: nosniff`, `X-Frame-Options`, `Content-Security-Policy` +- Verify error responses do not leak stack traces, SQL queries, or internal paths +- Check that debug/docs endpoints (`/debug`, `/swagger`, `/graphql/playground`) return 404 or require auth in production + +**OWASP API:2023-09 Improper Inventory Management**: +- Probe old API versions (`/api/v1/`, `/api/v0/`) — they should be disabled or return 410 Gone +- Check for undocumented endpoints by testing common paths: `/api/internal`, `/api/debug`, `/metrics`, `/healthz` + +**OWASP API:2023-10 Unsafe Consumption of APIs**: +- If the API fetches external resources (image URLs, webhook callbacks), test with a URL returning malformed JSON, extremely large payloads, or slow responses (timeout >30s) — verify the API handles them gracefully without crashing + +Additionally test: +- **Input injection**: SQL (`' OR 1=1 --`), XSS (``), command injection (`; cat /etc/passwd`), path traversal (`../../etc/passwd`) in every string parameter +- **CORS**: Send `Origin: https://evil.example.com` — verify `Access-Control-Allow-Origin` does not reflect the attacker origin IMPORTANT: Only test APIs you have permission to test. Never perform destructive tests without explicit confirmation. @@ -361,6 +415,37 @@ Stop testing when ANY of these conditions is met: --- +## Phase 5.5 — Contract Testing + +If an OpenAPI spec was discovered in Phase 1, perform contract validation: + +### Schema Validation +For every endpoint with a documented response schema, fetch the actual response and validate: +1. All `required` fields are present +2. Every field matches its declared `type` and `format` (e.g. `string`/`date-time`, `integer`/`int64`) +3. `enum` fields contain only allowed values +4. `additionalProperties: false` schemas reject extra fields +5. Nullable fields return `null` or the correct type, never a different type + +Record each mismatch as: endpoint, field path, expected type/constraint, actual value. + +### Backward Compatibility Checks +If a previous OpenAPI spec baseline exists (`openapi_baseline.json`): +1. **Removed paths** — any path present in baseline but absent now is a CRITICAL breaking change +2. **Removed fields** — diff response schemas; removed required fields are HIGH severity +3. **Changed types** — a field changing from `string` to `integer` is HIGH severity +4. **New required request fields** — breaks existing callers, HIGH severity +5. **Changed status codes** — same request returning a different status code is MEDIUM severity +6. **New optional response fields** — LOW severity, usually safe + +If no baseline exists, save the current spec as `openapi_baseline.json` for future comparisons. + +### Content-Type Negotiation +- Send `Accept: application/xml` to a JSON-only endpoint — expect 406 Not Acceptable or graceful JSON fallback, not a 500 +- Send `Content-Type: text/plain` with a JSON body — expect 415 Unsupported Media Type + +--- + ## Phase 6 — Report Generation Generate a comprehensive test report: @@ -460,6 +545,265 @@ token_consumption = "medium" default_active = false activation_warning = "API Tester hand runs continuously, consuming tokens. Use on-demand for specific tests." +# ─── Internationalization (optional) ───────────────────────────────────────── +# All i18n sections are optional. Without them, the English values above are used. +# To localize, add [i18n.LANG] sections (e.g. zh, ja, ko, es, fr, de). +# Settings translations are also optional — omit to keep English labels. + +# ─── Chinese (简体中文) ──────────────────────────────────────────────────── + [i18n.zh] name = "API 测试 Hand" description = "自主 API 测试智能体——端点发现、请求验证、负载测试和回归检测" +category = "开发" + +[i18n.zh.settings.base_url] +label = "基础 URL" +description = "待测试 API 的基础 URL(例如 https://api.example.com/v1)" + +[i18n.zh.settings.auth_type] +label = "认证方式" +description = "API 请求的认证方式" + +[i18n.zh.settings.auth_token] +label = "认证令牌 / API 密钥" +description = "Bearer 令牌、API 密钥或 Base64 编码的凭据,取决于认证方式" + +[i18n.zh.settings.test_mode] +label = "测试模式" +description = "执行的 API 测试类型" + +[i18n.zh.settings.openapi_spec_url] +label = "OpenAPI 规范 URL" +description = "OpenAPI/Swagger 规范的 URL(例如 /openapi.json)。留空则自动发现。" + +[i18n.zh.settings.auto_schedule] +label = "自动定时" +description = "按计划自动运行测试" + +[i18n.zh.settings.test_frequency] +label = "测试频率" +description = "定时测试的执行频率" + +[i18n.zh.settings.fail_on_error] +label = "严格模式" +description = "将任何非 2xx 响应视为失败(而非允许预期的错误码)" + +[i18n.zh.settings.approval_mode] +label = "审批模式" +description = "将测试计划和破坏性请求写入队列文件供审核,而非直接执行" + +# ─── Japanese (日本語) ──────────────────────────────────────────────────── + +[i18n.ja] +name = "APIテスト Hand" +description = "自律型APIテストエージェント——エンドポイント検出、リクエスト検証、負荷テスト、リグレッション検出" +category = "開発" + +[i18n.ja.settings.base_url] +label = "ベースURL" +description = "テスト対象APIのベースURL(例: https://api.example.com/v1)" + +[i18n.ja.settings.auth_type] +label = "認証方式" +description = "APIリクエストの認証方法" + +[i18n.ja.settings.auth_token] +label = "認証トークン / APIキー" +description = "認証方式に応じたBearerトークン、APIキー、またはBase64エンコードされた資格情報" + +[i18n.ja.settings.test_mode] +label = "テストモード" +description = "実行するAPIテストの種類" + +[i18n.ja.settings.openapi_spec_url] +label = "OpenAPI仕様URL" +description = "OpenAPI/Swagger仕様のURL(例: /openapi.json)。空欄にすると自動検出します。" + +[i18n.ja.settings.auto_schedule] +label = "自動スケジュール" +description = "スケジュールに基づいてテストを自動実行する" + +[i18n.ja.settings.test_frequency] +label = "テスト頻度" +description = "定期テストの実行頻度" + +[i18n.ja.settings.fail_on_error] +label = "厳格モード" +description = "2xx以外のレスポンスをすべて失敗として扱う(期待されるエラーコードを許容しない)" + +[i18n.ja.settings.approval_mode] +label = "承認モード" +description = "テスト計画や破壊的リクエストを直接実行せず、レビュー用のキューファイルに書き出す" + +# ─── Spanish (Español) ──────────────────────────────────────────────────── + +[i18n.es] +name = "Hand de Pruebas API" +description = "Agente autónomo de pruebas de API — descubrimiento de endpoints, validación de peticiones, pruebas de carga y detección de regresiones" +category = "Desarrollo" + +[i18n.es.settings.base_url] +label = "URL base" +description = "URL base de la API a probar (ej. https://api.example.com/v1)" + +[i18n.es.settings.auth_type] +label = "Tipo de autenticación" +description = "Cómo autenticar las peticiones a la API" + +[i18n.es.settings.auth_token] +label = "Token de autenticación / Clave API" +description = "Token Bearer, clave API o credenciales codificadas en Base64 según el tipo de autenticación" + +[i18n.es.settings.test_mode] +label = "Modo de prueba" +description = "Qué tipo de pruebas de API realizar" + +[i18n.es.settings.openapi_spec_url] +label = "URL de especificación OpenAPI" +description = "URL de la especificación OpenAPI/Swagger (ej. /openapi.json). Dejar vacío para descubrimiento automático." + +[i18n.es.settings.auto_schedule] +label = "Programación automática" +description = "Ejecutar pruebas automáticamente según un calendario" + +[i18n.es.settings.test_frequency] +label = "Frecuencia de pruebas" +description = "Con qué frecuencia ejecutar las pruebas programadas" + +[i18n.es.settings.fail_on_error] +label = "Modo estricto" +description = "Tratar cualquier respuesta no 2xx como un fallo (en lugar de permitir códigos de error esperados)" + +[i18n.es.settings.approval_mode] +label = "Modo de aprobación" +description = "Escribir planes de prueba y peticiones destructivas en un archivo de cola para revisión en lugar de ejecutarlos directamente" + +# ─── French (Français) ──────────────────────────────────────────────────── + +[i18n.fr] +name = "Hand de Test API" +description = "Agent autonome de test d'API — découverte de points de terminaison, validation de requêtes, tests de charge et détection de régression" +category = "Développement" + +[i18n.fr.settings.base_url] +label = "URL de base" +description = "URL de base de l'API à tester (ex. https://api.example.com/v1)" + +[i18n.fr.settings.auth_type] +label = "Type d'authentification" +description = "Méthode d'authentification des requêtes API" + +[i18n.fr.settings.auth_token] +label = "Jeton d'authentification / Clé API" +description = "Jeton Bearer, clé API ou identifiants encodés en Base64 selon le type d'authentification" + +[i18n.fr.settings.test_mode] +label = "Mode de test" +description = "Type de tests API à exécuter" + +[i18n.fr.settings.openapi_spec_url] +label = "URL de spécification OpenAPI" +description = "URL de la spécification OpenAPI/Swagger (ex. /openapi.json). Laisser vide pour la découverte automatique." + +[i18n.fr.settings.auto_schedule] +label = "Planification automatique" +description = "Exécuter automatiquement les tests selon un calendrier" + +[i18n.fr.settings.test_frequency] +label = "Fréquence des tests" +description = "Fréquence d'exécution des tests planifiés" + +[i18n.fr.settings.fail_on_error] +label = "Mode strict" +description = "Traiter toute réponse non 2xx comme un échec (au lieu d'autoriser les codes d'erreur attendus)" + +[i18n.fr.settings.approval_mode] +label = "Mode d'approbation" +description = "Écrire les plans de test et requêtes destructives dans un fichier d'attente pour révision au lieu de les exécuter directement" + +# ─── German (Deutsch) ──────────────────────────────────────────────────── + +[i18n.de] +name = "API-Test-Hand" +description = "Autonomer API-Test-Agent — Endpunkt-Erkennung, Anfrage-Validierung, Lasttests und Regressionserkennung" +category = "Entwicklung" + +[i18n.de.settings.base_url] +label = "Basis-URL" +description = "Basis-URL der zu testenden API (z.B. https://api.example.com/v1)" + +[i18n.de.settings.auth_type] +label = "Authentifizierungstyp" +description = "Authentifizierungsmethode für API-Anfragen" + +[i18n.de.settings.auth_token] +label = "Authentifizierungstoken / API-Schlüssel" +description = "Bearer-Token, API-Schlüssel oder Base64-kodierte Anmeldedaten je nach Authentifizierungstyp" + +[i18n.de.settings.test_mode] +label = "Testmodus" +description = "Art der durchzuführenden API-Tests" + +[i18n.de.settings.openapi_spec_url] +label = "OpenAPI-Spezifikations-URL" +description = "URL der OpenAPI/Swagger-Spezifikation (z.B. /openapi.json). Leer lassen für automatische Erkennung." + +[i18n.de.settings.auto_schedule] +label = "Automatische Planung" +description = "Tests automatisch nach Zeitplan ausführen" + +[i18n.de.settings.test_frequency] +label = "Testhäufigkeit" +description = "Ausführungshäufigkeit der geplanten Tests" + +[i18n.de.settings.fail_on_error] +label = "Strikter Modus" +description = "Jede Nicht-2xx-Antwort als Fehler behandeln (anstatt erwartete Fehlercodes zuzulassen)" + +[i18n.de.settings.approval_mode] +label = "Genehmigungsmodus" +description = "Testpläne und destruktive Anfragen in eine Warteschlange zur Überprüfung schreiben, anstatt sie direkt auszuführen" + +# ─── Korean (한국어) ──────────────────────────────────────────────────── + +[i18n.ko] +name = "API 테스트 Hand" +description = "자율 API 테스트 에이전트 — 엔드포인트 탐색, 요청 검증, 부하 테스트 및 회귀 감지" +category = "개발" + +[i18n.ko.settings.base_url] +label = "기본 URL" +description = "테스트할 API의 기본 URL (예: https://api.example.com/v1)" + +[i18n.ko.settings.auth_type] +label = "인증 방식" +description = "API 요청의 인증 방식" + +[i18n.ko.settings.auth_token] +label = "인증 토큰 / API 키" +description = "인증 방식에 따른 Bearer 토큰, API 키 또는 Base64 인코딩 자격 증명" + +[i18n.ko.settings.test_mode] +label = "테스트 모드" +description = "수행할 API 테스트 유형" + +[i18n.ko.settings.openapi_spec_url] +label = "OpenAPI 스펙 URL" +description = "OpenAPI/Swagger 스펙의 URL (예: /openapi.json). 비워두면 자동 탐색합니다." + +[i18n.ko.settings.auto_schedule] +label = "자동 일정" +description = "일정에 따라 자동으로 테스트 실행" + +[i18n.ko.settings.test_frequency] +label = "테스트 빈도" +description = "정기 테스트 실행 주기" + +[i18n.ko.settings.fail_on_error] +label = "엄격 모드" +description = "모든 비-2xx 응답을 실패로 처리 (예상된 오류 코드 허용 안 함)" + +[i18n.ko.settings.approval_mode] +label = "승인 모드" +description = "테스트 계획 및 파괴적 요청을 직접 실행하지 않고 큐 파일에 기록하여 검토" diff --git a/hands/apitester/SKILL.md b/hands/apitester/SKILL.md index 58fc634..71126e4 100644 --- a/hands/apitester/SKILL.md +++ b/hands/apitester/SKILL.md @@ -237,3 +237,713 @@ First Seen: 2025-01-15 run Previous Value: string (email format) Current Value: field absent ``` + +--- + +## Worked Examples + +### Example 1: Testing a REST API CRUD Endpoint + +Full test suite for a `/api/users` resource covering create, read, update, delete, and edge cases. + +**Setup — Create a test user**: +```bash +# POST /api/users — create +RESPONSE=$(curl -s -w "\n%{http_code}" -X POST \ + -H "Content-Type: application/json" \ + -H "Authorization: Bearer $TOKEN" \ + -d '{"name": "Ada Lovelace", "email": "ada@example.com", "role": "engineer"}' \ + "https://api.example.com/api/users") + +BODY=$(echo "$RESPONSE" | sed '$d') +STATUS=$(echo "$RESPONSE" | tail -1) + +# Expect 201 Created +[ "$STATUS" = "201" ] && echo "PASS: Create user" || echo "FAIL: Expected 201, got $STATUS" + +# Extract ID for subsequent tests +USER_ID=$(echo "$BODY" | python3 -c "import sys,json; print(json.load(sys.stdin)['id'])") +``` + +**Read operations**: +```bash +# GET /api/users — list all +curl -s -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/users" | python3 -m json.tool + +# GET /api/users/:id — single user +curl -s -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/users/$USER_ID" | python3 -m json.tool + +# GET /api/users/nonexistent-id — expect 404 +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/users/00000000-0000-0000-0000-000000000000") +[ "$STATUS" = "404" ] && echo "PASS: 404 for missing user" || echo "FAIL: Expected 404, got $STATUS" +``` + +**Update operations**: +```bash +# PUT /api/users/:id — full update +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X PUT \ + -H "Content-Type: application/json" -H "Authorization: Bearer $TOKEN" \ + -d '{"name": "Ada Lovelace", "email": "ada.updated@example.com", "role": "lead"}' \ + "https://api.example.com/api/users/$USER_ID") +[ "$STATUS" = "200" ] && echo "PASS: Full update" || echo "FAIL: Expected 200, got $STATUS" + +# PATCH — partial update (expect 200); also test invalid data (expect 400/422) +``` + +**Delete and verify**: +```bash +# DELETE /api/users/:id +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X DELETE \ + -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/users/$USER_ID") +[ "$STATUS" = "204" ] || [ "$STATUS" = "200" ] && echo "PASS: Delete user" || echo "FAIL: Expected 2xx, got $STATUS" + +# GET deleted user — expect 404 or 410 +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/users/$USER_ID") +[ "$STATUS" = "404" ] || [ "$STATUS" = "410" ] && echo "PASS: Deleted user gone" || echo "FAIL: Expected 404/410, got $STATUS" + +# DELETE again — idempotency check +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X DELETE \ + -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/users/$USER_ID") +[ "$STATUS" = "404" ] || [ "$STATUS" = "204" ] && echo "PASS: Idempotent delete" || echo "FAIL: Got $STATUS" +``` + +**Edge cases to test**: duplicate create (expect 409), empty body (expect 400/422), extra unknown fields (verify ignored or rejected, not persisted). + +### Example 2: Testing an Authenticated API with Rate Limiting + +Scenario: API uses Bearer tokens, tokens expire after 1 hour, rate limit is 100 requests/minute. + +**Token lifecycle testing**: +```bash +# Step 1: Obtain token +AUTH_RESPONSE=$(curl -s -X POST \ + -H "Content-Type: application/json" \ + -d '{"client_id": "myapp", "client_secret": "secret", "grant_type": "client_credentials"}' \ + "https://api.example.com/oauth/token") + +ACCESS_TOKEN=$(echo "$AUTH_RESPONSE" | python3 -c "import sys,json; print(json.load(sys.stdin)['access_token'])") +EXPIRES_IN=$(echo "$AUTH_RESPONSE" | python3 -c "import sys,json; print(json.load(sys.stdin)['expires_in'])") +echo "Token obtained, expires in ${EXPIRES_IN}s" + +# Step 2: Use token — expect 200 +STATUS=$(curl -s -o /dev/null -w "%{http_code}" \ + -H "Authorization: Bearer $ACCESS_TOKEN" \ + "https://api.example.com/api/protected") +[ "$STATUS" = "200" ] && echo "PASS: Valid token accepted" || echo "FAIL: Got $STATUS" + +# Step 3: Use expired/invalid token — expect 401 +STATUS=$(curl -s -o /dev/null -w "%{http_code}" \ + -H "Authorization: Bearer expired.token.here" \ + "https://api.example.com/api/protected") +[ "$STATUS" = "401" ] && echo "PASS: Expired token rejected" || echo "FAIL: Got $STATUS" + +# Step 4: Missing Authorization header — expect 401 +STATUS=$(curl -s -o /dev/null -w "%{http_code}" \ + "https://api.example.com/api/protected") +[ "$STATUS" = "401" ] && echo "PASS: No auth rejected" || echo "FAIL: Got $STATUS" + +# Step 5: Malformed header — expect 401 +STATUS=$(curl -s -o /dev/null -w "%{http_code}" \ + -H "Authorization: NotBearer $ACCESS_TOKEN" \ + "https://api.example.com/api/protected") +[ "$STATUS" = "401" ] && echo "PASS: Bad scheme rejected" || echo "FAIL: Got $STATUS" +``` + +**Rate limit testing**: +```bash +# Hit the endpoint rapidly and watch for 429 +RESULTS_FILE=$(mktemp) +for i in $(seq 1 120); do + curl -s -o /dev/null -w "%{http_code}\n" \ + -H "Authorization: Bearer $ACCESS_TOKEN" \ + "https://api.example.com/api/data" >> "$RESULTS_FILE" & +done +wait + +# Count status codes +echo "=== Rate Limit Results ===" +sort "$RESULTS_FILE" | uniq -c | sort -rn +# Expected: ~100 x 200, ~20 x 429 + +# Check rate limit headers on a single request +curl -s -D- -o /dev/null \ + -H "Authorization: Bearer $ACCESS_TOKEN" \ + "https://api.example.com/api/data" | grep -i "x-ratelimit" +# Expected headers: +# X-RateLimit-Limit: 100 +# X-RateLimit-Remaining: 99 +# X-RateLimit-Reset: 1700000060 + +rm "$RESULTS_FILE" +``` + +**Backoff strategy**: On 429, respect `Retry-After` header. Use exponential backoff (1s, 2s, 4s...) as fallback. Verify the API returns `X-RateLimit-Reset` for client scheduling. + +### Example 3: Testing a Webhook Endpoint + +Scenario: Your API accepts webhook callbacks at `POST /webhooks/payment` with HMAC-SHA256 signature verification. + +**Payload and signature generation**: +```bash +WEBHOOK_SECRET="whsec_test_secret_key_12345" +PAYLOAD='{"event":"payment.completed","data":{"id":"pay_123","amount":4999,"currency":"usd"}}' +TIMESTAMP=$(date +%s) +SIGNATURE=$(printf "%s.%s" "$TIMESTAMP" "$PAYLOAD" | openssl dgst -sha256 -hmac "$WEBHOOK_SECRET" | awk '{print $2}') + +# Valid webhook delivery +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X POST \ + -H "Content-Type: application/json" \ + -H "X-Webhook-Signature: t=$TIMESTAMP,v1=$SIGNATURE" \ + -H "X-Webhook-Id: wh_evt_001" \ + -d "$PAYLOAD" \ + "https://api.example.com/webhooks/payment") +[ "$STATUS" = "200" ] || [ "$STATUS" = "204" ] && echo "PASS: Valid webhook accepted" || echo "FAIL: Got $STATUS" +``` + +**Signature verification tests**: +```bash +# Wrong signature — expect 401 or 403 +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X POST \ + -H "Content-Type: application/json" \ + -H "X-Webhook-Signature: t=$TIMESTAMP,v1=badsignaturevalue" \ + -d "$PAYLOAD" \ + "https://api.example.com/webhooks/payment") +[ "$STATUS" = "401" ] || [ "$STATUS" = "403" ] && echo "PASS: Bad signature rejected" || echo "FAIL: Got $STATUS" + +# Missing signature header — expect 401 +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X POST \ + -H "Content-Type: application/json" \ + -d "$PAYLOAD" \ + "https://api.example.com/webhooks/payment") +[ "$STATUS" = "401" ] && echo "PASS: Missing signature rejected" || echo "FAIL: Got $STATUS" + +# Stale timestamp (replay attack) — expect 403 +OLD_TIMESTAMP=$((TIMESTAMP - 600)) +OLD_SIGNATURE=$(printf "%s.%s" "$OLD_TIMESTAMP" "$PAYLOAD" | openssl dgst -sha256 -hmac "$WEBHOOK_SECRET" | awk '{print $2}') +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X POST \ + -H "Content-Type: application/json" \ + -H "X-Webhook-Signature: t=$OLD_TIMESTAMP,v1=$OLD_SIGNATURE" \ + -d "$PAYLOAD" \ + "https://api.example.com/webhooks/payment") +[ "$STATUS" = "403" ] && echo "PASS: Stale timestamp rejected" || echo "FAIL: Got $STATUS" +``` + +**Also test**: idempotency (same `X-Webhook-Id` sent twice — should be processed once), invalid/empty payloads (expect 400). + +--- + +## Authentication Testing Patterns + +### OAuth 2.0 Flow Testing + +**Authorization Code flow**: +```bash +# Step 1: Initiate authorization — verify redirect +AUTHORIZE_URL="https://api.example.com/oauth/authorize?response_type=code&client_id=myapp&redirect_uri=https://myapp.example.com/callback&scope=read+write&state=random_state_123" +STATUS=$(curl -s -o /dev/null -w "%{http_code}" "$AUTHORIZE_URL") +[ "$STATUS" = "302" ] || [ "$STATUS" = "200" ] && echo "PASS: Auth endpoint reachable" || echo "FAIL: Got $STATUS" + +# Step 2: Exchange authorization code for token +TOKEN_RESPONSE=$(curl -s -X POST \ + -H "Content-Type: application/x-www-form-urlencoded" \ + -d "grant_type=authorization_code&code=AUTH_CODE_HERE&redirect_uri=https://myapp.example.com/callback&client_id=myapp&client_secret=secret" \ + "https://api.example.com/oauth/token") +echo "$TOKEN_RESPONSE" | python3 -m json.tool +# Verify: access_token, refresh_token, expires_in, token_type present + +# Step 3: Use invalid authorization code — expect 400 +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X POST \ + -H "Content-Type: application/x-www-form-urlencoded" \ + -d "grant_type=authorization_code&code=INVALID_CODE&redirect_uri=https://myapp.example.com/callback&client_id=myapp&client_secret=secret" \ + "https://api.example.com/oauth/token") +[ "$STATUS" = "400" ] && echo "PASS: Invalid code rejected" || echo "FAIL: Got $STATUS" + +# Step 4: Reuse authorization code — must fail (codes are single-use) +# Use the same AUTH_CODE_HERE again +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X POST \ + -H "Content-Type: application/x-www-form-urlencoded" \ + -d "grant_type=authorization_code&code=AUTH_CODE_HERE&redirect_uri=https://myapp.example.com/callback&client_id=myapp&client_secret=secret" \ + "https://api.example.com/oauth/token") +[ "$STATUS" = "400" ] && echo "PASS: Code reuse rejected" || echo "FAIL: Got $STATUS" +``` + +**Client Credentials flow**: Same pattern as above with `grant_type=client_credentials`. Test: valid credentials (expect `access_token`), invalid secret (expect 401), invalid `grant_type` (expect 400). + +**Refresh Token flow**: Exchange `grant_type=refresh_token` with `refresh_token=$REFRESH_TOKEN`. Verify: new `access_token` returned, old refresh token invalidated if rotation is enabled (reuse should return 400/401). + +### JWT Validation Testing + +Test each type of JWT failure independently: + +| Test Case | Token Modification | Expected Status | Expected Error | +|-----------|-------------------|-----------------|----------------| +| Expired token | Set `exp` to past timestamp | 401 | `token_expired` | +| Not-yet-valid | Set `nbf` to future timestamp | 401 | `token_not_yet_valid` | +| Wrong signature | Sign with different key | 401 | `invalid_signature` | +| Malformed token | Remove a segment | 401 | `malformed_token` | +| Missing `sub` claim | Remove `sub` from payload | 401 | `missing_claims` | +| Wrong audience | Set `aud` to different app | 401 | `invalid_audience` | +| Wrong issuer | Set `iss` to unknown issuer | 401 | `invalid_issuer` | +| Algorithm none attack | Set `alg: none`, remove signature | 401 | `invalid_algorithm` | + +```bash +# Generate a test JWT with wrong signature (using python3 as a helper) +HEADER=$(echo -n '{"alg":"HS256","typ":"JWT"}' | base64 | tr -d '=' | tr '+/' '-_') +PAYLOAD=$(echo -n '{"sub":"user123","exp":9999999999}' | base64 | tr -d '=' | tr '+/' '-_') +BAD_SIG=$(echo -n "fakesignature" | base64 | tr -d '=' | tr '+/' '-_') +BAD_JWT="${HEADER}.${PAYLOAD}.${BAD_SIG}" + +STATUS=$(curl -s -o /dev/null -w "%{http_code}" \ + -H "Authorization: Bearer $BAD_JWT" \ + "https://api.example.com/api/protected") +[ "$STATUS" = "401" ] && echo "PASS: Bad JWT signature rejected" || echo "FAIL: Got $STATUS" + +# Algorithm "none" attack +NONE_HEADER=$(echo -n '{"alg":"none","typ":"JWT"}' | base64 | tr -d '=' | tr '+/' '-_') +NONE_JWT="${NONE_HEADER}.${PAYLOAD}." +STATUS=$(curl -s -o /dev/null -w "%{http_code}" \ + -H "Authorization: Bearer $NONE_JWT" \ + "https://api.example.com/api/protected") +[ "$STATUS" = "401" ] && echo "PASS: alg:none attack blocked" || echo "FAIL: Got $STATUS — SECURITY RISK" +``` + +### API Key Testing Patterns + +```bash +# Valid API key in header +STATUS=$(curl -s -o /dev/null -w "%{http_code}" \ + -H "X-API-Key: valid_key_abc123" \ + "https://api.example.com/api/data") +[ "$STATUS" = "200" ] && echo "PASS: Valid API key" || echo "FAIL: Got $STATUS" +``` + +**Also test**: key in query param (if supported), revoked key (expect 401/403), empty key (expect 401), read-only key attempting write (expect 403). + +### Session-Based Auth Testing + +Test pattern: login (capture `Set-Cookie`), use cookie for authenticated request (expect 200), logout, reuse cookie (expect 401). Also verify session fixation prevention — session ID should rotate on login. + +--- + +## Contract Testing + +### Schema Validation Techniques + +Validate API responses against a JSON Schema using `python3 -c "from jsonschema import validate; ..."`: + +```bash +# Fetch response and validate against schema file +curl -s -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/users/user_001" | python3 -c " +import sys, json +from jsonschema import validate, ValidationError +schema = json.load(open('/tmp/user_schema.json')) +try: + validate(instance=json.load(sys.stdin), schema=schema) + print('PASS: Schema valid') +except ValidationError as e: + print(f'FAIL: {e.message}') +" +``` + +Schema should define `required` fields, property `type`/`format`/`enum` constraints, and `additionalProperties: false` for strict mode. + +### Breaking Change Detection + +Compare current response structure against a recorded baseline: + +```bash +# Helper: extract JSON shape as "path: type" lines +extract_shape() { + curl -s -H "Authorization: Bearer $TOKEN" "$1" | python3 -c " +import sys, json +def shape(obj, prefix=''): + s = {} + if isinstance(obj, dict): + for k, v in obj.items(): + p = f'{prefix}.{k}' if prefix else k + s[p] = type(v).__name__; s.update(shape(v, p)) + elif isinstance(obj, list) and obj: + s[f'{prefix}[]'] = type(obj[0]).__name__; s.update(shape(obj[0], f'{prefix}[]')) + return s +for p, t in sorted(shape(json.load(sys.stdin)).items()): print(f'{p}: {t}') +" +} + +# Record baseline once, then diff against current +extract_shape "https://api.example.com/api/users/user_001" > /tmp/api_baseline.txt +# ... later ... +extract_shape "https://api.example.com/api/users/user_001" > /tmp/api_current.txt +diff /tmp/api_baseline.txt /tmp/api_current.txt && echo "PASS: No schema changes" || echo "WARN: Schema changed" +``` + +### Backward Compatibility Checklist + +When a new API version is deployed, verify that existing consumers are not broken: + +| Check | How to Test | Severity | +|-------|------------|----------| +| Removed fields | Diff response shape against baseline | **HIGH** — breaks consumers | +| Renamed fields | Diff response keys | **HIGH** — breaks consumers | +| Changed field type | Compare type of each field | **HIGH** — breaks deserialization | +| New required request field | Send old-format request | **HIGH** — breaks callers | +| Changed enum values | Check if old values still accepted | **MEDIUM** — breaks validation | +| Changed error format | Compare error response structure | **MEDIUM** — breaks error handlers | +| Changed status codes | Compare response codes for same input | **MEDIUM** — breaks status checks | +| New optional fields | Verify response still parses | **LOW** — usually safe | +| Pagination format change | Test with existing page params | **MEDIUM** — breaks pagination loops | + +### Consumer-Driven Contract Testing + +Concept: Each API consumer defines the minimum contract they need (required fields, forbidden fields, expected status codes). The provider runs all consumer contracts in CI. + +```json +{ + "consumer": "mobile-app-v2", + "provider": "user-service", + "interactions": [ + { + "description": "get user profile", + "request": {"method": "GET", "path": "/api/users/me", "headers": {"Authorization": "Bearer valid_token"}}, + "response": {"status": 200, "body_contains": ["id", "name", "email"], "body_must_not_contain": ["password", "internal_id"]} + } + ] +} +``` + +Runner approach: iterate interactions, execute each request with curl, verify status code matches and required/forbidden fields are present/absent in the response body. + +--- + +## Performance Testing Deep Dive + +### Load Test Types + +| Type | Purpose | Pattern | +|------|---------|---------| +| **Soak** | Detect memory leaks, connection pool exhaustion | Steady traffic (e.g., 5 req/s) for hours; compare first-quarter vs last-quarter response times | +| **Spike** | Verify graceful handling of sudden bursts | Baseline → 10x-20x burst → recovery; check error rate and recovery time | +| **Stress** | Find the breaking point | Incrementally increase concurrency until errors begin | + +### Stress Testing (Representative Example) + +Incrementally increase load until errors begin — adapt the same pattern for soak (fixed concurrency, long duration) or spike (sudden burst) testing: + +```bash +echo "concurrency,success_rate,avg_time,p95_time" > /tmp/stress_results.csv +for CONCURRENCY in 10 25 50 100 200 500; do + RESULTS=$(mktemp) + for i in $(seq 1 $CONCURRENCY); do + curl -s -o /dev/null -w "%{http_code} %{time_total}\n" \ + -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/data" >> "$RESULTS" & + done + wait + + TOTAL=$(wc -l < "$RESULTS") + SUCCESS=$(grep -c "^200" "$RESULTS") + AVG_TIME=$(awk '{sum+=$2; n++} END {printf "%.3f", sum/n}' "$RESULTS") + P95_TIME=$(awk '{print $2}' "$RESULTS" | sort -n | awk -v p=0.95 'NR==1{n=0} {a[n++]=$1} END {print a[int(n*p)]}') + + echo "$CONCURRENCY,$((SUCCESS*100/TOTAL))%,$AVG_TIME,$P95_TIME" >> /tmp/stress_results.csv + echo "Concurrency $CONCURRENCY: ${SUCCESS}/${TOTAL} success, avg=${AVG_TIME}s, p95=${P95_TIME}s" + + rm "$RESULTS" + sleep 3 # Let the server recover between steps +done + +echo "=== Stress Test Summary ===" +column -t -s',' /tmp/stress_results.csv +``` + +### Latency Percentile Analysis + +Collect many response times (e.g., 1000 with concurrency capped at 20), then compute p50/p75/p90/p95/p99 percentiles. Compare first-quarter vs last-quarter averages to detect degradation over time. + +```bash +# Collect response times +TIMES_FILE=$(mktemp) +for i in $(seq 1 1000); do + curl -s -o /dev/null -w "%{time_total}\n" \ + -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/data" >> "$TIMES_FILE" & + [ $((i % 20)) -eq 0 ] && wait +done +wait +# Sort and compute percentiles with: sort -n "$TIMES_FILE" | python3 ... +rm "$TIMES_FILE" +``` + +### Connection Pool Testing + +- **Keep-alive reuse**: Send multiple URLs in one curl call with `Connection: keep-alive`; second/third requests should show near-zero `time_connect`. +- **Connection exhaustion**: Open 500 concurrent keep-alive connections; watch for 503 or connection refused errors. + +--- + +## Common API Bugs & How to Find Them + +### N+1 Query Detection + +Response time should not scale linearly with data size. If fetching 10 items takes 100ms but 100 items takes 1000ms, the API likely has an N+1 query problem. + +```bash +# Compare response times for different page sizes +for SIZE in 1 10 50 100; do + TIME=$(curl -s -o /dev/null -w "%{time_total}" \ + -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/orders?per_page=$SIZE") + echo "page_size=$SIZE time=${TIME}s" +done +# Expected (healthy): Times should NOT scale linearly +# page_size=1 time=0.045s +# page_size=10 time=0.052s +# page_size=50 time=0.078s +# page_size=100 time=0.110s +# Red flag (N+1): Times scale roughly linearly +# page_size=1 time=0.045s +# page_size=10 time=0.350s +# page_size=50 time=1.600s +# page_size=100 time=3.200s +``` + +### Race Condition Testing + +```bash +# Concurrent counter increment — final value should equal attempt count +curl -s -X PUT -H "Content-Type: application/json" -H "Authorization: Bearer $TOKEN" \ + -d '{"value": 0}' "https://api.example.com/api/counters/counter_001" + +for i in $(seq 1 50); do + curl -s -X POST -H "Content-Type: application/json" -H "Authorization: Bearer $TOKEN" \ + -d '{"increment": 1}' "https://api.example.com/api/counters/counter_001/increment" & +done +wait + +FINAL=$(curl -s -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/counters/counter_001" | python3 -c "import sys,json; print(json.load(sys.stdin)['value'])") +[ "$FINAL" = "50" ] && echo "PASS: No race condition" || echo "FAIL: Lost $((50 - FINAL)) increments" +``` + +**Optimistic locking test**: Two concurrent PUTs with same `If-Match` ETag — one should get 200, the other 409 Conflict. + +### Pagination Edge Cases + +| Input | Expected Behavior | +|-------|------------------| +| `page=0` | 400, or treat as page 1 | +| `page=-1` | 400 | +| `page=99999` (beyond data) | 200 with empty array, not error | +| `per_page=0` | 400 or use default | +| `per_page=100000` | Capped to server max (e.g., 100) | +| Delete item mid-pagination | No items skipped or duplicated on next page | + +### Timezone Handling Bugs + +Test that equivalent timestamps in different offset formats are stored identically: + +```bash +# All four represent the same moment — stored values should be equivalent +for TZ in "2025-06-15T10:00:00Z" "2025-06-15T10:00:00+00:00" "2025-06-15T18:00:00+08:00" "2025-06-15T05:00:00-05:00"; do + STORED=$(curl -s -X POST -H "Content-Type: application/json" -H "Authorization: Bearer $TOKEN" \ + -d "{\"title\": \"tz_test\", \"scheduled_at\": \"$TZ\"}" \ + "https://api.example.com/api/events" | python3 -c "import sys,json; print(json.load(sys.stdin).get('scheduled_at','ERROR'))") + echo "Input: $TZ -> Stored: $STORED" +done +``` + +**Also test**: date range filters across timezone boundaries, midnight boundary inclusion/exclusion behavior. + +### Character Encoding Issues + +Test that the API correctly round-trips various Unicode inputs. Key test values: + +| Category | Example | What Breaks | +|----------|---------|-------------| +| Emoji | `Hello 🌍🚀` | UTF-8 4-byte sequences, database column width | +| CJK | `你好世界` | Multi-byte encoding, string length vs byte length | +| Diacritics | `café` (composed vs decomposed) | Unicode normalization (NFC vs NFD) | +| Zero-width | `test\u200Bword` | Invisible characters in search/comparison | +| Null byte | `test\u0000value` | String termination in C-based systems | + +```bash +# Round-trip test pattern: POST a value, verify GET returns the same +for VALUE in "Hello 🌍🚀" "你好世界" "café"; do + RESPONSE=$(curl -s -X POST -H "Content-Type: application/json; charset=utf-8" \ + -H "Authorization: Bearer $TOKEN" \ + -d "{\"name\": \"$VALUE\"}" \ + "https://api.example.com/api/items") + RETURNED=$(echo "$RESPONSE" | python3 -c "import sys,json; print(json.load(sys.stdin).get('name','ERROR'))") + [ "$VALUE" = "$RETURNED" ] && echo "PASS: $VALUE" || echo "FAIL: sent='$VALUE' got='$RETURNED'" +done +``` + +--- + +## Advanced curl Patterns + +### File Upload Testing + +```bash +# Single file upload +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X POST \ + -H "Authorization: Bearer $TOKEN" \ + -F "file=@/path/to/document.pdf" \ + -F "description=Test upload" \ + "https://api.example.com/api/uploads") +echo "Single file upload: $STATUS" + +# Multiple file upload +STATUS=$(curl -s -o /dev/null -w "%{http_code}" -X POST \ + -H "Authorization: Bearer $TOKEN" \ + -F "files[]=@/path/to/file1.png" \ + -F "files[]=@/path/to/file2.png" \ + -F "category=images" \ + "https://api.example.com/api/uploads/batch") +echo "Multi-file upload: $STATUS" +``` + +**Edge cases to also test**: oversized files (expect 413), wrong content type (e.g., `script.sh` declared as `image/png`), zero-byte files (expect 400). + +### Multipart Form Data + +```bash +# Mixed multipart: file + JSON metadata +curl -s -X POST \ + -H "Authorization: Bearer $TOKEN" \ + -F "metadata={\"title\":\"Report Q4\",\"tags\":[\"finance\",\"quarterly\"]};type=application/json" \ + -F "file=@/path/to/report.pdf" \ + "https://api.example.com/api/documents" + +# Form-encoded data (not JSON) +curl -s -X POST \ + -H "Content-Type: application/x-www-form-urlencoded" \ + -d "username=testuser&password=testpass&remember=true" \ + "https://api.example.com/auth/login" +``` + +### Cookie-Based Session Testing + +```bash +# Full session lifecycle with cookie jar +COOKIE_JAR=$(mktemp) + +# Login — store cookies +curl -s -c "$COOKIE_JAR" -X POST \ + -H "Content-Type: application/json" \ + -d '{"username": "testuser", "password": "testpass"}' \ + "https://api.example.com/auth/login" + +# Authenticated request — send cookies +curl -s -b "$COOKIE_JAR" -c "$COOKIE_JAR" \ + "https://api.example.com/api/profile" + +# Logout and verify session invalidated +curl -s -b "$COOKIE_JAR" -c "$COOKIE_JAR" -X POST \ + "https://api.example.com/auth/logout" +STATUS=$(curl -s -b "$COOKIE_JAR" -o /dev/null -w "%{http_code}" \ + "https://api.example.com/api/profile") +[ "$STATUS" = "401" ] && echo "PASS: Session invalidated" || echo "FAIL: Got $STATUS" +rm "$COOKIE_JAR" +``` + +**Also verify**: HttpOnly/Secure/SameSite cookie attributes, session ID rotation on login (session fixation prevention). + +### Following Redirects + +```bash +# Follow redirects automatically +curl -s -L -o /dev/null -w "final_url:%{url_effective} status:%{http_code} redirects:%{num_redirects}\n" \ + "https://api.example.com/old-endpoint" + +# Don't follow — inspect redirect target +curl -s -D- -o /dev/null \ + "https://api.example.com/old-endpoint" | grep -i "location:" + +# Open redirect vulnerability test +LOCATION=$(curl -s -D- -o /dev/null \ + "https://api.example.com/redirect?url=https://evil.example.com" | grep -i "location:" | tr -d '\r') +echo "$LOCATION" | grep -q "evil.example.com" && echo "FAIL: Open redirect vulnerability" || echo "PASS: Redirect restricted" + +# HTTP to HTTPS redirect check +STATUS=$(curl -s -o /dev/null -w "%{http_code}" "http://api.example.com/api/data") +[ "$STATUS" = "301" ] || [ "$STATUS" = "308" ] && echo "PASS: HTTP redirects to HTTPS" || echo "WARN: No HTTPS redirect (got $STATUS)" +``` + +### HEAD, OPTIONS, and CORS + +```bash +# HEAD request — verify no body returned +curl -s -I -w "status:%{http_code} size:%{size_download}\n" \ + -H "Authorization: Bearer $TOKEN" \ + "https://api.example.com/api/data" + +# OPTIONS request — check CORS and allowed methods +curl -s -X OPTIONS -D- -o /dev/null \ + -H "Origin: https://myapp.example.com" \ + -H "Access-Control-Request-Method: POST" \ + "https://api.example.com/api/data" | grep -iE "(allow|access-control)" +``` + +--- + +## Chaos & Fault Injection Patterns + +| Fault | How to Inject | Expected Behavior | +|-------|--------------|-------------------| +| Slow client | `curl --limit-rate 1k` | Server does not hold connection indefinitely; times out gracefully | +| Partial body | Pipe truncated JSON via `echo '{"name":' \| curl -d @-` | 400 Bad Request, not 500 | +| Huge header | `-H "X-Pad: $(python3 -c 'print("A"*16000)')"` | 431 Request Header Fields Too Large or 400 | +| Concurrent duplicate | Fire same POST with idempotency key 50x in parallel | Exactly one resource created; others get 409 or identical response | +| Connection reset | `curl --max-time 0.001` (client aborts mid-response) | Server logs show no crash; subsequent requests succeed | +| Malformed encoding | Send `Content-Type: application/json; charset=iso-8859-1` with UTF-8 body | API rejects or correctly transcodes; no mojibake in stored data | + +--- + +## API Versioning Test Strategies + +When an API exposes multiple versions, verify isolation and deprecation handling: + +| Test | Method | Expected | +|------|--------|----------| +| Old version still works | `GET /api/v1/resource` | 200 with v1 schema (or 410 if sunset) | +| New version returns new schema | `GET /api/v2/resource` | 200 with v2 fields present | +| Version via header | `Accept: application/vnd.api.v2+json` | Response matches v2 schema | +| Unsupported version | `GET /api/v99/resource` | 404 or 400, not fallback to latest | +| Sunset header | Check `Sunset:` and `Deprecation:` headers on old versions | Headers present with valid dates | +| Cross-version mutation | Create in v1, read in v2 and vice versa | Data accessible in both; fields map correctly | + +--- + +## GraphQL-Specific Testing Patterns + +When the target exposes a GraphQL endpoint (`POST /graphql`): + +- **Introspection**: Send `{ __schema { types { name } } }` — should be disabled in production (expect error), or return schema if intentionally public +- **Query depth attack**: Nest a query 15+ levels deep (e.g. `{ user { friends { friends { ... } } } }`) — expect a depth-limit error, not a timeout +- **Batch attack**: Send an array of 100 queries in one request — expect rejection or rate limiting, not 100x execution cost +- **Field suggestion leak**: Send a query with a typo (e.g. `{ usr { name } }`) — verify the error does not suggest valid field names in production +- **Alias-based DoS**: Query the same expensive field 50 times using aliases (`a1: expensiveField, a2: expensiveField, ...`) — expect query complexity rejection +- **Mutation authorization**: Execute mutations for other users' resources — expect authorization errors identical to REST BOLA checks +- **N+1 detection**: Query a list with nested relations (`{ users { orders { items } } }`) — linear response time scaling signals N+1 + +--- + +## Webhook Reliability Testing Patterns + +Beyond signature verification (covered in worked examples), test delivery reliability: + +| Scenario | How to Simulate | What to Verify | +|----------|----------------|----------------| +| Slow consumer | Respond with 200 after 25s delay | Sender respects timeout >30s; does not mark as failed prematurely | +| Consumer down | Return 503 for first 3 deliveries | Sender retries with exponential backoff; check `X-Retry-Count` | +| Duplicate delivery | Verify same `X-Webhook-Id` arrives twice | Consumer handles idempotently — no duplicate side effects | +| Out-of-order events | Process events t2 before t1 | Consumer uses event timestamp, not arrival order, for state | +| Oversized payload | Trigger event producing >1MB payload | Sender truncates or sends reference URL instead of inline data | +| Replay attack | Accept delivery with timestamp >5min old | Consumer rejects stale deliveries to prevent replay | diff --git a/hands/browser/HAND.toml b/hands/browser/HAND.toml index 3ff4f86..1fd7e0d 100644 --- a/hands/browser/HAND.toml +++ b/hands/browser/HAND.toml @@ -142,6 +142,59 @@ description = "Automatically take a screenshot after every click/navigate for vi setting_type = "toggle" default = "false" +[[settings]] +key = "cookie_persistence" +label = "Cookie Persistence" +description = "Persist cookies across tasks in the same session to maintain login state and preferences" +setting_type = "toggle" +default = "true" + +[[settings]] +key = "user_agent" +label = "User Agent" +description = "Browser user-agent string sent with requests — affects how websites identify the browser" +setting_type = "select" +default = "chrome_desktop" + +[[settings.options]] +value = "chrome_desktop" +label = "Chrome Desktop (most compatible)" + +[[settings.options]] +value = "firefox_desktop" +label = "Firefox Desktop" + +[[settings.options]] +value = "chrome_mobile" +label = "Chrome Mobile (Android)" + +[[settings.options]] +value = "safari_mobile" +label = "Safari Mobile (iOS)" + +[[settings]] +key = "viewport_size" +label = "Viewport Size" +description = "Browser window dimensions — affects responsive layout and which version of a site is served" +setting_type = "select" +default = "1920x1080" + +[[settings.options]] +value = "1920x1080" +label = "1920x1080 (Full HD desktop)" + +[[settings.options]] +value = "1366x768" +label = "1366x768 (Laptop)" + +[[settings.options]] +value = "390x844" +label = "390x844 (Mobile)" + +[[settings.options]] +value = "1024x768" +label = "1024x768 (Tablet)" + # ─── Agent configuration ───────────────────────────────────────────────────── [agent] @@ -157,114 +210,153 @@ system_prompt = """You are Browser Hand — an autonomous web browser agent that ## Core Capabilities -You can navigate to URLs, click buttons/links, fill forms, read page content, and take screenshots. You have a real browser session that persists across tool calls within a conversation. +You can navigate to URLs, click buttons/links, fill forms, read page content, and take screenshots. You have a real browser session that persists across tool calls within a conversation. Cookies and login state carry over between actions unless the session is explicitly closed. ## Multi-Phase Pipeline -### Phase 1 — Understand the Task -Parse the user's request and plan your approach: +### Phase 1 — Understand & Plan +Parse the user's request and build an execution plan: - What website(s) do you need to visit? - What information do you need to find or what action do you need to perform? - What are the success criteria? +- Is the target likely a SPA (single-page app) or a traditional server-rendered site? +- Will login or cookie consent be needed before reaching the goal? ### Phase 2 — Navigate & Observe 1. Use `browser_navigate` to go to the target URL -2. Read the page content to understand the layout -3. Identify the relevant elements (buttons, links, forms, search boxes) +2. Use `browser_read_page` to understand the page structure +3. Identify page type: static HTML, SPA framework, or hybrid +4. Handle blocking overlays immediately (cookie banners, modals, age gates) +5. Verify you are on the correct domain and the page loaded completely +6. If content appears empty or minimal, wait 3-5 seconds and re-read — SPAs often render asynchronously -### Phase 3 — Interact -1. Use `browser_click` for buttons and links (use CSS selectors or visible text) +### Phase 3 — Detect & Adapt to Page Technology +Detect the page technology to choose the right interaction strategy: + +**SPA detection signals** (any of these means client-side rendering): +- Page has a single `
` or `
` with most content nested inside +- URL changes do not trigger full page reloads (hash routes like `#/page` or history API routes) +- Content appears after a delay with loading spinners or skeleton screens +- Page source is minimal HTML with large JS bundles + +**SPA interaction rules:** +- After every click that changes the view, wait 1-3 seconds before reading the page +- Look for loading indicators: `[aria-busy="true"]`, `.loading`, `.spinner`, `.skeleton` +- If `browser_read_page` returns stale content, wait and retry (up to 3 attempts) +- Prefer clicking visible UI elements over direct URL navigation (SPAs may not support deep links) + +**Iframe handling:** +- If target content is inside an iframe, note that `browser_read_page` may not capture iframe contents +- Try navigating directly to the iframe's `src` URL if you need to interact with its content +- For embedded widgets (payment forms, third-party logins), inform the user if interaction is blocked + +**Shadow DOM:** +- Some web components use shadow DOM which hides elements from normal selectors +- If a known element is not found, it may be inside a shadow root +- Use `browser_screenshot` to visually confirm the element exists, then try interacting by visible text + +### Phase 4 — Interact & Verify +1. Use `browser_click` for buttons and links — prefer these selector strategies in order: + a. `[data-testid="..."]` or `[data-test="..."]` — most stable, survives UI redesigns + b. `[aria-label="..."]` or `[role="button"]` — accessibility-based, framework-independent + c. `#id` — unique but may be auto-generated in SPAs + d. Visible text content — reliable fallback when selectors fail + e. CSS class selectors — least stable, use only as last resort 2. Use `browser_type` for filling form fields -3. Use `browser_read_page` after each action to see the updated state -4. Use `browser_screenshot` when you need visual verification +3. Use `browser_read_page` after each action to verify the expected state change occurred +4. Use `browser_screenshot` when text content alone is ambiguous or for visual verification +5. If an action produces no visible change, check for overlays, disabled states, or incomplete page loads before retrying -### Phase 4 — MANDATORY Purchase/Payment Approval +### Phase 5 — Error Recovery & Retry +When an interaction fails, follow this decision tree: + +1. **Element not found:** + a. Re-read the page — DOM may have changed since last read + b. Try alternative selectors: data-testid > aria-label > role > visible text > class + c. Scroll the page to trigger lazy loading, then re-read + d. Take a screenshot to see the actual page state + e. If still not found after 3 attempts, report to user with what was tried + +2. **Click has no effect:** + a. Check for overlays blocking the element (cookie banners, modals, chat widgets) + b. Dismiss overlays: look for "Accept", "Close", "X", or `[aria-label="Close"]` buttons + c. Check if the element is disabled (`[disabled]`, `[aria-disabled="true"]`, `.disabled`) + d. Try clicking a more specific child element (e.g., the `` inside a `