调试 dags
- 作者仓库星标 39
- 作者仓库 awesome-omni-skill
DAG Diagnosis
You are a data engineer debugging a failed Airflow DAG. Follow this systematic approach to identify the root cause and provide actionable remediation.
Running the CLI
Run all af commands using uvx (no installation required):
uvx --from astro-airflow-mcp af <command>
Throughout this document, af is shorthand for uvx --from astro-airflow-mcp af.
Step 1: Identify the Failure
If a specific DAG was mentioned:
- Run
af runs diagnose <dag_id> <dag_run_id>(if run_id is provided) - If no run_id specified, run
af dags statsto find recent failures
If no DAG was specified:
- Run
af healthto find recent failures across all DAGs - Check for import errors with
af dags errors - Show DAGs with recent failures
- Ask which DAG to investigate further
Step 2: Get the Error Details
Once you have identified a failed task:
- Get task logs using
af tasks logs <dag_id> <dag_run_id> <task_id> - Look for the actual exception - scroll past the Airflow boilerplate to find the real error
- Categorize the failure type:
- Data issue: Missing data, schema change, null values, constraint violation
- Code issue: Bug, syntax error, import failure, type error
- Infrastructure issue: Connection timeout, resource exhaustion, permission denied
- Dependency issue: Upstream failure, external API down, rate limiting
Step 3: Check Context
Gather additional context to understand WHY this happened:
- Recent changes: Was there a code deploy? Check git history if available
- Data volume: Did data volume spike? Run a quick count on source tables
- Upstream health: Did upstream tasks succeed but produce unexpected data?
- Historical pattern: Is this a recurring failure? Check if same task failed before
- Timing: Did this fail at an unusual time? (resource contention, maintenance windows)
Use af runs get <dag_id> <dag_run_id> to compare the failed run against recent successful runs.
On Astro
If you're running on Astro, these additional tools can help with diagnosis:
- Deployment activity log: Check the Astro UI for recent deploys — a failed deploy or recent code change is often the cause of sudden failures
- Astro alerts: Configure alerts in the Astro UI for proactive failure monitoring (DAG failure, task duration, SLA miss)
- Observability: Use the Astro observability dashboard to track DAG health trends and spot recurring issues
On OSS Airflow
- Airflow UI: Use the DAGs page, Graph view, and task logs to inspect recent runs and failures
Step 4: Provide Actionable Output
Structure your diagnosis as:
Root Cause
What actually broke? Be specific - not "the task failed" but "the task failed because column X was null in 15% of rows when the code expected 0%".
Impact Assessment
- What data is affected? Which tables didn't get updated?
- What downstream processes are blocked?
- Is this blocking production dashboards or reports?
Immediate Fix
Specific steps to resolve RIGHT NOW:
- If it's a data issue: SQL to fix or skip bad records
- If it's a code issue: The exact code change needed
- If it's infra: Who to contact or what to restart
Prevention
How to prevent this from happening again:
- Add data quality checks?
- Add better error handling?
- Add alerting for edge cases?
- Update documentation?
Quick Commands
Provide ready-to-use commands:
- To clear and rerun the entire DAG run:
af runs clear <dag_id> <run_id> - To clear and rerun specific failed tasks:
af tasks clear <dag_id> <run_id> <task_ids> -D - To delete a stuck or unwanted run:
af runs delete <dag_id> <run_id>
- 流狐分类
- 运维部署
- 作者声明 Agent
- 未找到明确声明;不据此推断已兼容或已测试
- 静态检查
- 88 / 100 · 启发式扫描,不代表运行安全
- 作者 / 版本 / 许可
- @diegosouzapw · 未声明 license
- 流狐 Token 估算
- 低消耗
- 流狐接入估算
- 即装即用
- 是否需要外部 API Key
- 未发现要求
- 检测到的系统要求
- Windows
- 底层运行要求
- 未声明
- 检测到的文件与系统行为
-
- 只读
- 允许写入 / 修改
- 检测到的网络行为
- 仅限本地
- 安装命令数
- 无(仅作为资料)
档案由构建时根据 SKILL.md 与安装命令自动衍生,可能与作者实际意图存在差异。
需要注意: 未限定 allowed-tools,默认拥有全部工具权限。
# Step 4: Provide Actionable Output
Structure your diagnosis as: If a specific DAG was mentioned: Run af runs diagnose <dagid> <dagrunid> (if runid is provided) If no runid specified, run af dags stats to find recent failures
Once you have identified a failed task: Get task logs using af tasks logs <dagid> <dagrunid> <taskid> Look for the actual exception - scroll past the Airflow boilerplate to find the real error
Gather additional context to understand WHY this happened: Recent changes: Was there a code deploy? Check git history if available Data volume: Did data volume spike? Run a quick count on source tables
Structure your diagnosis as:
# DAG Diagnosis
You are a data engineer debugging a failed Airflow DAG. Follow this systematic approach to identify the root cause and provide actionable remediation.
## Running the CLI
Run all `af` commands using uvx (no installation required):
```bash
uvx --from astro-airflow-mcp af <command>
```
Throughout this document, `af` is shorthand for `uvx --from astro-airflow-mcp af`.
---
## Step 1: Identify the Failure
If a specific DAG was mentioned:
- Run `af runs diagnose <dag_id> <dag_run_id>` (if run_id is provided)
- If no run_id specified, run `af dags stats` to find recent failures
If no DAG was specified:
- Run `af health` to find recent failures across all DAGs
- Check for import errors with `af dags errors`
- Show DAGs with recent failures
- Ask which DAG to investigate further
## Step 2: Get the Error Details
Once you have identified a failed task:
1. **Get task logs** using `af tasks logs <dag_id> <dag_run_id> <task_id>`
2. **Look for the actual exception** - scroll past the Airflow boilerplate to find the real error
3. **Categorize the failure type**:
- **Data issue**: Missing data, schema change, null values, constraint violation
- **Code issue**: Bug, syntax error, import failure, type error
- **Infrastructure issue**: Connection timeout, resource exhaustion, permission denied
- **Dependency issue**: Upstream failure, external API down, rate limiting
## Step 3: Check Context
Gather additional context to understand WHY this happened:
1. **Recent changes**: Was there a code deploy? Check git history if available
2. **Data volume**: Did data volume spike? Run a quick count on source tables
3. **Upstream health**: Did upstream tasks succeed but produce unexpected data?
… 作者原文负责流程事实;流狐只索引当前章节、要点、文件与命令。
章节 -> Running the CLI → Step 1: Identify the Failure → Step 2: Get the Error Details → Step 3: Check Context → On Astro → On OSS Airflow
要点 -> Get task logs · Look for the actual exception · Categorize the failure type · Data issue · Code issue · Infrastructure issue · Dependency issue · Recent changes
文件/命令 -> uvx --from astro-airflow-mcp af · af runs diagnose <dagid> <dagrunid> · af dags stats · af health · af dags errors · af tasks logs <dagid> <dagrunid> <taskid> · af runs get <dagid> <dagrunid> · af runs clear <dagid> <runid>
内容 SHA-256 -> a4a3508d2f66
方法与流程
适用与边界
原文中的明确线索
uvx --from astro-airflow-mcp af、af runs diagnose <dagid> <dagrunid>、af dags stats、af health、af dags errors、af tasks logs <dagid> <dagrunid> <taskid>、af runs get <dagid> <dagrunid>、af runs clear <dagid> <runid>