Docker 助手
- 作者仓库星标 0
- 作者仓库 skills-registry
Docker Databricks Lab Ops
Overview
Use this skill for the operational loop of this repository: bring up Docker services, generate source-table mutations, reset Databricks tables, execute a Databricks notebook or job, and verify whether the notebook run finished successfully.
Prefer the bundled scripts over rewriting shell commands. They encode the repository-specific paths and the expected sequence.
Workflow
1. Inspect the repo inputs first
- Confirm
docker-compose.yml,postgres-connector.json, and the target notebook paths exist. - Read references/repo-workflow.md if you need the repo-specific sequence or parameters.
2. Bring up the local CDC stack
- Use
scripts/start_stack.sh. - This runs
docker compose up -dfrom the repository root. - If the user asked for verification, follow with
docker compose psor service-specific health checks. - If Kafka must be reachable from Databricks through an internal network boundary, first use
scripts/prepare_ngrok_kafka.pyand then start Compose with the discoveredKAFKA_EXTERNAL_HOSTandKAFKA_EXTERNAL_PORT.
3. Register Debezium connector if CDC ingestion needs to be exercised
- Use
scripts/register_connector.sh. - Only do this after Kafka Connect is accepting requests.
- If the connector already exists, report that clearly instead of treating it as a fatal failure.
4. Run load generators
- Use
scripts/run_generators.sh. - The first argument is film update iterations, the second is rental/payment iterations.
- Prefer bounded runs for verification by passing
ITERATIONS; avoid indefinite generators unless the user asked for sustained load.
Example:
skills/docker-databricks-lab-ops/scripts/run_generators.sh 20 40
This runs 20 film mutations and 40 rental/payment mutations.
5. Reset Databricks tables (when starting fresh)
- Use
scripts/reset_databricks_tables.pyto drop all Bronze/Silver/Gold tables and clear streaming checkpoints before a fresh load. - Requires
--cluster-idor will submit via git source. - Use
--dry-runto preview what would be dropped without dropping.
python3 skills/docker-databricks-lab-ops/scripts/reset_databricks_tables.py \
--cluster-id <cluster-id> \
--catalog workspace
# preview only:
python3 skills/docker-databricks-lab-ops/scripts/reset_databricks_tables.py \
--cluster-id <cluster-id> \
--dry-run
6. Trigger a Databricks notebook job
- Use
scripts/run_databricks_notebook.py. - Provide either:
--job-idto run an existing Databricks job, or--notebook-pathand--cluster-idto submit a one-off notebook run.
- For dynamic Kafka exposure, pass
--notebook-param KAFKA_BOOTSTRAP=<ngrok-host:port>so the current tunnel endpoint is used at run time instead of a stale static value. - This script uses
DATABRICKS_HOSTandDATABRICKS_TOKENfrom the environment.
Examples:
python3 skills/docker-databricks-lab-ops/scripts/run_databricks_notebook.py \
--job-id 123 \
--notebook-param KAFKA_BOOTSTRAP=0.tcp.eu.ngrok.io:12345
python3 skills/docker-databricks-lab-ops/scripts/run_databricks_notebook.py \
--notebook-path /Workspace/agent/notebook \
--cluster-id 0123-456abc-cluster
7. Verify notebook behavior
- Treat the job as successful only when lifecycle is terminal and result is
SUCCESS. - On failure, capture:
run_id- lifecycle state
- result state
- state message
- notebook path or job id
- If the run succeeded, summarize which notebook or job was exercised and what evidence was collected.
8. Report the outcome
- State what was started locally.
- State whether load generation ran and with what iteration counts.
- State which Databricks job or notebook was executed.
- State whether notebook verification passed or failed.
- If it failed, include the exact failure message and the next corrective step.
9. Use the smoke test when the user wants one-command verification
- Use
scripts/smoke_test_notebooks.py. - It discovers or starts an ngrok tunnel, restarts Docker with the correct advertised Kafka listener, registers the connector, runs bounded load generators, triggers the Databricks job, and waits for terminal results.
- Pass
--reset --cluster-id <id>to drop all tables and checkpoints before the smoke run.
# Full smoke test with table reset:
python3 skills/docker-databricks-lab-ops/scripts/smoke_test_notebooks.py \
--job-id 123 \
--reset \
--cluster-id <cluster-id>
# Smoke test without reset (append to existing data):
python3 skills/docker-databricks-lab-ops/scripts/smoke_test_notebooks.py \
--job-id 123
Scripts
scripts/start_stack.sh: start Docker Compose services for the labscripts/prepare_ngrok_kafka.py: discover or start an ngrok TCP tunnel for Kafka and print the current public bootstrapscripts/register_connector.sh: register the Debezium connector frompostgres-connector.jsonscripts/run_generators.sh: run film and rental/payment generatorsscripts/reset_databricks_tables.py: drop all Bronze/Silver/Gold Delta tables and clear streaming checkpointsscripts/run_databricks_notebook.py: launch or submit a Databricks run and poll to completionscripts/smoke_test_notebooks.py: run the end-to-end smoke test with dynamic ngrok bootstrap handlingscripts/migrate_and_run.py: full migration script — drops legacy tables, updates the Databricks job via API, resets dvdrental tables, starts Docker+connector+generators, and triggers the job end-to-end
References
references/repo-workflow.md: repo-specific execution order, assumptions, and parameters
<!-- tomevault:4.0:skill_md:2026-05-22 -->Source: alexeyban/databricks-lab — distributed by TomeVault.
- 流狐分类
- 运维部署
- 作者声明 Agent
- 未找到明确声明;不据此推断已兼容或已测试
- 静态检查
- 88 / 100 · 启发式扫描,不代表运行安全
- 作者 / 版本 / 许可
- @tomevault-io · 未声明 license
- 流狐 Token 估算
- 低消耗
- 流狐接入估算
- 需手动接入
- 是否需要外部 API Key
- 未发现要求
- 检测到的系统要求
- Docker
- 底层运行要求
- Docker
- 检测到的文件与系统行为
-
- 只读
- 允许写入 / 修改
- Shell 执行
- 检测到的网络行为
- 仅限本地
- 安装命令数
- 无(仅作为资料)
档案由构建时根据 SKILL.md 与安装命令自动衍生,可能与作者实际意图存在差异。
需要注意: 未限定 allowed-tools,默认拥有全部工具权限。
作者没有在当前 SKILL.md 中定义固定输出样例。 Use this skill for the operational loop of this repository: bring up Docker services, generate source-table mutations, reset Databricks tables, execute a Databricks notebook or job, and verify whether the notebook run finished successfully.
Workflow
Confirm docker-compose.yml, postgres-connector.json, and the target notebook paths exist. Read references/repo-workflow.md if you need the repo-specific sequence or parameters.
Use scripts/startstack.sh. This runs docker compose up -d from the repository root. If the user asked for verification, follow with docker compose ps or service-specific health checks.
Use scripts/registerconnector.sh. Only do this after Kafka Connect is accepting requests. If the connector already exists, report that clearly instead of treating it as a fatal failure.
Use scripts/rungenerators.sh. The first argument is film update iterations, the second is rental/payment iterations. Prefer bounded runs for verification by passing ITERATIONS; avoid indefinite generators unless the user asked for sustained load.
# Docker Databricks Lab Ops
## Overview
Use this skill for the operational loop of this repository: bring up Docker services, generate source-table mutations, reset Databricks tables, execute a Databricks notebook or job, and verify whether the notebook run finished successfully.
Prefer the bundled scripts over rewriting shell commands. They encode the repository-specific paths and the expected sequence.
## Workflow
### 1. Inspect the repo inputs first
- Confirm `docker-compose.yml`, `postgres-connector.json`, and the target notebook paths exist.
- Read [references/repo-workflow.md](./references/repo-workflow.md) if you need the repo-specific sequence or parameters.
### 2. Bring up the local CDC stack
- Use `scripts/start_stack.sh`.
- This runs `docker compose up -d` from the repository root.
- If the user asked for verification, follow with `docker compose ps` or service-specific health checks.
- If Kafka must be reachable from Databricks through an internal network boundary, first use `scripts/prepare_ngrok_kafka.py` and then start Compose with the discovered `KAFKA_EXTERNAL_HOST` and `KAFKA_EXTERNAL_PORT`.
### 3. Register Debezium connector if CDC ingestion needs to be exercised
- Use `scripts/register_connector.sh`.
- Only do this after Kafka Connect is accepting requests.
- If the connector already exists, report that clearly instead of treating it as a fatal failure.
### 4. Run load generators
- Use `scripts/run_generators.sh`.
- The first argument is film update iterations, the second is rental/payment iterations.
- Prefer bounded runs for verification by passing `ITERATIONS`; avoid indefinite generators unless the user asked for sustained load.
Example:
```bash
skills/docker-databricks-lab-ops/scripts/run_generators.sh 20 40
```
… 作者原文负责流程事实;流狐只索引当前章节、要点、文件与命令。
章节 -> Overview → Workflow → 1. Inspect the repo inputs first → 2. Bring up the local CDC stack → 3. Register Debezium connector if CDC ingestion needs to be exercised → 4. Run load generators
要点 -> Prefer the bundled scripts over rewriting shell commands. · - Confirm docker-compose.yml, postgres-connector.json, and the target notebook paths exist. · - Use scripts/startstack.sh. · - Use scripts/registerconnector.sh. · - Use scripts/rungenerators.sh. · This runs 20 film mutations and 40 rental/payment mutations. · - Use scripts/resetdatabrickstables.py to drop all Bronze/Silver/Gold tables and clear streaming checkpoints before a fresh load. · - Use scripts/rundatabricksnotebook.py.
文件/命令 -> docker-compose.yml · postgres-connector.json · scripts/startstack.sh · docker compose up -d · docker compose ps · scripts/preparengrokkafka.py · KAFKAEXTERNALHOST · KAFKAEXTERNALPORT
内容 SHA-256 -> a44f2c351990
方法与流程
适用与边界
原文中的明确线索
docker-compose.yml、postgres-connector.json、scripts/startstack.sh、docker compose up -d、docker compose ps、scripts/preparengrokkafka.py、KAFKAEXTERNALHOST、KAFKAEXTERNALPORT