docker-databricks-lab-ops
- Repo stars 0
- Author repo skills-registry
Docker Databricks Lab Ops
Overview
Use this skill for the operational loop of this repository: bring up Docker services, generate source-table mutations, reset Databricks tables, execute a Databricks notebook or job, and verify whether the notebook run finished successfully.
Prefer the bundled scripts over rewriting shell commands. They encode the repository-specific paths and the expected sequence.
Workflow
1. Inspect the repo inputs first
- Confirm
docker-compose.yml,postgres-connector.json, and the target notebook paths exist. - Read references/repo-workflow.md if you need the repo-specific sequence or parameters.
2. Bring up the local CDC stack
- Use
scripts/start_stack.sh. - This runs
docker compose up -dfrom the repository root. - If the user asked for verification, follow with
docker compose psor service-specific health checks. - If Kafka must be reachable from Databricks through an internal network boundary, first use
scripts/prepare_ngrok_kafka.pyand then start Compose with the discoveredKAFKA_EXTERNAL_HOSTandKAFKA_EXTERNAL_PORT.
3. Register Debezium connector if CDC ingestion needs to be exercised
- Use
scripts/register_connector.sh. - Only do this after Kafka Connect is accepting requests.
- If the connector already exists, report that clearly instead of treating it as a fatal failure.
4. Run load generators
- Use
scripts/run_generators.sh. - The first argument is film update iterations, the second is rental/payment iterations.
- Prefer bounded runs for verification by passing
ITERATIONS; avoid indefinite generators unless the user asked for sustained load.
Example:
skills/docker-databricks-lab-ops/scripts/run_generators.sh 20 40
This runs 20 film mutations and 40 rental/payment mutations.
5. Reset Databricks tables (when starting fresh)
- Use
scripts/reset_databricks_tables.pyto drop all Bronze/Silver/Gold tables and clear streaming checkpoints before a fresh load. - Requires
--cluster-idor will submit via git source. - Use
--dry-runto preview what would be dropped without dropping.
python3 skills/docker-databricks-lab-ops/scripts/reset_databricks_tables.py \
--cluster-id <cluster-id> \
--catalog workspace
# preview only:
python3 skills/docker-databricks-lab-ops/scripts/reset_databricks_tables.py \
--cluster-id <cluster-id> \
--dry-run
6. Trigger a Databricks notebook job
- Use
scripts/run_databricks_notebook.py. - Provide either:
--job-idto run an existing Databricks job, or--notebook-pathand--cluster-idto submit a one-off notebook run.
- For dynamic Kafka exposure, pass
--notebook-param KAFKA_BOOTSTRAP=<ngrok-host:port>so the current tunnel endpoint is used at run time instead of a stale static value. - This script uses
DATABRICKS_HOSTandDATABRICKS_TOKENfrom the environment.
Examples:
python3 skills/docker-databricks-lab-ops/scripts/run_databricks_notebook.py \
--job-id 123 \
--notebook-param KAFKA_BOOTSTRAP=0.tcp.eu.ngrok.io:12345
python3 skills/docker-databricks-lab-ops/scripts/run_databricks_notebook.py \
--notebook-path /Workspace/agent/notebook \
--cluster-id 0123-456abc-cluster
7. Verify notebook behavior
- Treat the job as successful only when lifecycle is terminal and result is
SUCCESS. - On failure, capture:
run_id- lifecycle state
- result state
- state message
- notebook path or job id
- If the run succeeded, summarize which notebook or job was exercised and what evidence was collected.
8. Report the outcome
- State what was started locally.
- State whether load generation ran and with what iteration counts.
- State which Databricks job or notebook was executed.
- State whether notebook verification passed or failed.
- If it failed, include the exact failure message and the next corrective step.
9. Use the smoke test when the user wants one-command verification
- Use
scripts/smoke_test_notebooks.py. - It discovers or starts an ngrok tunnel, restarts Docker with the correct advertised Kafka listener, registers the connector, runs bounded load generators, triggers the Databricks job, and waits for terminal results.
- Pass
--reset --cluster-id <id>to drop all tables and checkpoints before the smoke run.
# Full smoke test with table reset:
python3 skills/docker-databricks-lab-ops/scripts/smoke_test_notebooks.py \
--job-id 123 \
--reset \
--cluster-id <cluster-id>
# Smoke test without reset (append to existing data):
python3 skills/docker-databricks-lab-ops/scripts/smoke_test_notebooks.py \
--job-id 123
Scripts
scripts/start_stack.sh: start Docker Compose services for the labscripts/prepare_ngrok_kafka.py: discover or start an ngrok TCP tunnel for Kafka and print the current public bootstrapscripts/register_connector.sh: register the Debezium connector frompostgres-connector.jsonscripts/run_generators.sh: run film and rental/payment generatorsscripts/reset_databricks_tables.py: drop all Bronze/Silver/Gold Delta tables and clear streaming checkpointsscripts/run_databricks_notebook.py: launch or submit a Databricks run and poll to completionscripts/smoke_test_notebooks.py: run the end-to-end smoke test with dynamic ngrok bootstrap handlingscripts/migrate_and_run.py: full migration script — drops legacy tables, updates the Databricks job via API, resets dvdrental tables, starts Docker+connector+generators, and triggers the job end-to-end
References
references/repo-workflow.md: repo-specific execution order, assumptions, and parameters
<!-- tomevault:4.0:skill_md:2026-05-22 -->Source: alexeyban/databricks-lab — distributed by TomeVault.
- Fluxly category
- DevOps
- Author-declared agents
- No explicit declaration found; this is not inferred or tested compatibility
- Static check
- 88 / 100 · heuristic scan, not runtime safety proof
- Author / version / license
- @tomevault-io · no license declared
- Fluxly token estimate
- Lean
- Fluxly setup estimate
- Manual integration
- External API key
- No requirement detected
- Detected OS requirements
- Docker
- Runtime requirements
- Docker
- Detected file/system behavior
-
- Read-only
- Write / modify
- Shell exec
- Detected network behavior
- Local-only
- Install commands
- None (reference only)
Profile is derived at build time from SKILL.md and install vectors. Subject to drift from author intent.
Heads up: 未限定 allowed-tools,默认拥有全部工具权限。
The current SKILL.md does not define a fixed output example. Use this skill for the operational loop of this repository: bring up Docker services, generate source-table mutations, reset Databricks tables, execute a Databricks notebook or job, and verify whether the notebook run finished successfully.
Workflow
Confirm docker-compose.yml, postgres-connector.json, and the target notebook paths exist. Read references/repo-workflow.md if you need the repo-specific sequence or parameters.
Use scripts/startstack.sh. This runs docker compose up -d from the repository root. If the user asked for verification, follow with docker compose ps or service-specific health checks.
Use scripts/registerconnector.sh. Only do this after Kafka Connect is accepting requests. If the connector already exists, report that clearly instead of treating it as a fatal failure.
Use scripts/rungenerators.sh. The first argument is film update iterations, the second is rental/payment iterations. Prefer bounded runs for verification by passing ITERATIONS; avoid indefinite generators unless the user asked for sustained load.
# Docker Databricks Lab Ops
## Overview
Use this skill for the operational loop of this repository: bring up Docker services, generate source-table mutations, reset Databricks tables, execute a Databricks notebook or job, and verify whether the notebook run finished successfully.
Prefer the bundled scripts over rewriting shell commands. They encode the repository-specific paths and the expected sequence.
## Workflow
### 1. Inspect the repo inputs first
- Confirm `docker-compose.yml`, `postgres-connector.json`, and the target notebook paths exist.
- Read [references/repo-workflow.md](./references/repo-workflow.md) if you need the repo-specific sequence or parameters.
### 2. Bring up the local CDC stack
- Use `scripts/start_stack.sh`.
- This runs `docker compose up -d` from the repository root.
- If the user asked for verification, follow with `docker compose ps` or service-specific health checks.
- If Kafka must be reachable from Databricks through an internal network boundary, first use `scripts/prepare_ngrok_kafka.py` and then start Compose with the discovered `KAFKA_EXTERNAL_HOST` and `KAFKA_EXTERNAL_PORT`.
### 3. Register Debezium connector if CDC ingestion needs to be exercised
- Use `scripts/register_connector.sh`.
- Only do this after Kafka Connect is accepting requests.
- If the connector already exists, report that clearly instead of treating it as a fatal failure.
### 4. Run load generators
- Use `scripts/run_generators.sh`.
- The first argument is film update iterations, the second is rental/payment iterations.
- Prefer bounded runs for verification by passing `ITERATIONS`; avoid indefinite generators unless the user asked for sustained load.
Example:
```bash
skills/docker-databricks-lab-ops/scripts/run_generators.sh 20 40
```
… Author text anchors workflow facts; Fluxly only indexes current sections, terms, files, and commands.
sections -> Overview → Workflow → 1. Inspect the repo inputs first → 2. Bring up the local CDC stack → 3. Register Debezium connector if CDC ingestion needs to be exercised → 4. Run load generators
terms -> Prefer the bundled scripts over rewriting shell commands. · - Confirm docker-compose.yml, postgres-connector.json, and the target notebook paths exist. · - Use scripts/startstack.sh. · - Use scripts/registerconnector.sh. · - Use scripts/rungenerators.sh. · This runs 20 film mutations and 40 rental/payment mutations. · - Use scripts/resetdatabrickstables.py to drop all Bronze/Silver/Gold tables and clear streaming checkpoints before a fresh load. · - Use scripts/rundatabricksnotebook.py.
files/cmd -> docker-compose.yml · postgres-connector.json · scripts/startstack.sh · docker compose up -d · docker compose ps · scripts/preparengrokkafka.py · KAFKAEXTERNALHOST · KAFKAEXTERNALPORT
body sha256 -> a44f2c351990
Decide Fit First
Design Intent
How To Use It
Boundaries And Review