docetl

Engineering Community
Interpretation is structured for decision-making; original keeps the upstream SKILL.md unchanged.

Decide Fit First

  • Core job: Build and run LLM-powered data processing pipelines with DocETL. Use when users say "docetl", want to analyze unstructured data…
  • Best fit: Use it when the task has reusable inputs, steps, and validation criteria rather than a one-off answer.
  • Avoid forcing it: If the source lacks commands, platform support, or external-service evidence, keep those fields unknown instead of guessing.

Design Intent

  • Structure: The skill is organized around “Workflow Overview: Iterative Data Analysis”, “Phase 1: Data Collection”, “Phase 2: Pipeline Development”, “Phase 3: Visualization & Presentation”, showing how the author expects the agent to judge fit, collect context, and produce verifiable output.
  • Trigger evidence: Prioritize the author’s wording around when to use it, what context to collect, and what output shape to produce.
  • Evidence boundary: Author text states facts, repository files prove commands and paths, and Fluxly only adds fit, limits, and usage judgment.

How To Use It

  • Inputs: Provide target material, scope, expected result, forbidden changes, and validation method.
  • Invocation: Name docetl directly; if the source includes slash commands, start with the command and then add task context.
  • Validation: Start small and check whether the result follows “Workflow Overview: Iterative Data Analysis / Phase 1: Data Collection / Phase 2: Pipeline Development” before expanding.

Boundaries And Review

  • Dependencies: Prepare OpenAI / Anthropic / Gemini API keys before running a full task.
  • Permissions: Declared permissions include read / write / shell-exec / env-read; ask the agent to state file, command, and rollback boundaries before acting.
  • Quality bar: A useful result names the deliverable, evidence, and next action. Generic prose means the task needs tighter context.
Fluxly profile Author and license come from source; runtime, permissions, and network are Fluxly detections or estimates
Fluxly category
Engineering
Author-declared agents
No explicit declaration found; this is not inferred or tested compatibility
Static check
94 / 100 · heuristic scan, not runtime safety proof
Author / version / license
@ucbepic · MIT
Fluxly token estimate
Heavy
Fluxly setup estimate
Manual integration
External API key
Required · OpenAI / Anthropic / Gemini
Detected OS requirements
Unspecified
Runtime requirements
Python
Detected file/system behavior
  • Read-only
  • Write / modify
  • Shell exec
  • Env read
Detected network behavior
External requests
Install commands
None (reference only)

Profile is derived at build time from SKILL.md and install vectors. Subject to drift from author intent.

Heads up: 未限定 allowed-tools,默认拥有全部工具权限。

Output preview docetl.preview
# Bad outputs

- Read more input data examples to improve prompt specificity
- Add `validate` rules with retries
- Simplify output schema
- Add concrete examples to prompt

Discussion

Powered by GitHub Discussions. Sign in with GitHub to comment, react, or subscribe.