sn da Excel 工作流

数据 社区
解读按原文结构重写,命令、链接、术语均保留;右侧可核对作者原始 SKILL.md

方法与流程

  • Workflow:原文以此为独立章节。
  • Step 1 — Count rows across all sheets (lightweight, no full load):Count rows per sheet without loading data into memory. Use openpyxl readonly mode — this works for any file size. ⚠️ Do NOT use pd.readexcel() to count rows — it loads all data into
  • Step 2 — Large file gate (CRITICAL — choose strategy by row count):totalrows · Strategy · What to do < 10k · Direct read · df = pd.readexcel(filepath, sheetname=targetsheet) 10k – 100k · Parquet cache · pd.readexcel() once → df.toparquet() → all later reads from Parquet

适用与边界

  • 当前原文没有单列适用场景。
  • Key rules:Always count rows first — gate large-file logic on the 10k threshold. >= 100k rows → MUST load sn-da-large-file-analysis skill — do not attempt to handle with pd.readexcel(). Column names may contain spaces (e.g. '是否通 过') — use exact string indexing.

原文中的明确线索

  • 要点:「without loading data into memory」、「Do NOT use pd.readexcel() to count rows」、「>= 100k」、「STOP. Load sn-da-large-file-analysis skill」、「Do NOT use pd.readexcel() at all」、「For >= 100k rows」、「For 10k – 100k rows (only)」、「Large file rule」
  • 文件与命令readonlypd.readexcel()excel-reading/multi-sheet-readingdf = pd.readexcel(filepath, sheetname=targetsheet)df.toparquet()sn-da-large-file-analysisstreamexceltoparquet()iterrows

流狐整理:以上内容来自当前 SKILL.md 的章节与原词;未补写作者没有声明的工具、兼容性或能力。

流狐档案 作者与许可取自来源;运行、权限和网络为流狐检测或估算
流狐分类
数据
作者声明 Agent
未找到明确声明;不据此推断已兼容或已测试
静态检查
88 / 100 · 启发式扫描,不代表运行安全
作者 / 版本 / 许可
@OpenSenseNova · 未声明 license
流狐 Token 估算
低消耗
流狐接入估算
需简单配置
是否需要外部 API Key
未发现要求
检测到的系统要求
未声明
底层运行要求
Python
检测到的文件与系统行为
  • 只读
  • Shell 执行
检测到的网络行为
仅限本地
安装命令数
无(仅作为资料)

档案由构建时根据 SKILL.md 与安装命令自动衍生,可能与作者实际意图存在差异。

需要注意: 未限定 allowed-tools,默认拥有全部工具权限。

输出预览 sn-da-excel-workflow.preview
# Step 6 — Export results

output_path = "/mnt/data/result.xlsx"
result_df.to_excel(output_path, index=False)
print(f"[Download](sandbox:{output_path})")

讨论

基于 GitHub Discussions。登录 GitHub 即可参与讨论、点赞、订阅更新。