pdf 代码审计

数据 社区
解读按原文结构重写,命令、链接、术语均保留;右侧可核对作者原始 SKILL.md

原文结构

  • Key Principles:Preserve the logical structure of the document: headings, sections, lists, and table relationships. When extracting data, maintain the original ordering and hierarchy unless the user requests a different organization. Clearly distinguish between exact text…
  • Extraction Techniques:For text-based PDFs, extract content while preserving paragraph boundaries and section headings. For scanned PDFs, use OCR tools (tesseract, pdf2image + OCR, or cloud OCR APIs) and note the confidence level. For tables, reconstruct the row/column structure.…
  • Analysis Patterns:Summarization: Provide a hierarchical summary — one-line overview, then section-by-section breakdown. Data extraction: Pull specific data points (dates, amounts, names, addresses) into structured formats. Comparison: When comparing multiple PDFs, align them by…

适用与边界

  • 当前原文没有单列适用场景。
  • Pitfalls to Avoid:Do not assume all text in a PDF is selectable — some documents are scanned images. Do not ignore headers, footers, and page numbers that may interfere with content flow. Do not merge table cells incorrectly — verify row/column alignment before presenting…

原文中的明确线索

  • 要点:「Summarization」、「Data extraction」、「Comparison」、「Search」、「Metadata」
  • 文件与命令tesseractpdf2imagerow/columnCSV/JSONInvoices/receipts

流狐整理:以上内容来自当前 SKILL.md 的章节与原词;未补写作者没有声明的工具、兼容性或能力。

流狐档案 作者与许可取自来源;运行、权限和网络为流狐检测或估算
流狐分类
数据
作者声明 Agent
未找到明确声明;不据此推断已兼容或已测试
静态检查
88 / 100 · 启发式扫描,不代表运行安全
作者 / 版本 / 许可
@RightNow-AI · 未声明 license
流狐 Token 估算
低消耗
流狐接入估算
即装即用
是否需要外部 API Key
未发现要求
检测到的系统要求
未声明
底层运行要求
未声明
检测到的文件与系统行为
  • 只读
检测到的网络行为
仅限本地
安装命令数
无(仅作为资料)

档案由构建时根据 SKILL.md 与安装命令自动衍生,可能与作者实际意图存在差异。

需要注意: 未限定 allowed-tools,默认拥有全部工具权限。

输出预览 pdf-reader.preview
作者没有在当前 SKILL.md 中定义固定输出样例。

讨论

基于 GitHub Discussions。登录 GitHub 即可参与讨论、点赞、订阅更新。