e 代码审查
- 作者仓库星标 1,112
- 作者仓库 web-agent
E-commerce Extraction
General Patterns
- Check for sitemap.xml first -- many stores list all product URLs
- Look for /products.json, /api/products, or similar API endpoints before scraping HTML
- Product listing pages usually paginate: look for ?page=N, ?offset=N, or "Load more" buttons
- Always check the total count shown on the page vs what you've extracted
Product Data Checklist
- Name, brand, SKU/ID
- Price (current, original/compare-at, currency)
- Variants (size, color, etc.) with per-variant pricing and availability
- Images (primary + gallery)
- Description (short + long)
- Category/breadcrumb path
- Availability/stock status
- Ratings and review count
Pagination
- Check for: next/prev links, page numbers, "showing X of Y" text
- For infinite scroll: use interact to scroll and load more
- Keep count: if page says "200 products" and you have 24, keep going
JS-Heavy Sites
- Use interact for sites that require JavaScript rendering
- Look for XHR/API calls in the page that return JSON -- often easier than parsing HTML
- Single-page apps often have internal APIs at predictable paths
Site-specific playbooks are in the sites/ directory
- 流狐分类
- 数据
- 作者声明 Agent
- 未找到明确声明;不据此推断已兼容或已测试
- 静态检查
- 88 / 100 · 启发式扫描,不代表运行安全
- 作者 / 版本 / 许可
- @firecrawl · 未声明 license
- 流狐 Token 估算
- 低消耗
- 流狐接入估算
- 即装即用
- 是否需要外部 API Key
- 未发现要求
- 检测到的系统要求
- 未声明
- 底层运行要求
- 未声明
- 检测到的文件与系统行为
-
- 只读
- 检测到的网络行为
- 仅限本地
- 安装命令数
- 无(仅作为资料)
档案由构建时根据 SKILL.md 与安装命令自动衍生,可能与作者实际意图存在差异。
需要注意: 未限定 allowed-tools,默认拥有全部工具权限。
作者没有在当前 SKILL.md 中定义固定输出样例。 Check for sitemap.xml first -- many stores list all product URLs Look for /products.json, /api/products, or similar API endpoints before scraping HTML Product listing pages usually paginate: look for ?page=N, ?offset=N, or "Load more" buttons
Name, brand, SKU/ID Price (current, original/compare-at, currency) Variants (size, color, etc.) with per-variant pricing and availability
Check for: next/prev links, page numbers, "showing X of Y" text For infinite scroll: use interact to scroll and load more Keep count: if page says "200 products" and you have 24, keep going
Use interact for sites that require JavaScript rendering Look for XHR/API calls in the page that return JSON -- often easier than parsing HTML Single-page apps often have internal APIs at predictable paths
Site-specific playbooks are in the sites/ directory
# E-commerce Extraction
## General Patterns
- Check for sitemap.xml first -- many stores list all product URLs
- Look for /products.json, /api/products, or similar API endpoints before scraping HTML
- Product listing pages usually paginate: look for ?page=N, ?offset=N, or "Load more" buttons
- Always check the total count shown on the page vs what you've extracted
## Product Data Checklist
- Name, brand, SKU/ID
- Price (current, original/compare-at, currency)
- Variants (size, color, etc.) with per-variant pricing and availability
- Images (primary + gallery)
- Description (short + long)
- Category/breadcrumb path
- Availability/stock status
- Ratings and review count
## Pagination
- Check for: next/prev links, page numbers, "showing X of Y" text
- For infinite scroll: use interact to scroll and load more
- Keep count: if page says "200 products" and you have 24, keep going
## JS-Heavy Sites
- Use interact for sites that require JavaScript rendering
- Look for XHR/API calls in the page that return JSON -- often easier than parsing HTML
- Single-page apps often have internal APIs at predictable paths
## Site-specific playbooks are in the sites/ directory 作者原文负责流程事实;流狐只索引当前章节、要点、文件与命令。
章节 -> General Patterns → Product Data Checklist → Pagination → JS-Heavy Sites → Site-specific playbooks are in the sites/ directory
要点 -> 原文未标出关键词
文件/命令 -> sitemap.xml · products.js · api/products · SKU/ID · original/compare-at · Category/breadcrumb · Availability/stock · next/prev
内容 SHA-256 -> c9869d491964
原文结构
适用与边界
原文中的明确线索
sitemap.xml、products.js、api/products、SKU/ID、original/compare-at、Category/breadcrumb、Availability/stock、next/prev