Harness 会话验证

AI 智能 社区 v0.2.0
解读按原文结构重写,命令、链接、术语均保留;右侧可核对作者原始 SKILL.md

方法与流程

  • Workflow:原文以此为独立章节。
  • 4. Design Tests:Each test follows this pattern: For detailed code templates and the makeengine/collect helpers, consult references/test-patterns.md.

适用与边界

  • 当前原文没有单列适用场景。
  • Common Pitfalls:Testing on OpenHarness itself — agent modifies its own running code Using mocks — misses serialization and API compatibility bugs Single-turn only — misses context accumulation and compaction bugs

原文中的明确线索

  • 要点:「Test on an unfamiliar project」、「Use real API calls」、「Multi-turn conversations」、「Combine features」、「Verify tool execution」、「references/test-patterns.md」、「references/feature-matrix.md」
  • 文件与命令maxturns200srtpwdmakeenginecollectreferences/test-patterns.mdAutoAgent

流狐整理:以上内容来自当前 SKILL.md 的章节与原词;未补写作者没有声明的工具、兼容性或能力。

流狐档案 作者与许可取自来源;运行、权限和网络为流狐检测或估算
流狐分类
AI 智能
作者声明 Agent
未找到明确声明;不据此推断已兼容或已测试
静态检查
92 / 100 · 启发式扫描,不代表运行安全
作者 / 版本 / 许可
@HKUDS · v0.2.0 · 未声明 license
流狐 Token 估算
低消耗
流狐接入估算
需简单配置
是否需要外部 API Key
需要 · Anthropic
检测到的系统要求
macOS · Linux · Windows
底层运行要求
Python
检测到的文件与系统行为
  • 只读
  • 允许写入 / 修改
  • Shell 执行
检测到的网络行为
允许外网请求
安装命令数
无(仅作为资料)

档案由构建时根据 SKILL.md 与安装命令自动衍生,可能与作者实际意图存在差异。

需要注意: 未限定 allowed-tools,默认拥有全部工具权限。

输出预览 harness-eval.preview
# 7. Interpret Results

- First check whether `max_turns` was manually set too low
- If `max_turns=200` and the run still fails, the next suspect is wall-clock timeout, not turn count
- Distinguish environment failures from product failures
  - Example: missing dependency in the unfamiliar target repo is not automatically an OpenHarness regression
  - Example: missing `srt`/`bwrap`/`rg` is an eval environment issue

讨论

基于 GitHub Discussions。登录 GitHub 即可参与讨论、点赞、订阅更新。