Harness 会话验证
- 作者仓库星标 13,484
- 作者仓库 OpenHarness
Harness Eval — End-to-End Feature Validation
Validate OpenHarness features by running real agent loops against an unfamiliar codebase with actual LLM API calls. Every test exercises the full stack: API client → model → tool calls → execution → result.
Core Principles
- Test on an unfamiliar project — never test on OpenHarness itself (the agent modifies its own code). Clone a real project as the workspace.
- Use real API calls — no mocks. Configure a real LLM endpoint.
- Multi-turn conversations — always test 2+ turns where the model needs prior context.
- Combine features — test hooks+skills+agent loop together, not in isolation.
- Verify tool execution — inspect tool call lists and output files, not just model text.
Workflow
1. Prepare Workspace
Clone an unfamiliar project (do not use OpenHarness):
git clone https://github.com/HKUDS/AutoAgent /tmp/eval-workspace
2. Configure Environment
export ANTHROPIC_API_KEY=sk-xxx
export ANTHROPIC_BASE_URL=https://api.moonshot.cn/anthropic # or any provider
export ANTHROPIC_MODEL=kimi-k2.5
For long-running real evals, do not artificially lower max_turns. Use the product default (200) unless the user explicitly wants a tighter bound.
3. Prepare Real Sandbox Runtime When Relevant
If the task is validating sandbox behavior, install and verify the actual runtime before running agent loops:
npm install -g @anthropic-ai/sandbox-runtime
sudo apt-get update
sudo apt-get install -y bubblewrap ripgrep
which srt
which bwrap
which rg
srt --version
Then run a minimal smoke check through OpenHarness, not just raw srt, so you verify the real adapter path:
from pathlib import Path
from openharness.config.settings import Settings, SandboxSettings, save_settings
from openharness.tools.bash_tool import BashTool
cfg = Path("/tmp/openharness-sandbox-settings.json")
save_settings(Settings(sandbox=SandboxSettings(enabled=True, fail_if_unavailable=True)), cfg)
# Point config loader at this file, then run BashTool on a tiny command such as `pwd`.
If sandbox dependencies are missing, treat that as an environment/setup failure, not a feature regression.
4. Design Tests
Each test follows this pattern:
engine = make_engine(system_prompt="...", cwd=UNFAMILIAR_PROJECT)
evs1 = [ev async for ev in engine.submit_message("Read X, analyze Y")]
r1 = collect(evs1) # text, tools, turns, tokens
evs2 = [ev async for ev in engine.submit_message("Based on what you found...")]
r2 = collect(evs2)
assert "grep" in r1["tools"] # verify tools ran
For detailed code templates and the make_engine/collect helpers, consult references/test-patterns.md.
5. Prefer Long-Horizon, Real Agent Loops
For meaningful end-to-end validation, prefer unfamiliar-repo tasks that force multiple turns, context reuse, and mixed tool usage.
Recommended pattern:
- Use a real external workspace such as
AutoAgent - Use real provider credentials and the actual target model
- Keep
max_turns=200 - Use per-prompt timeouts large enough for real exploration, such as
240-600s - Require at least 2 turns per scenario
- Verify both text quality and tool traces
- Keep polling long-running sessions until they finish; do not abandon a run after the first long pause
Recommended long-horizon scenarios:
architecture_multiturn- Turn 1: map architecture, shell/subprocess surfaces, and test entrypoints
- Turn 2: identify top risks and propose refactors
- Turn 3: condense into onboarding or remediation actions
- Success:
bash,glob,grep,read_fileall appear; no timeout; noMaxTurnsExceeded
hook_block_and_recover- Force the model to try
bash - Block it with a real pre-tool hook
- Verify the model adapts with
glob/grep/read_file
- Force the model to try
sandbox_multiturn- Enable real sandbox settings with
fail_if_unavailable=true - First prompt must start with exactly one shell command such as
pwd && ls -la - Second prompt must explicitly reuse the prior shell findings
- Success:
bashexecutes via sandbox, non-shell tools continue the task, and the agent recovers from incidental repo errors
- Enable real sandbox settings with
When a scenario fails, classify it before changing code:
MaxTurnsExceeded: likely eval harness misconfiguration ifmax_turnswas manually loweredtimeout: task is too broad or per-prompt timeout is too small- sandbox unavailable: environment missing
srt,bwrap, orrg - tool error with task still completed: feature may still be healthy; inspect recovery behavior
6. Run Tests
python tests/test_merged_prs_on_autoagent.py # PR feature tests
python tests/test_real_large_tasks.py # large multi-step tasks
python tests/test_hooks_skills_plugins_real.py # hooks/skills/plugins
python -m pytest tests/ -q -k "not autoagent" # unit tests (no API)
For ad hoc long-horizon validation, it is acceptable to run a temporary Python driver script as long as it:
- uses real OpenHarness engine/tool objects
- targets an unfamiliar repository
- prints per-scenario JSON summaries
- records tools, errors, turns, and token usage
- stays attached until completion
7. Interpret Results
| Result | Meaning | Action |
|---|---|---|
| PASS with tool calls | Feature works end-to-end | Done |
| PASS without tool calls | Model answered from knowledge | Rewrite prompt to force tool use |
| FAIL with exception | Code bug | Read traceback |
| FAIL with wrong output | Model behavior issue | Check system prompt and tool schemas |
| Timeout | Task too complex | Increase max_turns or simplify prompt |
For long-running real evals, refine the timeout guidance:
- First check whether
max_turnswas manually set too low - If
max_turns=200and the run still fails, the next suspect is wall-clock timeout, not turn count - Distinguish environment failures from product failures
- Example: missing dependency in the unfamiliar target repo is not automatically an OpenHarness regression
- Example: missing
srt/bwrap/rgis an eval environment issue
Feature Coverage Checklist
- Engine: multi-turn memory, tool chaining, parallel tools, error recovery, auto-compaction
- Swarm: InProcessBackend lifecycle, concurrent teammates, coordinator+notifications
- Hooks: pre_tool_use blocking → model adapts, post_tool_use firing
- Skills: skill tool invocation → model follows instructions
- Plugins: plugin-provided skill loaded and used in agent loop
- Memory: YAML frontmatter parsing, body content search, context injection
- Session: save → load → resume with context preserved
- Providers: Anthropic client, OpenAI client (with reasoning_content), multi-turn
- Cost: token accumulation across turns
Common Pitfalls
- Testing on OpenHarness itself — agent modifies its own running code
- Using mocks — misses serialization and API compatibility bugs
- Single-turn only — misses context accumulation and compaction bugs
- Artificially lowering
max_turnsduring real evals — can create false failures that do not reflect product defaults - Not checking tool call list — model may claim tool use without calling it
- Hardcoding paths — use
WORKSPACEvariable, skip in CI withpytest.mark.skipif - Declaring sandbox “tested” after only checking raw
srt— verify the OpenHarness adapter path too - Abandoning long tasks too early — some real tasks pause for minutes before the next event arrives
Additional Resources
Reference Files
references/test-patterns.md— Complete code templates formake_engine,collect, and each feature categoryreferences/feature-matrix.md— Detailed test cases for every OpenHarness module
Existing Test Files
Working test suites in the repo:
tests/test_merged_prs_on_autoagent.py— PR feature validationtests/test_real_large_tasks.py— Large multi-step taskstests/test_hooks_skills_plugins_real.py— Hooks/skills/plugins in agent loopstests/test_untested_features.py— Module-level integration tests
- 流狐分类
- AI 智能
- 作者声明 Agent
- 未找到明确声明;不据此推断已兼容或已测试
- 静态检查
- 92 / 100 · 启发式扫描,不代表运行安全
- 作者 / 版本 / 许可
- @HKUDS · v0.2.0 · 未声明 license
- 流狐 Token 估算
- 低消耗
- 流狐接入估算
- 需简单配置
- 是否需要外部 API Key
- 需要 · Anthropic
- 检测到的系统要求
- macOS · Linux · Windows
- 底层运行要求
- Python
- 检测到的文件与系统行为
-
- 只读
- 允许写入 / 修改
- Shell 执行
- 检测到的网络行为
- 允许外网请求
- 安装命令数
- 无(仅作为资料)
档案由构建时根据 SKILL.md 与安装命令自动衍生,可能与作者实际意图存在差异。
需要注意: 未限定 allowed-tools,默认拥有全部工具权限。
# 7. Interpret Results
- First check whether `max_turns` was manually set too low
- If `max_turns=200` and the run still fails, the next suspect is wall-clock timeout, not turn count
- Distinguish environment failures from product failures
- Example: missing dependency in the unfamiliar target repo is not automatically an OpenHarness regression
- Example: missing `srt`/`bwrap`/`rg` is an eval environment issue Test on an unfamiliar project — never test on OpenHarness itself (the agent modifies its own code). Clone a real project as the workspace. Use real API calls — no mocks. Configure a real LLM endpoint. Multi-turn conversations — always test 2+ turns where the…
Workflow
Clone an unfamiliar project (do not use OpenHarness):
For long-running real evals, do not artificially lower maxturns. Use the product default (200) unless the user explicitly wants a tighter bound.
If the task is validating sandbox behavior, install and verify the actual runtime before running agent loops: Then run a minimal smoke check through OpenHarness, not just raw srt, so you verify the real adapter path: If sandbox dependencies are missing, treat…
Each test follows this pattern: For detailed code templates and the makeengine/collect helpers, consult references/test-patterns.md.
# Harness Eval — End-to-End Feature Validation
Validate OpenHarness features by running real agent loops against an unfamiliar codebase with actual LLM API calls. Every test exercises the full stack: API client → model → tool calls → execution → result.
## Core Principles
1. **Test on an unfamiliar project** — never test on OpenHarness itself (the agent modifies its own code). Clone a real project as the workspace.
2. **Use real API calls** — no mocks. Configure a real LLM endpoint.
3. **Multi-turn conversations** — always test 2+ turns where the model needs prior context.
4. **Combine features** — test hooks+skills+agent loop together, not in isolation.
5. **Verify tool execution** — inspect tool call lists and output files, not just model text.
## Workflow
### 1. Prepare Workspace
Clone an unfamiliar project (do not use OpenHarness):
```bash
git clone https://github.com/HKUDS/AutoAgent /tmp/eval-workspace
```
### 2. Configure Environment
```bash
export ANTHROPIC_API_KEY=sk-xxx
export ANTHROPIC_BASE_URL=https://api.moonshot.cn/anthropic # or any provider
export ANTHROPIC_MODEL=kimi-k2.5
```
For long-running real evals, do not artificially lower `max_turns`. Use the product default (`200`) unless the user explicitly wants a tighter bound.
### 3. Prepare Real Sandbox Runtime When Relevant
If the task is validating sandbox behavior, install and verify the actual runtime before running agent loops:
```bash
npm install -g @anthropic-ai/sandbox-runtime
sudo apt-get update
sudo apt-get install -y bubblewrap ripgrep
which srt
which bwrap
which rg
srt --version
```
Then run a minimal smoke check through OpenHarness, not just raw `srt`, so you verify the real adapter path:
```python
from pathlib import Path
… 作者原文负责流程事实;流狐只索引当前章节、要点、文件与命令。
章节 -> Core Principles → Workflow → 1. Prepare Workspace → 2. Configure Environment → 3. Prepare Real Sandbox Runtime When Relevant → 4. Design Tests
要点 -> Test on an unfamiliar project · Use real API calls · Multi-turn conversations · Combine features · Verify tool execution · references/test-patterns.md · references/feature-matrix.md
文件/命令 -> maxturns · 200 · srt · pwd · makeengine · collect · references/test-patterns.md · AutoAgent
内容 SHA-256 -> 771a9ad41fcf
方法与流程
适用与边界
原文中的明确线索
maxturns、200、srt、pwd、makeengine、collect、references/test-patterns.md、AutoAgent