|
|
8 mesiacov pred | |
|---|---|---|
| .. | ||
| agents | 8 mesiacov pred | |
| framework | 8 mesiacov pred | |
| results | 8 mesiacov pred | |
| test_tmp | 8 mesiacov pred | |
| ARCHITECTURE.md | 8 mesiacov pred | |
| GETTING_STARTED.md | 8 mesiacov pred | |
| HOW_TESTS_WORK.md | 8 mesiacov pred | |
| README.md | 8 mesiacov pred | |
Comprehensive SDK-based evaluation framework for testing OpenCode agents with real execution, event streaming, and automated violation detection.
cd evals/framework
npm install
npm run build
# Run all tests (free model by default)
npm run eval:sdk
# Run specific agent
npm run eval:sdk -- --agent=openagent
npm run eval:sdk -- --agent=opencoder
# View results dashboard
cd ../results && ./serve.sh
📖 New to the framework? Start with GETTING_STARTED.md
| Agent | Tests | Pass Rate | Status |
|---|---|---|---|
| OpenAgent | 22 tests | 100% | ✅ Production Ready |
| Opencoder | 4 tests | 100% | ✅ Production Ready |
✅ Context Loading Tests - 5 comprehensive tests (3 simple, 2 complex multi-turn)
✅ Smart Timeout System - Activity monitoring with absolute max timeout
✅ Fixed Context Evaluator - Properly detects context files in multi-turn sessions
✅ Batch Test Runner - Run tests in controlled batches to avoid API limits
✅ Results Dashboard - Interactive web dashboard with filtering and charts
evals/
├── framework/ # Core evaluation framework
│ ├── src/
│ │ ├── sdk/ # SDK-based test runner
│ │ ├── collector/ # Session data collection
│ │ ├── evaluators/ # Rule violation detection
│ │ └── types/ # TypeScript types
│ ├── docs/ # Framework documentation
│ ├── scripts/utils/run-tests-batch.sh # Batch test runner
│ └── README.md # Framework docs
│
├── agents/ # Agent-specific test suites
│ ├── openagent/ # OpenAgent tests
│ │ ├── tests/
│ │ │ ├── context-loading/ # Context loading tests (NEW)
│ │ │ ├── developer/ # Developer workflow tests
│ │ │ ├── business/ # Business analysis tests
│ │ │ └── edge-case/ # Edge case tests
│ │ ├── CONTEXT_LOADING_COVERAGE.md
│ │ ├── IMPLEMENTATION_SUMMARY.md
│ │ └── README.md
│ │
│ ├── opencoder/ # Opencoder tests
│ │ ├── tests/developer/
│ │ └── README.md
│ │
│ └── shared/ # Shared test utilities
│
├── results/ # Test results & dashboard
│ ├── history/ # Historical results (60-day retention)
│ ├── index.html # Interactive dashboard
│ ├── serve.sh # One-command server
│ ├── latest.json # Latest test results
│ └── README.md
│
├── test_tmp/ # Temporary test files (auto-cleaned)
│
├── GETTING_STARTED.md # Quick start guide (START HERE)
├── HOW_TESTS_WORK.md # Detailed test execution guide
├── ARCHITECTURE.md # System architecture review
└── README.md # This file
@opencode-ai/sdk for real agent interactionopencode/grok-code-fast (OpenCode Zen)--model=provider/model| Document | Purpose | Audience |
|---|---|---|
| GETTING_STARTED.md | Quick start guide | New users |
| HOW_TESTS_WORK.md | Test execution details | Test authors |
| ARCHITECTURE.md | System architecture | Developers |
| framework/SDK_EVAL_README.md | Complete SDK guide | All users |
| framework/docs/test-design-guide.md | Test design philosophy | Test authors |
| agents/openagent/CONTEXT_LOADING_COVERAGE.md | Context loading tests | OpenAgent users |
| agents/openagent/IMPLEMENTATION_SUMMARY.md | Recent implementation | Developers |
| Feature | OpenAgent | Opencoder |
|---|---|---|
| Approval | Text-based + tool permissions | Tool permissions only |
| Workflow | Analyze→Approve→Execute→Validate | Direct execution |
| Context | Mandatory before execution | On-demand |
| Test Style | Multi-turn (approval flow) | Single prompt |
| Timeout | 300s (smart timeout) | 60s (standard) |
# All tests with free model
npm run eval:sdk
# Specific category
npm run eval:sdk -- --pattern="context-loading/*.yaml"
# Custom model
npm run eval:sdk -- --model=anthropic/claude-3-5-sonnet-20241022
# Debug single test
npm run eval:sdk -- --pattern="ctx-simple-coding-standards.yaml" --debug
# Batch execution (avoid API limits)
./scripts/utils/run-tests-batch.sh openagent 3 10
# Interactive dashboard (one command!)
cd results && ./serve.sh
# View JSON
cat results/latest.json
# Historical results
ls results/history/2025-11/
# Example: context-loading/my-test.yaml
id: my-test-001
name: "My Test"
description: What this test validates
category: developer
agent: openagent
model: anthropic/claude-sonnet-4-5
prompt: "Your test prompt here"
behavior:
mustUseTools: [read]
requiresContext: true
minToolCalls: 1
expectedViolations:
- rule: context-loading
shouldViolate: false
severity: error
approvalStrategy:
type: auto-approve
timeout: 60000
tags:
- context-loading
See GETTING_STARTED.md for more examples.
serve.sh)# Behavior expectations (what agent should do)
behavior:
mustUseTools: [read, write] # Required tools
mustUseAnyOf: [[bash], [list]] # Alternative tools
requiresApproval: true # Must ask for approval
requiresContext: true # Must load context
minToolCalls: 2 # Minimum tool calls
# Expected violations (what rules to check)
expectedViolations:
- rule: approval-gate
shouldViolate: false # Should NOT violate
severity: error
- rule: context-loading
shouldViolate: false
severity: error
Context Loading Tests (5 tests, 100% passing)
Smart Timeout System
Fixed Context Loading Evaluator
tool.data.state.input.filePath)Batch Test Runner
run-tests-batch.sh scriptResults Dashboard
✅ Full SDK integration with @opencode-ai/sdk@1.0.90
✅ Real-time event streaming (12+ events per test)
✅ 5 evaluators integrated and working
✅ YAML-based test definitions with Zod validation
✅ CLI runner with detailed reporting
✅ Free model by default (no API costs)
✅ Model-agnostic test design
✅ Both positive and negative test support
✅ Smart timeout with activity monitoring
✅ Context loading validation (100% coverage)
✅ Results tracking and visualization
✅ Batch execution support
Status: ✅ Production-ready for OpenAgent & Opencoder evaluation
See ../docs/contributing/CONTRIBUTING.md
MIT
Last Updated: 2025-11-26
Framework Version: 0.1.0
Test Coverage: 26 tests (22 OpenAgent, 4 Opencoder)
Pass Rate: 100%