Bottom line: The WorkBuddy Bench framework validates coding agents across four practical domains with contamination-resistant task construction and full reproducibility through open publication.
Tencent has released WorkBuddy Bench, an open evaluation framework for coding agents across four work domains: repository engineering, web frontend, office workflows, and security. The benchmark is protected against data contamination by reconstructing tasks from actual commits and pull requests and converting them into new prompts.
Architecture and methodology: WorkBuddy Bench consists of four domain-specific subsets: (1) repository-level engineering for code changes, (2) frontend development for web-based tasks, (3) office and business workflows, (4) red-team and blue-team scenarios for security tasks. Each task is not simply adapted from public issue texts, but reverse-engineered from actual commits, pull requests, or business scenarios and reformulated as a brief, colloquial request. This makes it impossible to rediscover the original source through web search.
Contamination resistance through transparency: Rather than achieving security through secrecy, the benchmark relies on complete disclosure: task directories, environment images, evaluation harness, tests, and reference solutions are publicly available. Contamination resistance is based on the construction methodology combined with explicit data versioning. This makes the benchmark end-to-end reproducible and directly verifiable — any party can re-execute tasks and inspect contents.
Evaluation and results: The framework runs under a uniform, reproducible protocol on two agent harnesses: CodeBuddy Code and Claude Code. Each of the four subsets uses its own scoring instrument, which means that scores are not directly comparable across domains. The benchmark publishes a cross-model leaderboard across multiple model families, but deliberately refrains from providing a suite-wide average.
Source: arxiv.org · Published July 22, 2026
Lumi AI News — AI-assisted curation in accordance with Article 50 EU AI Act. Paraphrase and classification by Lumi News Pipeline v1.7.3.