Claude Opus required 60 hours and repeated human guidance to discover vulnerabilities in HAWK and AES variants—a proof-of-concept for LLM-assisted security research, but only at substantial cost and resource expense.
Arbor enables AI-driven research through systematic hypothesis management and achieved an average of 2.5x higher improvements than existing code models on six test tasks.
Arbor coordinates autonomous AI agents via persistent hypothesis trees and achieved 2.5× better results than Codex and Claude Code on six research tasks.
Multi-agent coordination with task decomposition and parallelization substantially improves computer-use agents and solves complex long-horizon tasks where single agents fail.