Verified Success Rate
accepted tasks / all tasksEvidence before confidence
How VOLY separates an agent’s completion claim from an accepted engineering result.
Take the survey04 measures →
The core economic unit is neither run cost nor generated code volume. It is the cost of a verified completed task.
accepted tasks / all tasksreported done, but rejectedtotal cost / accepted tasksruns with unrequested changesResearch hypotheses
A multi-agent workflow is not better by definition. It must demonstrate an advantage in quality, result cost, or error recovery.
Reported success is systematically higher than the result verified by tests, diff, and code review.
Architect, developer, and reviewer roles must prove an advantage over one general-purpose agent.
Plan gates, protected paths, and file limits are tested as controls against unrequested changes.
Architecture ≠ outcome
Each architecture receives the same starting commit, task brief, time budget, and pre-frozen acceptance criteria.
1–2 min survey · 30 min interview
Which AI coding tools does the team use, how often, and for which kinds of tasks?
What counts as done: an agent response, diff, tests, review, merge, or a production metric?
How often does the agent say done, change too much, or create a regression?
Can the team see task cost, and what happens when a quota runs out?
What evidence is required to accept a result without fully repeating the work by hand?
Editorial honesty
In our experiment, N of M tasks passed the acceptance criteria.
On the selected test bank, the role workflow showed…
According to participant self-reports…
Most companies… based on a non-representative sample.
Multi-agent is always better and cheaper.
Present a roadmap or outside figure as a VOLY result.