VOLY open practical research

The AI agent said “done.”
But is the task complete?

We compare more than persuasive model output. We study results a team can verify and accept: git diff, tests, scope compliance, manual fixes, and total cost.

Real cases · controlled experiments · independent verification

30 reproducible tasks

What the test bank contains

Thirty tasks span six engineering disciplines. These are not tutorial exercises: each task runs in a fixed environment and ends in a verifiable result.

B01–B05

Backend

API, idempotency, concurrency, DB pool

F01–F05

Frontend

state, accessibility, forms, responsive UI

T01–T05

Testing

regression, flaky tests, properties, queues

D01–D05

DevOps

Docker, healthcheck, CI cache, rollback

S01–S05

Security

traversal, secrets, commands, protected paths

X01–X05

Cross-stack

DB → API → UI, flags, timezones, audit

Every task defines its starting state, allowed scope, reference commit, verification commands, and independent acceptance.

MethodologyReady
Test bank30 tasks
PilotIn progress
ResultsPending

Why this matters

“Done” can mean five different things.

The agent reported success. Code changed. Tests passed. Requirements were met without scope drift. A human or independent check accepted the result—these events are not interchangeable.

VOLY studies where task meaning is lost, where the executor falls short, and where the verification process fails to detect it.

Signals from the field

What engineering teams are debating now

@hrswatigupta · XAug 10, 2026

“Don’t prompt Claude. Build a system that prompts itself.” Attention is moving from a single chat to agent systems.

VOLYWe test conditions where specialization and independent verification produce a measurable gain—not whether multi-agent is popular.

@anupamrjp · XAug 3, 2026

Claude Code users discuss usage limits and manual workarounds: multiple accounts, proxy CLIs, and alternative harnesses.

VOLYThis supports the Job: finish a task without quota failure, with transparent fallback and total result cost.

@ClaudeDevs · XAug 8, 2026

Native Managed Agents already introduce session budgets; basic orchestration is rapidly becoming a platform capability.

VOLYVOLY differentiates on cross-agent task budgets across tools and fallback boundaries—not roles alone.

@PrimeIntellect · XAug 8, 2026

Large multi-agent systems increase the need for observable environments, bounded roles, and verifiable episodes.

VOLYCollect trustworthy episodes and role metrics first; defer self-play until the evidence is reliable.

Metrics are snapshots of public activity when checked. They provide market context, not VOLY research results.

Choose your depth

One question.
Three paths.

If you have two minutes, start with the survey. To inspect the claims, open the methodology or the test bank.