Agents Lab

An R&D system that orchestrates specialized AI agents through a software-development workflow, from project discovery to review, on a sandboxed copy of an existing project. The interesting part is not the agents but the guarantees around them.

Status
Working R&D system
Period
2026
Stack
Python · FastAPI · WebSocket · React · TypeScript · OpenAI Agents SDK

Independent R&D project; not a product and not production-ready. The source code is private, and only sanitized results are shown here. Runs and evidence can be walked through during an interview.

Problem

Language-model agents report work they have not done: a file "updated" that never changed, tests "passing" that never ran, a review written without reading the code.

A multi-agent development workflow is only useful if each claim can be checked, if agents cannot act outside their role, and if a run cannot loop or spend without limit.

My role

Independent project. I defined the problem, the workflow, the role boundaries and the guarantees the system enforces, built the system and ran the real validation runs. The decisions below are mine.

Architecture

How a run moves through the system

  1. 01Existing projectSandboxed copy
  2. 02Project discoveryCode, not a model
  3. 03ArchitectProposes the design
  4. 04PlannerOrdered tasks, acceptance criteria
  5. 05Tech LeadChooses task order
  6. 06DeveloperOnly role that writes
  7. 07QARuns approved tests
  8. 08ReviewerReads, approves or rejects
  9. Correction loopQA or the Reviewer can return the work to the Developer for correction, up to a revision limit.
  10. 09CompletedOnly through evidence gates
Correction loopQA or the Reviewer can return the work to the Developer for correction, up to a revision limit.

Across every stage

  • Orchestrator & run stateEnforces stage order and persists the state of the run
  • Tool executionEach role gets only its tools; every call emits an event
  • Request & token budgetsChecked before every model call
  • Workspace isolationChecked paths, no shell, allowlisted test commands
  • Tool-failure recoveryErrors return to the agent; limits end the run cleanly
Conceptual view, simplified. Outlined stages are deterministic code; the others are model-driven agents. The correction loop is available in every run but is not needed in every run.
  • A Python core exposes REST and WebSocket APIs; a React and TypeScript interface shows the run as it happens.
  • Agents run on the OpenAI Agents SDK. The model provider sits behind an adapter; one provider is operational today.
  • The orchestrator owns the run state and decides which role may act. Each agent receives only the tools its role allows: read, list, search, write or run tests.
  • Every tool call goes through a workspace sandbox and emits an event. Events flow to an event store, a projection and, over WebSocket, to the interface.
  • A run never touches the original project: it works on a sandboxed copy.

Agent boundaries

Project discovery
Deterministic code, not an agent. Reads manifests and documentation and produces a project profile for the other roles.
Architect
Proposes the design. Cannot write code.
Planner
Turns the design into ordered tasks with dependencies and acceptance criteria.
Tech Lead
Coordinates and chooses the order of tasks. Never implements; code validates its choice before using it.
Developer
The only role with write access to the workspace.
QA
Runs tests, only from an allowlist of commands defined in code. Cannot edit files.
Reviewer
Read-only. Every finding must quote text that exists in the reviewed file.

Orchestration and state

  • Hybrid orchestration: code enforces the invariants of the workflow (stage order, which role may act, when a run may complete) and the Tech Lead agent only proposes the order of tasks.
  • The state of each run is persisted, and every stage transition and tool call is recorded as an event, so a run can be inspected after it ends.
  • Request and token budgets are checked before every model call, so a call that would exceed the limit is never sent. Each role also has a turn limit, and revisions are capped.
  • Execution and presentation are separate: the interface renders a projection of events and never drives the run.

Failure recovery and feedback loops

  • A failed tool call is returned to the agent as a result, so it can correct course instead of ending the run.
  • Hitting a turn or budget limit ends the run with a structured failure that says which limit was reached, instead of hanging or overspending.
  • QA failures and Reviewer rejections can send the work back to the Developer for correction, up to a revision limit. The run stops if revisions make no progress.
  • Missing evidence is treated as blocked, not as a rejection, so the Developer is never sent to change code that may already be correct.

Key architectural decisions

Lead decision

Evidence gates, enforced in code

The first real run failed because agents reported work they had not done. The fix went into code, not into prompts. A task passes only with evidence from the tools: write receipts with file hashes for the implementation, a real test run whose exit code overrides the QA verdict, and a recorded read of every relevant file for the review. A final gate checks that nothing changed after QA.

  1. Discovery as deterministic code

    Understanding the project is done by code that reads manifests and documentation, not by a model. It is predictable, testable and costs no model calls.

  2. One joint QA and review pass

    QA and review run once after all tasks instead of after each task, because test-writing tasks depend on earlier implementation tasks.

  3. Limits taken from real runs

    Turn limits were set from what real runs actually needed, plus a margin, rather than guessed.

  4. No automatic passes

    A task can count as already satisfied only with a stated reason and real evidence.

Trade-offs

  • Hybrid orchestration gives the Tech Lead less autonomy than a fully agent-managed run, in exchange for invariants an agent cannot argue its way around. A deterministic pipeline and a manager mode were also built; hybrid is the default.
  • Evidence gates are strict: a run can stop for missing evidence even when the work was right. That is preferred over accepting claims that cannot be verified.
  • A joint QA and review pass keeps task dependencies simple, but a problem is found later and can cost a larger revision.
  • Strict budgets and turn limits protect cost and prevent loops, but a limit set too low stops a legitimate run. That happened on a larger project.
  • Tasks run one at a time: state and evidence are easier to reason about, at the cost of speed.
  • No shell and allowlisted test commands make execution safer, but QA can only run what was configured.

Validation and evidence

Offline, automated test suites exercise the real orchestration, tools and test runner with a scripted model. The evidence that matters is a real run on an existing project:

  • Real calls to a language model, with no mocks.
  • Real discovery of the existing project.
  • Real file modifications by the Developer.
  • Real QA test execution.
  • A real review that approved the change.
  • Recovery from a real tool error during the run.
  • End-to-end execution that reached RUN_COMPLETED.

An independent check afterwards repeated the tests and functional checks outside the system and confirmed that the change touched only the expected files.

That run did not need a correction cycle. The QA and review loops are implemented, and real runs have exercised them only partially.

Earlier real runs failed. Those failures are what led to the evidence gates and the turn limits.

Current limitations

  • Successful real runs so far are on a small, controlled project. A pilot on a real monorepo stopped at the architecture stage: discovery only looks one level deep, and the Architect ran out of turns.
  • Only one model provider is operational.
  • An interrupted run cannot be resumed, tasks run sequentially, and there is no git integration.
  • It is a single-user tool that runs on a local machine, with no authentication.