All projects
AI

Industrialized agentic development pipeline

Engineer first: AI executes; I design, decide, review and prove.

Key figures

testing campaigns
5
test cases run
800+
agents in parallel, at peak
11

Context

Delivering a complete information system single-handedly, at the quality level of a full team. Generating code is not enough: every change has to be scoped, reviewed and proven.

My role

I scope, orchestrate, review every change and keep the final say, fully autonomously.

Solution

Everything starts with a spec-driven workflow: each need becomes a versioned roadmap, split into deliverables with proof and verification status. Every repository carries its own AGENTS.md contract. Execution is orchestrated by reusable skills, subagents whose model and effort are calibrated to risk, and parallel worktrees.

Proof comes from multi-agent acceptance testing in a real browser, via MCP: verdicts are re-measured and a defect log links every bug to its fix. Knowledge is captured in a self-hosted long-term memory, fed by hooks. This pipeline delivered the e-commerce platform, SSO across 9 applications and the supplier import running in production.

Key features

  • Versioned specs and roadmaps

    Each need becomes a roadmap split into deliverables, with proof and verification status.

  • Multi-agent browser testing

    Up to 11 subagents run acceptance tests in a real browser, via MCP.

  • Re-measured verdicts

    No agent verdict is taken at face value: each one is re-measured, and traceability is checked by script.

  • Results in production

    E-commerce platform, SSO across 9 applications and supplier import, all delivered with this pipeline.

Engineering challenges

  1. 1

    The engineer keeps the final say

    A validated plan before any code, a review of every change: nothing gets merged unless it is understood.

  2. 2

    Risk-calibrated agents

    Each subagent's model and effort level are chosen according to what is at stake.

  3. 3

    Failing test first

    Every fix starts with a spec and a test that reproduces the defect before it is corrected.

  4. 4

    Automatic knowledge capture

    Hooks feed the long-term memory, shared by Claude Code and Codex.

Tech stack

AI
Claude CodeCodexMCP (Chrome DevTools)SkillsHooksSubagentsHindsight
Infrastructure
Git worktreesDocker
Quality
Python

A project of this scale?

Let's talk about your context: I'll tell you frankly what is feasible, and how long it takes.