Industrialized agentic development pipeline
Engineer first: AI executes; I design, decide, review and prove.
Key figures
- testing campaigns
- 5
- test cases run
- 800+
- agents in parallel, at peak
- 11
Context
Delivering a complete information system single-handedly, at the quality level of a full team. Generating code is not enough: every change has to be scoped, reviewed and proven.
My role
I scope, orchestrate, review every change and keep the final say, fully autonomously.
Solution
Everything starts with a spec-driven workflow: each need becomes a versioned roadmap, split into deliverables with proof and verification status. Every repository carries its own AGENTS.md contract. Execution is orchestrated by reusable skills, subagents whose model and effort are calibrated to risk, and parallel worktrees.
Proof comes from multi-agent acceptance testing in a real browser, via MCP: verdicts are re-measured and a defect log links every bug to its fix. Knowledge is captured in a self-hosted long-term memory, fed by hooks. This pipeline delivered the e-commerce platform, SSO across 9 applications and the supplier import running in production.
Key features
Versioned specs and roadmaps
Each need becomes a roadmap split into deliverables, with proof and verification status.
Multi-agent browser testing
Up to 11 subagents run acceptance tests in a real browser, via MCP.
Re-measured verdicts
No agent verdict is taken at face value: each one is re-measured, and traceability is checked by script.
Results in production
E-commerce platform, SSO across 9 applications and supplier import, all delivered with this pipeline.
Engineering challenges
- 1
The engineer keeps the final say
A validated plan before any code, a review of every change: nothing gets merged unless it is understood.
- 2
Risk-calibrated agents
Each subagent's model and effort level are chosen according to what is at stake.
- 3
Failing test first
Every fix starts with a spec and a test that reproduces the defect before it is corrected.
- 4
Automatic knowledge capture
Hooks feed the long-term memory, shared by Claude Code and Codex.
Tech stack
- AI
- Claude CodeCodexMCP (Chrome DevTools)SkillsHooksSubagentsHindsight
- Infrastructure
- Git worktreesDocker
- Quality
- Python
A project of this scale?
Let's talk about your context: I'll tell you frankly what is feasible, and how long it takes.