AI agents for real engineering work.
How to build and evaluate them. What fails in practice. And what it takes to make the results trustworthy.
Latest essay
The Distance to the Edge
Why organisational AI capability is nonlinear and why the edge is better understood as a learning regime than a position on a maturity ladder.
Read the essay →Where to start
Understand the harness
How tools, verifiers and control flow turn model capability into useful engineering work.
Harness engineering →Evaluate an agent
Start with measured results, then examine what scores miss and what a trustworthy task needs.
Agent evaluation →Govern organisational use
Bring agents into architecture, engineering and construction with room to experiment and clear limits on authority.
AI in AEC →Recent writing
All 16 essays →-
A World Worth Learning From
Before an agent can learn from experience, its environment has to produce experience worth learning from. A sixteen-run engineering study tested what survives after the agent commits.
-
Broad Creation, Narrow Authority
An open-ended approach to AI-enabled software: let practitioners explore, embed controls in the platform, and govern the moment an experiment acquires organisational consequence.
-
The Attacker Moves Second. So Did I.
My benchmark optimiser found the same seam an adversary would: it rewrote the world its own grader consumed. A design note on provenance ledgers—agent memory where authority comes from evidence, not persuasion.
-
Fluent, But Unsafe
How 150 supposedly finished tasks and perfect model scores hid a weak engineering benchmark—and how auditable reviews exposed what the numbers missed.
-
Mediation, Not Intermediation
Why the 'fix your foundations before AI' message has it backwards: agentic workflows are the way out of legacy data, and governance worth having is co-designed from practice, not committees.