M1Spec Join the waitlist

White paper

Executable Intent

Managing Trust, Verification Costs, and Architectural Drift in Generative AI Engineering

AI can write code faster than any of us. That part is solved. What isn’t solved is trusting it. This paper follows the gap between how much teams use AI and how little they trust it, shows where “just describe it” breaks down, and makes the case that the answer isn’t less AI but more structure: say what you want in plain language, enforce it with rules a machine can check, and treat the generated code as the part that can be replaced.

Start reading

01

Everyone uses it. Hardly anyone trusts it.

1.1 A tool we can’t stop using — and can’t stop checking

Picture a Tuesday morning in almost any engineering team. Someone asks their coding assistant for a new endpoint, and thirty seconds later it’s there: clean, well named, with tests. A year ago that took an afternoon. Then the pull request lands in review, and a senior engineer spends the rest of the morning working out whether it’s actually right.

That scene is the whole story of AI in software engineering right now, and the numbers say the same thing. In Google’s DORA 2025 study, 90% of technology professionals use AI at work, more than 80% say it has made them more productive, and 59% see a positive effect on the quality of their code. The 2025 Stack Overflow Developer Survey agrees: 84% of developers use AI tools or plan to, and 51% of professional developers use them every day.

How much we use it

AI workplace usage DORA 202590%
Productivity gain DORA 2025>80%
Positive quality perception DORA 202559%
Using or planning to use Stack Overflow 202584%
Daily professional use Stack Overflow 202551%

How much we trust it

Distrust output accuracy Stack Overflow 202546%
Trust output accuracy Stack Overflow 202533%
Refrain from unconstrained “vibe coding” Stack Overflow 202572%

But ask the same people whether they trust what the tools produce, and the picture flips:

  • More doubt than trust. In the same Stack Overflow survey, 46% of developers distrust the accuracy of AI output. Only 33% trust it.
  • Hands stay on the wheel. 72% of professional developers stay away from unconstrained “vibe coding” — describing what they want in plain language and shipping whatever comes back, without reading it line by line.

So we have a strange situation. Teams adopt AI because it drafts code fast and the code looks good. Then they keep tight manual control over it, because nobody is sure it handles the edge cases or fits the architecture. The assistant removes the chore of writing the first draft — and hands the team a new chore: reading it.

When writing code costs almost nothing, writing is no longer the bottleneck. Checking is.

1.2 Cheap to write, expensive to trust

For most of software’s history, writing code was the expensive part. It’s what senior engineers spent their days on. Generative AI turns that upside down: a first draft now costs next to nothing. What it doesn’t make cheaper is knowing whether that draft is correct. If anything, that gets more expensive.

The generation–verification gap

Generation cost → approaches zero Verification cost → escalates AI adoption / code volume →

Conceptual illustration.

The reason is the particular way these models fail. They are rarely wildly wrong; they are almost right. In the 2025 Stack Overflow survey, 66% of developers named “almost right, but not quite” as their biggest frustration, and 45% said debugging AI-generated code takes longer than writing it themselves. Tidy, plausible code is exactly the kind that hides an off-by-one, a missed boundary condition or a security hole. To find it, the reviewer has to rebuild in their head all the context the model never had — for code they didn’t write.

Multiply that by a CI pipeline full of generated pull requests and the economics get uncomfortable. If checking a generated change takes a senior engineer longer than writing it would have, the team has paid twice and gained nothing.

DORA’s follow-up research in 2026 points the same way at the level of whole delivery pipelines: more AI means more throughput, but also more instability. Left unconstrained, generation quietly moves engineers from designing systems to auditing them and fighting fires.

What generation speeds upWhat verification has to pay for
Cost per lineClose to zero — the model writes it.Rising — someone has to rebuild the context and read it.
FrictionAlmost none; a prompt becomes a draft in seconds.A lot; the defects hide in code that is almost right.
When problems showLate; elegant code sails through a quick look.In production, as runtime failures or silent architectural drift.
Effect on deliveryMore pull requests, more throughput.More instability, and senior time pulled from design into review.

That gap between fast writing and slow checking is where the real failures start.

02

Where “just describe it” breaks down

2.1 One sentence, a dozen possible systems

More and more, the way we build software is to describe what we want in plain language and let a model write the implementation. Call it Implementation-as-Spec. It’s a wonderful interface — until you notice how much a sentence leaves out.

Here is a perfectly normal request to a coding assistant:

“Orders should be processed asynchronously and reliably.”

A product manager reads that and nods. A distributed-systems architect reads it and starts asking questions — because every one of them changes the code that gets written:

  • Delivery: at-least-once with idempotent consumers, or genuinely exactly-once with distributed transactions?
  • Ordering: must one customer’s orders be handled strictly in sequence, or can they arrive out of order across partitions?
  • Retries: what backoff formula, how much jitter, and after how many attempts does a message go to the dead-letter queue?
  • Poison messages: how is a payload nobody can parse set aside, flagged and fixed without blocking everything behind it?
  • Duplicates: how are repeated submissions recognised and dropped across distributed stores when the network splits?
  • Service levels: what latency is acceptable (say, p99 under 150 ms), how fast must it recover, and which W3C trace headers must travel with each message?

The model will answer all of these questions — silently, by picking something. The only way to make sure it picks what you need is to write the decision down in a form a machine can check. For example, the rule that order logic must never talk to infrastructure directly can be an executable test rather than a hope:

The same intent, written as a rule a build can enforce (ArchUnit)
@ArchTest
static final ArchRule enforce_asynchronous_order_isolation = classes()
    .that().resideInAPackage("..order.domain..")
    .should().onlyDependOnClassesThat()
    .resideInAnyPackage("..order.domain..", "java..")
    .because("Domain logic must not directly invoke infrastructure messaging or DB drivers");

This is the heart of the problem. A compiler is boring in the best way — the same input gives the same output, every time:

Source Code + Deterministic Compiler = Invariant Machine Artifact

A language model is not:

Ambiguous Prompt + LLM Context Window = Probabilistic Candidate Implementation

Change the prompt slightly, the files in context, the tools, or the model version, and you get a different system. A prompt on its own can’t be the definition of how software behaves. Without hard edges around it, the vagueness leaks straight into the design.

2.2 Passing the tests, losing the architecture

Now give that assistant a task and a test suite. Its goal, in practice, is simple: change the code until the tests pass. It sees the files in its context window, not the shape of the whole system or the reasons it looks the way it does. And very often the quickest way to green tests is a shortcut that quietly damages the architecture.

A fast assistant with a narrow viewIts goal: “make the tests pass”
Boundary violationsReaching into the database from the UI or domain layer
Duplication and couplingCopying logic across services and tying them together at runtime
Governance bypassesHard-coded secrets, personal data in logs, compliance ignored

The damage comes in three familiar shapes:

  1. Boundary violations. To make a data-fetching test pass, the assistant imports a database driver straight into a domain entity — skipping the layers the team built on purpose.
  2. Duplication and coupling. Instead of finding the shared concept, it copies the logic into another service. Now two services share a hidden dependency, a schema neither owns, and assumptions about timing nobody wrote down.
  3. Governance bypasses. The easiest path to a passing test may ignore data-residency rules, send unencrypted personal data to the log aggregator, or hard-code a credential.

None of this changes who is responsible. An assistant may write 90% of a pull request, but when it breaks something, the accountability is still entirely human. More generated code doesn’t mean less ownership — it means we need clearer ownership, a full audit trail, and automatic checks that stop debt before it piles up.

That is why a counter-movement has been growing: put the architecture itself into code, and let machines guard it.

03

Turning architecture into guardrails

3.1 From advice to constraint

For a long time, architecture was advice. A diagram on a whiteboard, a wiki page that slowly went stale, a review board that met on Thursdays. People followed it when they remembered to. Architecture as Code changes that: the rules become executable, and the build enforces them whether anyone remembers or not.

How it used to work

Architecture as advice
  • Diagrams on a stale wiki
  • Subjective review boards
  • Rules nobody enforces

Architecture as Code

Architecture as constraint
  • Machine-readable models
  • Policies checked in CI/CD
  • Versioned, reviewable rules

It goes well beyond Infrastructure as Code:

  • Models as code. Tools like Structurizr keep the C4 model as text in the repository, so the system’s structure is versioned and diffable like everything else.
  • Fitness tests. Frameworks like ArchUnit turn layering rules, package visibility and dependency directions into tests that run on every build.
  • Policies and contracts. Security and compliance rules live in policy engines such as Open Policy Agent (Rego); interfaces live in OpenAPI; decisions live in versioned Architecture Decision Records.

Here is what “storage must never be public” looks like when it stops being a sentence in a document and becomes a rule:

A Rego policy that blocks publicly accessible storage
package architecture.compliance

deny[msg] {
    input.resource_type == "aws_s3_bucket"
    input.attributes.public_access == true
    msg := sprintf("Architectural Violation: S3 Bucket '%v' cannot be publicly accessible.", [input.name])
}

There is a second benefit that matters even more in the age of AI. Files like architecture/boundaries.dsl, policies/, decisions/ and api-contracts/ are exactly the kind of context a coding assistant can actually read. Institutional memory and old wiki pages are not.

3.2 Why it’s harder than it sounds

Nobody really argues against Architecture as Code. Doing it everywhere is another matter.

Firefly’s research on infrastructure shows how wide the gap is:

  • 2025: 89% of organizations had adopted Infrastructure as Code, but only 6% had it covering everything. 65% said their cloud had become more complex over the previous two years.
  • 2026: 90% said their IaC orchestration falls short; only 5% called their infrastructure fully self-healing. A third had already traced a costly production incident back to infrastructure drift.

The same research shows what this means for automation. 44% of organizations are testing or piloting AI for infrastructure work, yet only 34% trust an AI agent to make production changes on its own. Asked what holds them back, the most common answer — 42% — was missing guardrails.

Firefly 2026 · The AI automation paradox

Piloting / testing AI automation the interest is there44%
Trust autonomous changes the trust is not34%
Cite missing guardrails the top blocker42%

In practice, four things get in the way:

The representation tax

Translating what an architect means into DSLs, Rego policies and ArchUnit suites is real work. Writing and maintaining governance code can become a project as big as the application.

Judgement doesn’t fit in a rule

“UI code must not import database drivers” is easy to check. “Balance regional latency against data sovereignty” or “is independent deployment worth the overhead of another service?” is not — those are trade-offs, not yes-or-no questions.

Too many artifacts

Terraform or OpenTofu, C4 models, Rego files, YAML pipelines, OpenAPI specs, ADRs in Markdown, tests in every language — governance spreads across a dozen formats.

The room gets smaller

Once the rules live in code-heavy repositories, the people outside engineering — product, compliance, risk, leadership — can no longer read them, let alone shape them.

So we have one movement that is fast but vague, and another that is precise but heavy. The interesting question is what happens when you put them together.

04

Two movements, one destination

4.1 They only look opposed

At first glance, Implementation-as-Spec and Architecture-as-Code pull in opposite directions:

Abstract · Natural language specs / intent
↑

Implementation-as-Specmoves upward, toward plain-language intent

Executable intent
↓

Architecture-as-Codemoves downward, toward formal rules

Concrete · Executable rules / policy / code
  1. Architecture-as-Code moves down. It takes informal designs and turns them into precise, machine-checkable rules and tests.
  2. Implementation-as-Spec moves up. It takes hand-written code and lifts it into requirements, structured specifications and plans.

But they are solving the same problem from two ends: turning what people intend into what machines do. Architecture-as-Code supplies the hard edges that plain-language specs lack. And AI, the engine behind Implementation-as-Spec, can take on much of the boilerplate that makes Architecture-as-Code so expensive to write. Each one covers the other’s weakness.

Architecture as CodeImplementation as Spec
DirectionFrom abstract ideas to formal rules.From hand-written code to abstract intent.
What it protectsThe system from becoming structurally or legally invalid.The team’s time — it produces the behaviour they asked for, fast.
What it’s made ofModels, policies, schemas, tests.Intent, requirements, acceptance criteria.
The machine’s jobCheck, constrain, enforce.Understand, plan, generate.
The human’s jobDecide what must never be broken, and weigh the trade-offs.Say what the business needs, and judge whether the result is right.
The riskOver-formalising, and the upkeep that comes with it.Ambiguity, and code that is plausible but wrong.
How it failsCorrect rules that faithfully enforce the wrong architecture.Correct-looking code that does the wrong thing.
The goalA structure that can’t quietly erode.Software that needs as little hand-writing as possible.

Put together, they suggest a simple division of labour:

Natural language becomes the authoring interface. Formal artifacts become the enforcement interface. Generated code becomes the replaceable implementation layer.

4.2 Five layers, from intent to evidence

To use AI at scale without the architecture drifting away, it helps to think of the work in five layers:

  1. 1
    Human intentBusiness goals, domain language, trade-offs
  2. 2
    Structured specificationGitHub Spec Kit, a project constitution, scenarios, acceptance criteria
  3. 3
    Executable architecture & policyArchUnit rules, Rego policies, OpenAPI contracts, boundaries.dsl
  4. 4
    Generated implementationAI-written application code, IaC, migrations, boilerplate
  5. 5
    Verification & runtime evidenceCI tests, policy checks, telemetry, SLOs, drift detection
Layer 1 · Human intent
This is where people explain why the system exists: the vision, the language of the domain, the trade-offs the business is willing to make. Everything else follows from it.
Layer 2 · Structured specification
Loose intent becomes something testable: requirements, scenarios, edge cases and acceptance criteria — for example with GitHub Spec Kit. This layer also holds the project’s constitution: the principles that are not up for negotiation. Because the planning agent reads it before any code is written, the constitution is the bridge to Layer 3.
Layer 3 · Executable architecture & policy
The hard edges: dependency rules, ArchUnit suites, Rego policies, versioned ADRs and OpenAPI contracts. Coding agents read them as context, so what they generate stays inside the boundaries the organisation has agreed on.
Layer 4 · Generated implementation
The code itself — application logic, migrations, integration glue, infrastructure scripts — written largely by AI, but only within the space Layer 3 allows.
Layer 5 · Verification & runtime evidence
Proof, not promises: tests in CI, static policy checks, telemetry, SLOs and drift alerts. What this layer learns flows back up to the specification and the rules.

That stack is not just a diagram for me. It is how M1Spec is built:

LayerWhere it lives in M1Spec
1 · Human intentDomain stories told as numbered sentences, in a shared domain language per bounded context.
2 · Structured specificationThe generated event model, the backlog with acceptance criteria, and the API contract.
3 · Executable architecture & policyThe C4 model, technology catalog and enterprise patterns; correctness rules on every save; the readiness gate; your recorded decisions.
4 · Generated implementationThe desktop agent, driving Claude Code or GitHub Copilot CLI in your own repository, one commit per task.
5 · Verification & evidenceTest plans from the event model, your CI on the pull request, and a C4 change set confirmed against the code. Runtime telemetry and SLOs stay with your observability stack.
The five layers, mapped onto M1Spec.

05

What to do on Monday

5.1 The less we write by hand, the more we must write down

It sounds backwards at first: the more implementation we hand over to machines, the more formal our architecture has to become, not less. Follow the chain and it makes sense:

  1. Code gets written faster
  2. There are many more possible implementations
  3. Checking them gets harder
  4. Rules a machine can check become more valuable
  5. Architecture as Code spreads

Teams have always relied on things nobody wrote down. “We never call the database from this controller” lived in people’s heads and in review comments. An AI assistant doesn’t share that memory. It knows what’s in its context window, and it fills the rest with probability.

As generated pull requests multiply, so does the number of ways a change can be wrong, and the burden on reviewers grows with it. The only way out of that bottleneck is to take what the team knows implicitly and turn it into explicit, automatic checks — fitness functions a machine can run on every change.

We are abstracting away the implementation at exactly the same time that we are formalizing the architecture.

5.2 Four moves for engineering leaders

If you lead a team that is adopting AI, here is where I would start:

  1. Separate writing intent from enforcing it. Let people describe what they need in natural language and structured specs — GitHub Spec Kit is one way. But keep the architecture itself in formal, machine-readable models, contracts and policies. Give the project a constitution that the AI has to respect before it writes a single line.
  2. Make the build say no. Put architectural checks — ArchUnit for dependencies, Open Policy Agent for compliance — into the CI/CD pipeline, and let a violation fail the pull request automatically. A circuit breaker that trips before a human ever has to review is cheaper than any review.
  3. Feed your agents the right context. Keep the architecture where an AI can read it: textual C4 models in architecture/boundaries.dsl, decisions in decisions/, OpenAPI schemas and policy files in predictable places. The clearer the boundaries in the repository, the smaller the space the agent can wander into.
  4. Move your best people to verification. Stop spending senior engineers on reading generated code line by line and writing boilerplate. Point them at the things only they can do: designing the automated checks, sharpening the constitution, and encoding the rules the organisation actually depends on.

The cheaper code generation becomes, the more valuable explicit architectural constraints become. Architecture-as-code and implementation-as-spec move in opposite directions along the abstraction axis, but converge on a single objective: executable intent.

This paper is why M1Spec exists.

M1Spec turns executable intent into a working system: domain stories as the authoring interface, event models, contracts and a readiness gate as the enforcement interface, and code your own assistant generates as the replaceable layer — with the architecture updated from what was actually built.