Skip to content

Build1 publisher3 min readPublished

Agentic architecture review earns its place as advice, not as a gate

An InfoQ article puts AI judgment beside deterministic fitness functions rather than in front of them: scoped evidence, versioned rubrics, structured verdicts, humans on low confidence.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Agentic architecture review earns its place as advice, not as a gate
Generated illustration

What happened

  • Evolutionary architecture treats architecture not as a fixed target state but as a system of decisions that can evolve safely as business needs, technology, operating conditions and team structures change.
  • Fitness functions turn architectural intent into executable feedback; examples include dependency rules protecting package boundaries, contract tests protecting integration compatibility, latency budgets protecting performance and security scans protecting policy compliance.
  • Deterministic fitness functions should remain the primary enforcement mechanism for measurable invariants such as dependency direction, contract shape, latency budgets, security posture and policy checks.
  • Agentic fitness functions add value when architectural risk is evidence-bound but judgement-heavy, such as boundary fidelity, semantic contract drift, workflow coupling and stale ADR assumptions.
  • A production-ready implementation separates deterministic gates from agentic advisory signals, scopes evidence to the change, applies versioned rubrics, returns structured verdicts, and escalates low-confidence or high-blast-radius outcomes to humans.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

An article on InfoQ makes the case for adding AI agents to evolutionary architecture as "agentic fitness functions", and its load-bearing recommendation is a boundary rather than a capability: deterministic fitness functions stay the primary enforcement mechanism for measurable invariants such as dependency direction, contract shape, latency budgets, security posture and policy checks, while the agent emits advisory signals [3][5]. That split matters because the failure mode of putting a language model on the enforcement path is not a bad review, it is a build that blocks or passes for reasons nobody can reproduce a month later.

The argument for the layer starts with an honest limit. Deterministic checks only protect what can be reduced to a rule, a threshold, a schema or an executable experiment [14]. A dependency rule can show that a service interaction changed without telling you whether it is intentional collaboration or accidental coupling [7]. A schema diff can prove an API still parses without judging whether a new field preserves the semantic model or leaks a UI concern into a domain event [8]. A trace can reveal a new runtime path without determining whether that path bypasses workflow ownership [9]. According to the article, this is how architecture normally decays: individually reasonable changes that pass every written rule while moving the implementation away from the intent the team believed it had protected [12]. The historical answer was manual review, which does not scale down to every pull request, contract change and workflow trace [13].

What the agent is meant to do is narrower than "review the architecture". Calibrated against architecture decision records, ownership metadata, service boundaries, rubrics and historical examples, it evaluates bounded evidence and returns a structured judgment with score, confidence, rationale and escalation guidance [10]. The target concerns are the ones that are evidence-bound but judgement-heavy: boundary fidelity, semantic contract drift, workflow coupling, stale ADR assumptions [4]. It replaces neither the deterministic checks nor the architects [11].

The implementation pattern the article calls production-ready is five constraints, not one [5][17]. Separate deterministic gates from agentic advisory signals, so enforcement stays reproducible. Scope evidence to the change, which is the only practical control on what the model is reasoning about. Version the rubrics, so a shift in verdicts can be traced to a rubric edit rather than argued about. Return structured verdicts, so the output is storable and diffable instead of prose. Escalate low-confidence or high-blast-radius outcomes to humans, which is the admission that confidence is a routing signal and not a score to be optimised [5].

The payoff claimed is not better reviews in the moment. It is that architectural judgment becomes observable, calibratable and auditable, and easier to convert into deterministic guardrails once patterns repeat [6]. That is a ratchet worth having: findings that recur graduate out of the advisory lane and become rules, which is also the only mechanism that stops the advisory lane growing without limit. The context is that architectural decisions now surface inside pull requests, contract diffs, workflow traces and agent-invoked actions, alongside smaller changes, faster feedback and AI-generated code [16].

Worth noting: the supplied excerpt sets out a design pattern and carries no measured accuracy or false-positive rates for agentic verdicts [18].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories