Build1 distinct publisher3 min readUpdated
An InfoQ article puts AI judgment beside deterministic fitness functions rather than in front of them: scoped evidence, versioned rubrics, structured verdicts, humans on low confidence.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
An article on InfoQ makes the case for adding AI agents to evolutionary architecture as "agentic fitness functions", and its load-bearing recommendation is a boundary rather than a capability: deterministic fitness functions stay the primary enforcement mechanism for measurable invariants such as dependency direction, contract shape, latency budgets, security posture and policy checks, while the agent emits advisory signals [3][5]. That split matters because the failure mode of putting a language model on the enforcement path is not a bad review, it is a build that blocks or passes for reasons nobody can reproduce a month later.
The argument for the layer starts with an honest limit. Deterministic checks only protect what can be reduced to a rule, a threshold, a schema or an executable experiment [14]. A dependency rule can show that a service interaction changed without telling you whether it is intentional collaboration or accidental coupling [7]. A schema diff can prove an API still parses without judging whether a new field preserves the semantic model or leaks a UI concern into a domain event [8]. A trace can reveal a new runtime path without determining whether that path bypasses workflow ownership [9]. According to the article, this is how architecture normally decays: individually reasonable changes that pass every written rule while moving the implementation away from the intent the team believed it had protected [12]. The historical answer was manual review, which does not scale down to every pull request, contract change and workflow trace [13].
What the agent is meant to do is narrower than "review the architecture". Calibrated against architecture decision records, ownership metadata, service boundaries, rubrics and historical examples, it evaluates bounded evidence and returns a structured judgment with score, confidence, rationale and escalation guidance [10]. The target concerns are the ones that are evidence-bound but judgement-heavy: boundary fidelity, semantic contract drift, workflow coupling, stale ADR assumptions [4]. It replaces neither the deterministic checks nor the architects [11].
The implementation pattern the article calls production-ready is five constraints, not one [5][17]. Separate deterministic gates from agentic advisory signals, so enforcement stays reproducible. Scope evidence to the change, which is the only practical control on what the model is reasoning about. Version the rubrics, so a shift in verdicts can be traced to a rubric edit rather than argued about. Return structured verdicts, so the output is storable and diffable instead of prose. Escalate low-confidence or high-blast-radius outcomes to humans, which is the admission that confidence is a routing signal and not a score to be optimised [5].
The payoff claimed is not better reviews in the moment. It is that architectural judgment becomes observable, calibratable and auditable, and easier to convert into deterministic guardrails once patterns repeat [6]. That is a ratchet worth having: findings that recur graduate out of the advisory lane and become rules, which is also the only mechanism that stops the advisory lane growing without limit. The context is that architectural decisions now surface inside pull requests, contract diffs, workflow traces and agent-invoked actions, alongside smaller changes, faster feedback and AI-generated code [16].
Worth noting: the supplied excerpt sets out a design pattern and carries no measured accuracy or false-positive rates for agentic verdicts [18].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Evolutionary architecture treats architecture not as a fixed target state but as a system of decisions that can evolve safely as business needs, technology, operating conditions and team structures change.
Fitness functions turn architectural intent into executable feedback; examples include dependency rules protecting package boundaries, contract tests protecting integration compatibility, latency budgets protecting performance and security scans protecting policy compliance.
Deterministic fitness functions should remain the primary enforcement mechanism for measurable invariants such as dependency direction, contract shape, latency budgets, security posture and policy checks.
A production-ready implementation separates deterministic gates from agentic advisory signals, scopes evidence to the change, applies versioned rubrics, returns structured verdicts, and escalates low-confidence or high-blast-radius outcomes to humans.
A dependency rule can show that a service interaction changed but cannot always tell whether the change represents intentional collaboration or accidental coupling.
A schema diff can prove that an API still parses but cannot always judge whether the contract still expresses the right domain concept, or whether a new field preserves the semantic model or leaks a UI concern into a domain event.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source design argument, no measurements
One practitioner article from one publisher carries the entire cluster. Its descriptive and definitional content is directly and clearly attested, and its illustrations of where mechanical checks fall short are internally coherent. But every efficacy-bearing element — that agentic judges add value, that a calibrated agent returns reliable structured verdicts, that the pattern is production-ready — is asserted rather than evidenced: the excerpt contains no accuracy, precision, recall, variance or false-positive data, no reference implementation, and no external corroboration.
No adoption signal supplied
The supplied material contains no release, deployment, benchmark, pricing, licensing or usage disclosure. No tool, model, vendor or adopting organisation is named, and no team is reported running agentic fitness functions in a pipeline. Adoption cannot be estimated without inventing facts the source does not provide.
Near-aligned, mildly ahead of evidence
The article deliberately deflates its own AI claim: agentic checks are advisory, explicitly 'not an oracle', gain influence only after calibration proves acceptable precision and recall, and never displace deterministic gates or architects. That restraint keeps the gap close to zero. It stays slightly positive because a five-control pattern is called 'production-ready' while the excerpt supplies no measured accuracy and no adoption evidence at all, so the readiness framing runs a little ahead of what is shown.
Thought-leadership incentive, no visible product pitch
The supplied excerpt promotes no vendor, product, model or paid service; the only diagram is marked author-created, and the pattern is described generically enough to be implemented with any agent. The residual incentive is category-building: a contributed practitioner article on a developer-media outlet benefits its author's and publisher's standing by naming and framing a new pattern ('agentic fitness functions'), and the URL carries publisher campaign tracking parameters. Author affiliation is not supplied, so any commercial tie behind the framing cannot be assessed.
High on what is proposed, low on whether it works
Confidence is high that the cluster accurately represents the recommendation — the takeaways are explicit, unambiguous and self-consistent, and the advisory-not-gate posture is stated twice. Confidence is low on any operational conclusion: one publisher, one article, zero adoption observations, and no measured verdict quality, so the pattern's real-world reliability, cost and failure modes are unknown.
build
A green build only proves your agent was consistent with itself1 distinct publisher
build
Flux moves GitOps' source of truth into registries you own, and mirroring becomes the prerequisite1 distinct publisher
build
JDK 28 firms up: a JSON API in the incubator, and a deprecation notice for Intel Macs1 distinct publisher
build
One DGX Spark, four Macs, and an attempt to turn token billing into a capital purchase1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 17, 2026