The Liability of Code: Software Engineering After AI
Code is no longer proof of progress. In an AI-augmented ecosystem, every generated line carries two costs: an immediate token tax during reasoning and a perpetual long-term support tax after it ships. The strongest engineer is not the fastest code author. It is the system governor who narrows architecture, deletes maintenance surface area, and designs feedback loops that keep routine correction inside controlled boundaries.
Code is no longer proof of progress. In AI-assisted engineering, every generated line becomes an operating expense: it consumes context during future reasoning, expands review scope, increases dependency exposure, and creates another surface that must survive migration, incident response, and ownership transfer.
The scarce skill is deciding what should exist at all. That decision happens before generation: which interface owns the behavior, which data boundary may change, which failure mode is acceptable, which telemetry proves the change worked, and which code path should be deleted instead of patched.
In AI-assisted engineering, code is not the asset. The asset is the smallest stable system that can keep delivering value without expanding its own maintenance gravity.
1. The Code Liability Principle
Every line of code is an instruction the organization must keep understanding. It must be tested, secured, migrated, explained to new engineers, scanned by tools, reasoned over by AI, and revisited when the surrounding system changes. Code is useful only when the capability it provides exceeds that obligation.
AI makes this more visible because code now carries a dual tax. The first is the token tax: more files, more branches, more duplicate patterns, and more incidental complexity increase the cost and latency of every AI-assisted reasoning loop. The second is the LTS tax: every generated abstraction becomes future work for dependency upgrades, security patches, test maintenance, onboarding, and incident response. This is not theoretical. Public analysis of real-world pull requests has found AI-assisted changes producing 1.7x more issues, including increases in logic and security defects, when teams treat generation as a substitute for rigorous verification.
This changes the definition of engineering value. The question is not "how much code did we ship?" It is "how much durable business capability did we extract from the smallest maintainable system surface?" The highest-value code is often the code that never had to be written, or the code that was deleted after the right boundary made it unnecessary.
Metric
Volume-Oriented Engineering
Liability-Oriented Engineering
Token Tax
Context consumed every time an AI assistant reads, reasons over, or patches the system
Large codebase inflates latency, cost, and reasoning error
↓Smaller surface area keeps prompts focused and patches precise
LTS Tax
Ongoing maintenance from dependency drift, security fixes, and logic rot
Every generated branch becomes future operating expense
↓Durable interfaces reduce perpetual support burden
Review Load
Human effort required to validate generated changes
Volume overwhelms review and hides regressions
↓Small atomic changes make risk legible
AI Patch Accuracy
How reliably an assistant can understand and modify the repository
Bloated context produces shallow, redundant fixes
↑Smaller context reduces duplicate paths and ownership ambiguity
Capability per Maintained Surface
Value delivered per durable unit of system complexity
Output measured by code volume
↑Output measured by stable capability with minimal code
2. Upstream Architecture as Risk Mitigation
Architecture is not the diagram that comes after the implementation. In AI-assisted engineering, architecture is the filter that prevents generation from flooding the system with plausible but unnecessary code. The more precise the constraints, the less room there is for the model to invent surface area.
Constraint-based design starts before a prompt is written. It defines ownership boundaries, interfaces, failure modes, privacy limits, persistence rules, and observability expectations. Those constraints become the guardrails that allow AI to execute with high density instead of broad improvisation.
This is why declarative precision matters. Spending an hour defining what the system should do, what it must never do, and which existing patterns it must reuse can save days of auditing what an AI wrote. Specification is not ceremony. It is compression.
Before I let AI generate implementation, I want four constraints written down: the owning interface, the persistence boundary, the observable success signal, and the rollback path. If the change crosses API, storage, and UI boundaries in one generated diff, the work is already too wide. The model will fill ambiguity with plausible structure. The reviewer will inherit the real architecture decision after the code exists.
I refuse to build AI workflows that optimize for accepted suggestions while leaving review scope, ownership ambiguity, and remediation latency unchanged. That is not acceleration. It is a faster path to code nobody can confidently delete.
Prompt-Led Implementation
- 01Ask AI to build the featureThe model fills gaps with statistically likely structure.
- 02Review a large diffHumans discover architecture decisions after code exists.
- 03Patch regressionsFollow-up prompts add more code around unclear boundaries.
- 04Carry the debt forwardThe repository becomes harder for humans and AI to reason about.
Architecture-Led Generation
- 01Define constraints firstInterfaces, invariants, data flow, and failure behavior are explicit.
- 02Generate inside boundariesAI implements a narrower, testable slice.
- 03Validate against telemetryCI, logs, metrics, and tests verify behavior continuously.
- 04Delete redundant surface areaRefactoring reduces future token and LTS cost.
3. AI for High-Density Execution
The goal of AI is not to reach "feature complete" faster if the result is a larger, noisier, more fragile repository. The goal is a smaller change graph: fewer moving parts, fewer ownership crossings, stronger tests, and less code for the next engineer or model to load into context.
High-density execution means each generated change should carry more intent per line. A good AI patch reuses an existing extension point, deletes obsolete logic, tightens an interface, or converts repeated behavior into a single declarative rule. A bad patch simply moves implementation time from the author to the reviewer.
This makes refactoring a reduction discipline. The best refactor in an AI-enabled codebase is not the one that makes the structure more fashionable. It is the one that removes future prompt tokens, removes future test cases, removes future dependency exposure, and makes the next change obvious.
4. The Self-Looping Feedback Cycle
Once AI can generate and modify code, the central engineering question becomes governance. Who defines the loop? What evidence is allowed to trigger remediation? Which changes can be automated, which require review, and which require architectural redesign?
A mature AI engineering system has a read and observe loop. Production telemetry, logs, traces, metrics, incident reports, CI failures, and security findings flow into an orchestration layer that can classify routine issues, propose patches, run validations, and send bounded changes through review.
The point is not to remove the human engineer. It is to move the human out of repetitive typing and into the governance of self-correcting systems. Automated remediation should handle routine maintenance; humans should own the constraints, escalation paths, risk thresholds, and long-term system direction.
5. The Engineer as System Governor
The engineer's role is shifting from code author to system governor. That does not make engineering less technical. It makes the technical surface wider: architecture, observability, context efficiency, risk controls, evaluation design, and automated remediation all become part of the job.
In this paradigm, the best engineers will be judged by the systems they keep small, the interfaces they make durable, the loops they make observable, and the maintenance they prevent from existing. Code still matters. But code is not proof of value. It is a liability that must continuously justify its place in the system.
Any AI workflow that increases code volume without reducing review scope, ownership ambiguity, or remediation latency is not engineering acceleration. It is deferred maintenance with better autocomplete.
1.DORA 2024: Accelerate State of DevOps Report
2.GitClear: AI Copilot Code Quality Research 2025
3.DX: AI Measurement Framework
4.CodeRabbit: State of AI vs. Human Code Generation
