Start with evidence, not assumptions
A legacy system is often business-critical software whose behavior is only partly understood. Its true specification may be distributed across code, database procedures, job schedules, message formats, operator runbooks, vendor contracts, and the experience of a few people.
Begin by observing production safely. Inventory applications, interfaces, data stores, scheduled jobs, certificates, service accounts, network paths, upstream producers, downstream consumers, peak volumes, failure modes, and manual recovery steps. Record which system owns each business fact and who can approve a behavioral change.
The first architecture deliverable is a verified map. A polished target diagram is not useful if it omits the nightly file, shared table, or manual correction that keeps the business running.
Turn discovery into living documentation
Keep documentation beside the code and validate what can be validated. Use a system context diagram for people and external systems, container diagrams for deployable parts, sequence diagrams for critical transactions, and data-flow diagrams for sensitive or regulated information.
For every interface, document the protocol, schema, examples, authentication, ownership, availability expectation, timeout, retry policy, idempotency rule, ordering guarantee, error vocabulary, and versioning policy. Add data lineage, deployment topology, operational runbooks, service-level objectives, and architecture decision records.
Documentation becomes trustworthy when the pipeline checks links, schemas, examples, and generated API specifications. Assign an owner and review date to facts that cannot be checked automatically. Treat documentation changes as part of the same pull request as behavior changes.
Choose an integration boundary deliberately
Place an anti-corruption layer between modern domain concepts and legacy semantics. The adapter translates names, formats, error codes, and workflows so legacy rules do not leak across every new service.
- API facade: best for synchronous queries and commands when callers need immediate outcomes. Apply strict timeouts, bounded retries, circuit breaking, and rate limits.
- Messaging: best for asynchronous workflows and loose coupling. Design consumers for duplicate delivery, delayed messages, poison records, and replay.
- Change data capture: useful for propagating committed database changes without making the database a public API. Preserve ordering keys and monitor replication lag.
- Batch or managed file transfer: still appropriate for large scheduled exchanges. Add manifests, checksums, encryption, quarantine, replay, and reconciliation.
Use the Strangler Fig pattern to route selected capabilities through a facade and replace them one bounded slice at a time. Keep the old and new paths observable during coexistence. Avoid permanent direct access to legacy tables because it couples consumers to storage details, bypasses business rules, and makes ownership unclear.
Clients | Facade and routing layer |-----------------------| Anti-corruption adapters New capability | | API and events Legacy application Owned data | Outbox or CDC -> Event platform -> Consumers and reconciliation
Make data ownership and consistency explicit
Declare one system of record for every business entity during each migration phase. New components should not update legacy and modern stores independently inside one request. Prefer a transactional outbox or change data capture, idempotent consumers, and reconciliation over fragile distributed transactions.
Define stable business identifiers, mapping rules, retention, deletion, privacy classification, and acceptable replication delay. Measure record counts, totals, missing keys, duplicates, and field-level differences. When a multi-step workflow cannot be atomic, define compensating actions and make incomplete states visible to operators.
Database evolution should follow expand and contract: add backward-compatible structures, deploy readers and writers that tolerate both versions, migrate and verify data, then remove the obsolete structure only after all consumers have moved.
Build a layered automation testing strategy
- Characterization tests: capture current behavior before changing it. Golden-master samples are useful when the output is stable and reviewed, but mask timestamps and other nondeterministic values.
- Unit tests: exercise translation, validation, mapping, retry decisions, and domain rules quickly without infrastructure.
- Contract tests: verify that providers and consumers agree on requests, responses, events, optional fields, error cases, and compatibility. Publish contracts and verify them before deployment.
- Integration tests: run adapters against realistic databases, brokers, file stores, and protocol emulators. Use disposable environments and representative schemas instead of mocking away the risk.
- End-to-end tests: keep a small set for the most valuable business journeys, including retries, duplicate delivery, partial failure, and recovery.
- Non-functional tests: cover capacity, soak behavior, security, failover, backup restoration, and resilience to slow or unavailable dependencies.
- Migration tests: prove data conversion, reconciliation, restartability, rollback, and the ability to resume safely from a checkpoint.
Test data needs the same discipline as production code. Use synthetic or irreversibly masked data, version fixtures, state the expected owner of each dataset, and clean environments predictably. A passing test suite should provide evidence about a defined compatibility surface, not a promise that every device, dependency, or production condition is perfect.
Design CI/CD as a chain of evidence
A pull-request pipeline should lint code and documentation, validate API and event schemas, run unit tests, verify consumer contracts, scan dependencies and source, execute integration tests, and build one immutable artifact. Generate a software bill of materials, sign or attest the artifact where appropriate, and promote that exact artifact through environments.
The delivery pipeline should apply environment-specific configuration from a controlled source, retrieve secrets at runtime, execute backward-compatible migrations, deploy progressively, run smoke and synthetic transaction checks, and evaluate service-level indicators before continuing. Canary or blue-green deployment limits exposure. Feature flags separate deployment from business release.
Pull request -> docs and schema checks -> unit and contract tests -> integration, security, and migration tests -> build once, attest, publish -> deploy to test -> smoke and reconciliation -> canary production -> observe -> promote or stop
Define both rollback and roll-forward procedures. Rollback is safe only while data and contracts remain backward compatible. For irreversible changes, prefer disabling the new route, repairing forward, and replaying messages from a known checkpoint.
Make coexistence observable and operable
Propagate a correlation identifier across APIs, messages, batches, and legacy calls. Emit structured logs, business and technical metrics, and distributed traces where protocols allow them. Dashboards should show throughput, latency, errors, retries, circuit state, queue depth, dead letters, replication lag, reconciliation differences, and business completion rates.
Alerts must lead to an owned action and a tested runbook. Protect sensitive values in telemetry, preserve audit evidence, and test restoration and replay procedures. During migration, compare old and new outcomes with shadow traffic or dual reads when it is safe, but avoid double writes that can create competing sources of truth.
A phased modernization roadmap
- Discover: map dependencies, business journeys, risks, ownership, and measurable baselines.
- Stabilize: add observability, characterization tests, repeatable builds, and safe configuration management around the existing system.
- Encapsulate: introduce adapters, schemas, contract tests, and a controlled routing layer.
- Automate: build disposable test environments and a CI/CD path that produces traceable evidence.
- Extract: move one cohesive capability, migrate its data ownership, and release progressively.
- Retire: prove that traffic, data, controls, and operational responsibilities have moved before removing the old path.
Use exit criteria for every phase. Examples include a verified dependency map, zero unresolved contract incompatibilities, reconciliation within an agreed tolerance, successful recovery exercises, and a defined period with no traffic on the retired path.
Architecture review checklist
- Critical flows, dependencies, owners, and recovery procedures are documented.
- Every interface has a versioned contract and explicit compatibility policy.
- An anti-corruption layer contains legacy concepts and protocols.
- Each business entity has one declared system of record per migration phase.
- Retries are bounded, operations are idempotent, and poison work is recoverable.
- Automated tests cover behavior, contracts, infrastructure, migration, and failure recovery.
- The pipeline builds once and promotes an immutable, traceable artifact.
- Schema and database changes remain backward compatible during rollout.
- Dashboards and alerts expose both technical health and business completion.
- Rollback, roll-forward, replay, reconciliation, and retirement criteria are tested.
The strongest legacy integration architecture is not a dramatic rewrite. It is a controlled learning system: document what is true, isolate change behind stable boundaries, verify compatibility automatically, observe real outcomes, and replace capabilities in reversible increments.
References
- Strangler Fig patternMicrosoft Learn
- Anti-corruption Layer patternMicrosoft Learn
- OpenAPI SpecificationOpenAPI Initiative
- Introduction to contract testingPact
- OpenTelemetry signalsOpenTelemetry
- Continuous integration with GitHub ActionsGitHub Docs
Get the next Insight.
A short email when a new article is published. Nothing else.