Designs runtime and platform architecture inside a chosen solution. Use when deciding modular monolith vs services, consistency, resilience, or estate topology.
Use this skill for deep software and platform architecture decisions inside a known solution shape rather than implementation details within a single service or component.
If the question starts from a business workflow, system landscape, target state, or phased cross-system migration, use ../software-solution-architecture/SKILL.md first and come here for runtime, decomposition, and operability depth.
Treat estate modernization, platform engineering, and AI-native interoperability as optional deep dives. Do not load them unless the user is explicitly asking for those concerns.
AI-native architecture: RAG boundaries, tool gateways, and interoperability protocols when the request is architecture-level rather than tool/server implementation
Agent workflow implementation / MCP server implementation → ai-agents, agents-mcp
Boundary Rules
This skill owns runtime boundaries, deployable-unit decisions, data consistency tradeoffs, resilience internals, and platform defaults.
Start from the simplest architecture that satisfies the constraints; do not default to microservices, event sourcing, service mesh, or multi-agent splits without explicit evidence.
If the unresolved question is still "which systems participate, where is the system of record, or what is the target-state landscape?" route back to software-solution-architecture.
If the unresolved question is implementation of agent protocols, tool servers, or runtime-specific integrations, route to ai-agents or agents-mcp.
Decision Tree: Choosing Architecture Pattern
Primary question: [What kind of architecture problem is this?] ├─ Large estate with many repos/services and rising cognitive load? │ ├─ Runtime count is the main problem → Bounded-context platforms + selective consolidation
Decision Factors:
Default posture: prefer modular monolith over microservices unless independent deployment, ownership, and operability benefits are clear — see the explicit team-size/release-cadence/operational-maturity gates in modern-patterns.md § Modular Monolith vs. Microservices
Estate posture: optimize for fewer runtime units before fewer repos; repositories are collaboration units, runtimes are operational cost centers
Agent posture: prefer deterministic workflows or a single agent before introducing multi-agent coordination
Connectivity posture: prefer gateway plus application-library patterns until mTLS, traffic policy, or shared telemetry needs justify mesh complexity
Team structure (Conway's Law) — architecture mirrors org structure
Where to split: fracture planes (Skelton and Pais, Team Topologies, 2nd ed., 2025). A fracture plane is "a natural seam in the software system that allows the system to be split easily into two or more parts" — the stonemason's analogy. Candidate planes: business domain bounded context (the default, and the one most splits should map to), regulatory compliance, change cadence, team location, risk, performance isolation, technology, and user personas.
Litmus test, quoted: "Does the resulting architecture support more autonomous teams (less dependent teams) with reduced cognitive load (less disparate responsibilities)?" Concretely: after the split, can each team build, test, and deploy its part without coordinating with another team?
Composite rule: real boundaries usually combine planes — "we can and should break down a monolith by combining different types of fracture planes," and "often, a combination of fracture planes will be required."
Distributed-monolith warning: splitting the software without aligning the boundaries to teams and their release paths buys distribution cost with none of the autonomy. The book quotes Amy Phillips: "If you have microservices but you wait and do end-to-end testing of a combination of them before a release, what you have is a distributed monolith." Coupling also creeps in below the service boundary — shared databases, coupled builds and releases.
Make segments team-sized: "it is essential to make software segments team sized so that teams can effectively own and evolve their software in a sustainable way."
The references in this skill are background knowledge for you — absorb the patterns and present them as your own expertise. Do not cite internal reference file names (e.g., "from data-architecture-patterns.md") in user-facing output. Users don't know these files exist.
Every architecture recommendation must cover the following; skip elements only with explicit justification:
Simplest sufficient topology — state the least-complex architecture that still satisfies requirements
Concrete technology picks — name specific technologies (e.g., "Temporal.io for workflow orchestration", not just "an orchestrator")
Recommended option + rejected alternatives — what was considered, why alternatives lost
What NOT to build — explicitly defer or exclude premature scope
Team and process alignment — CODEOWNERS, deployment ownership, on-call boundaries
Repo and runtime model — for multi-repo estates, distinguish repo count from deployable count
evals/evals.json — trigger, non-trigger, and near-boundary behavioral checks for this skill
Applied-Recipe Toolkits
references/decision-theory-applied.md — Decision-theory applied recipes for architecture: ADRs with EU + sensitivity, real-options for irreversible choices, VoI on spikes.
references/queueing-theory-applied.md — Queueing-theory applied recipes for architecture: service sizing, backpressure topology, tail-latency budget.
references/distributed-systems-applied.md — Distributed-systems primitives applied to architecture: CAP-conscious service boundaries, consensus algorithm selection, idempotency at API surfaces, leases-with-fencing for leader-elected jobs, quorum sizing, consistency-vs-latency ADR template.
docs-diagram-design — Whether a diagram earns its place, and what it must show
ai-agents — Agent system design, orchestration, evaluation
agents-mcp — MCP server/client patterns and integration
Freshness Protocol
When users ask version-sensitive questions about architecture patterns, platform engineering, or AI-native systems, verify current information before answering.
Trigger Conditions
"What's the best architecture for [use case]?"
"Microservices vs monolith — what's the current recommendation?"
"What's the latest in platform engineering / service mesh / AI architecture?"
"How do I modernize 50/100+ repos or reduce service sprawl?"
"Is [pattern] still recommended?"
How to Freshness-Check
Start from data/sources.json and prefer official docs, standards, release notes, and lifecycle pages.
Run a targeted web search for the specific architecture pattern or platform.
Use non-primary sources only as durable background, not as freshness authority.
Load only when the question explicitly involves current trends, vendor-specific constraints, AI-native architecture, or "what's the latest thinking on X?"
optional_ai_architecture — MCP/A2A protocols and architecture-level AI interoperability references
modern_architecture_2026 — ambient mesh and other version-sensitive platform patterns
If live web access is available, consult 2–3 authoritative sources from data/sources.json and fold findings into the recommendation. If not, answer with durable patterns and explicitly state assumptions that could change (vendor limits, pricing, managed-service capabilities, or lifecycle status).
Fact-Checking
Known bugs, regressions, framework/compiler/runtime footguns, and version-specific crash or workaround guidance must be verified against current primary web sources before being treated as current fact.
Use web search/web fetch to verify current external facts, versions, pricing, deadlines, regulations, or platform behavior before final answers.
Prefer primary sources; report source links and dates for volatile information.
If web access is unavailable, state the limitation and mark guidance as unverified.
Learnings Loop
Before applying this skill on a non-trivial task, read learnings.consolidated.md in this directory (and learnings.md if present).
After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.
│ ├─ Delivery inconsistency is the main problem → IDP + golden paths + scorecards
│ └─ Both are true → Platform-first modernization, then consolidate low-value runtime units
│
├─ Deterministic workflow, known steps?
│ ├─ Single deployable acceptable → Modular Monolith
│ ├─ Independent teams/capabilities required → Sequential or event-driven services
│ └─ Burst-driven or edge-triggered workload → Serverless / event-driven
│
├─ Adaptive workflow with tool use and reasoning?
│ ├─ One agent can own the task → Single-agent system
│ ├─ Keep data and writes together → Monolith or Modular Monolith
│ └─ Split only at stable bounded contexts → Microservices with owned data
│
└─ Need platform-level consistency across many teams?
├─ Repeated service creation / compliance needs → IDP + golden paths
└─ Cross-agent or cross-vendor interoperability → MCP for tools/context, A2A for agent-to-agent
Architecture design request -> Define quality attributes and system boundaries -> Map domain model, dependencies, and failure modes -> Choose architecture pattern and integration style -> Document rejected options and tradeoffs -> Define migration, observability, and verification checks -> Hand off implementable decisions and open risks