
Begin
14 pages · ~28 min
DevOps Platform Selection and Architecture
This training guides engineers and architects in selecting and designing a DevOps platform, covering architecture patterns and key tradeoffs for informed decisions.
A digital instructor presents all 14 pages. Hold “Ask” at any point and ask out loud — the answer comes from this course. No sign-up needed.
What you’ll learn
- 01DevOps Platform: Selection, Architecture, and TradeoffsWelcome. Over the next fourteen slides, we're going to treat your DevOps platform as a strategic decision, not a tooling purchase. Platform decisions now split three ways: build, buy, or assemble. So we'll work through three lenses. First, selection criteria you can defend. Second, a reference architecture you can actually run. Third, the tradeoffs stated plainly, so nobody is surprised later. Quick definitions, because the language drifts. A DevOps platform is the delivery system your teams depend on. An internal developer platform, or IDP, is the self-service layer on top of it, the golden paths that let developers ship without filing tickets. Capabilities span continuous integration and delivery, software supply chain security, environment provisioning, and cost governance. That last one matters more each year, because FinOps is shifting from dashboards to pre-deployment controls. And who has to be in the room? Platform engineering, security, operations, procurement, and finance. Leave any of them out, and you'll revisit the decision at renewal. Next, why platform choices reached the boardroom.
giiresearch.complatformengineering.orgsnsinsider.com+22 min - 02Why Platform Choices Reached the BoardroomSo why did platform choices suddenly land on the board agenda? Start with the numbers. By 2026, about eighty percent of large software organizations are expected to run platform teams, up from roughly forty-five percent in 2022. The internal developer platform market follows the same curve, growing from about ten point four billion dollars in 2026 toward thirty-one point six billion by 2031. That is budget moving, not just talk. Now the caution. Homegrown platforms stall most often on sponsorship, cited by forty-one percent, and unclear ownership, at thirty-seven percent. Nearly thirty percent of platform teams measure nothing, and roughly seventy percent of initiatives miss meaningful adoption without a golden path. So watch the overdue signals in your own organization: tool sprawl, cloud cost scrutiny, supply chain mandates, and ticket queues quietly replacing self-service. Those are the signals that move this from an engineering debate to an executive decision. Next, let us define what a DevOps platform must actually deliver. Capability Model: What a DevOps Platform Must Deliver.
giiresearch.complatformengineering.orgsnsinsider.com+22 min - 03Capability Model: What a DevOps Platform Must DeliverSo let's define what the platform actually has to deliver. Six domains: source and build, delivery, environments, observability, security, and the portal. Those are your table stakes, and the portal, service catalog, CI/CD templates, policy as code, secrets, and unified observability are the minimum bar. The golden path is the opinionated, supported default, with explicit escape hatches, because if it only works one way, you generate exceptions forever. Here's the decision point: score your platform against your organizational maturity, not a vendor checklist. Start with the highest-friction workflow every team repeats. Log a few weeks of requests, find the twenty percent of request types driving eighty percent of volume, and automate those first. That shift takes the platform team out of the loop on routine asks, and that is the real signal of self-service. Next, we'll look at the reference architecture: control plane, data plane, and integration surfaces.
cncf.iocubepath.comsystem-design.space+22 min - 04Reference Architecture: Control Plane, Data Plane, and Integration SurfacesLet's walk through the reference architecture. There are three planes to draw, and how you split them decides your blast radius. The control plane holds the service catalog, policy, orchestration, identity, audit, and desired state. The data plane runs the runners, agents, and clusters that serve traffic, and it keeps serving even if the control link drops. So choose a stricter control plane split when auditability matters more than operational simplicity. Integration surfaces are where the platform meets source control, registries, cloud APIs, Kubernetes, ticketing, and secrets. Treat each as a trust boundary. Register data planes over outbound mutual TLS so their API servers are never exposed inbound, isolate per tenant, and scope access to least privilege. For variants, use jurisdiction-scoped or air-gapped data planes when residency beats centralization. Next, Build, Buy, or Assemble: Sourcing Models and Lock-In Exposure.
cncf.iodocs.hashicorp.comarchstandard.org+22 min - 05Build, Buy, or Assemble: Sourcing Models and Lock-In ExposureLet's move to sourcing models and lock-in exposure. Three archetypes: a managed commercial platform, assembled open-source components, or a full internal build. Self-hosted Backstage typically needs two to four dedicated engineers. That's roughly four hundred fifty to eight hundred thousand dollars in year one, and two hundred fifty to four hundred fifty thousand every year after. A commercial platform at one hundred engineers runs about thirty-six to seventy-eight thousand dollars a year. Per-seat pricing only beats a platform team past five hundred to nine hundred engineers. So the real question is not build versus buy. It's whether owning this layer is strategically core. The lock-in that actually hurts is rarely the subscription. It's your golden paths, catalog metadata, and templates authored in a vendor's format, where delivery knowledge quietly accumulates. Before signing, ask the security reviewer and procurement lead one question: if this vendor disappeared tomorrow, what would migration truly cost? Build in-house only when compliance is genuinely bespoke, you have a standing platform team, and you're past roughly one hundred fifty to two hundred engineers. Next, we'll apply these tradeoffs through a weighted selection criteria framework.
developers.devsquareops.comblog.localops.co+22 min - 06Selection Criteria and Weighted Evaluation FrameworkNow, let's turn to how you actually score the options: a weighted evaluation framework. Define your criteria and pass or fail gates before any demo happens, because demos reward polish, not fit. Start with a weighting model across developer experience, security, integration, scalability, total cost of ownership, and vendor health. Score each vendor from zero to five, multiply by the weight, and sum to a total. Keep the gates binary. If a vendor fails your compliance or data residency gate, the score never gets calculated. Then weigh evidence over demos: run proof of concept work on your own workflows, take reference calls in your industry, and ask for architecture reviews. Watch four pitfalls. Feature count bias rewards breadth over the few capabilities that decide your outcome. Sunk cost pressure keeps a weak option alive because you already invested evaluation time. Single stakeholder dominance distorts weights toward one team's pain. And list pricing hides the real cost, so model seat, consumption, and infrastructure together. The payoff is a defensible recommendation procurement can negotiate and security can co-sign. Next, let's cover Security, Compliance, and Supply Chain Requirements.
developers.devsquareops.comblog.localops.co+22 min - 07Security, Compliance, and Supply Chain RequirementsNext, let's talk about security, compliance, and supply chain requirements, and how the platform becomes your control point. Instead of routing every change through a review queue, you enforce policy as code and admission control, so guardrails apply automatically. On supply chain, four things matter: provenance, SBOM, artifact signing, and build isolation. SLSA Build levels give you a shared vocabulary here. Level one means provenance exists, level two means it's authentic and signed, and level three means it's unforgeable and produced in an isolated environment. A security reviewer should decide how high you need to go based on the threat model, not on a vendor claim. Also eliminate shared administrator tokens; use short-lived credentials and least privilege for every build and deploy identity. Here's the tradeoff worth defending: strictness should cost platform teams configuration, not developers context switches. If a control forces engineers to stop and manually attest, adoption drops and people route around it. If it runs at admission time, developers keep their flow. Coming up next, architecture tradeoffs: autonomy, standardization, reliability, and cost.
giiresearch.complatformengineering.orgsnsinsider.com+22 min - 08Architecture Tradeoffs: Autonomy, Standardization, Reliability, and CostNow let's look at architecture tradeoffs: autonomy, standardization, reliability, and cost. On autonomy versus standardization, keep the golden path optional but clearly better. When compliance and audit readiness matter more than local freedom, centralize the control plane. When delivery speed matters more, federate ownership while the core team maintains shared primitives. On reliability, a data plane that keeps serving when the control plane link drops is the difference between a regional outage and a non-event. Cost has three levers: control plane overhead, idle capacity, and telemetry volume. Sequence showback before chargeback, because chargeback only works once allocation is accurate and budget owners can act. Finally, document each tradeoff as a decision with a named owner and a review date. That gives your platform engineer, security reviewer, and procurement lead a defensible record. Next, we'll look at FinOps and the Platform Cost Model.
cncf.iodocs.hashicorp.comarchstandard.org+22 min - 09FinOps and the Platform Cost ModelLet's turn to FinOps and the platform cost model. The core principle is that cost governance must be embedded at provisioning time, because retrospective chargeback cannot reverse spend you have already committed. That means attribution has to be mandatory. Team, product, environment, cost-center, and workload-type labels should be validated before the resource exists, blocked at admission if they are missing. For shared platform costs, you need a defensible split rule, usually usage share where you can measure it and headcount where you cannot. Untagged spend goes into a visible backstop the owning team must clear, never smeared quietly across everyone. Showback builds visibility and trust first; chargeback moves real money once allocation is accurate and budget owners can act. One more number worth knowing: AI and GPU workloads now drive roughly eighteen percent of cloud spend at AI-forward enterprises, up from four percent in twenty twenty-three. Those costs are volatile, so monthly reporting lags too far behind provisioning decisions. Next, we look at migration, adoption, and the operating model.
giiresearch.complatformengineering.orgsnsinsider.com+22 min - 10Migration, Adoption, and the Operating ModelLet's talk about migration, adoption, and the operating model. Pilot one golden path first, then migrate wave by wave, and keep a documented return route for exceptions so a justified deviation stays observable instead of becoming shadow IT. Budget fifteen to twenty-five percent of the platform budget for adoption, because adoption is work, not communications. On staffing, industry norms converge around one platform engineer per fifteen to twenty-five developers, roughly two to six percent of engineering headcount. Choose the lower ratio when your stack is polyglot and deployments are complex; accept the higher ratio when your architecture is standardized. Run platform as a product: a product manager, a roadmap, published SLOs, and escape hatches with owners and review dates. And handle the thirty percent long tail deliberately. If every exception lands in a ticket queue, the team that was meant to remove toil becomes the source of it. Next, measuring success: DORA, developer experience, and adoption.
cncf.iocubepath.comsystem-design.space+22 min - 11Measuring Success: DORA, Developer Experience, and AdoptionNow let's talk about measuring success, because a platform you can't measure is a platform you can't defend. Start with a week-zero baseline: lead time, ticket volume, and developer satisfaction. Without two data points, there is no delta, and your ROI story becomes an anecdote. Use the five DORA metrics: deployment frequency, lead time, change failure rate, recovery time, and rework rate, tracked per service, not blended across the org. Then watch golden path adoption. Above eighty percent voluntary is the bar. Mandated exclusivity cut throughput by six percent, so keep the path easiest, not the only route. Add developer experience signals: time to first deploy, platform Net Promoter Score, self-service completion, and ticket deflection. Expect a J-curve, with early gains, a dip, then recovery over six to twelve months, so lead with adoption metrics while delivery data matures. Let's move into the decision workshop, choosing your platform direction.
1 min - 12Decision Workshop: Choosing Your Platform DirectionNow let's put the model to work in a decision workshop. The rule is simple: score two or three realistic options against the same weighted criteria, using your real workloads and real cost data, not vendor demos. Bring platform, security, operations, procurement, and finance into the same room, because surface-level agreement usually hides disagreement on cost, compliance, and operating burden. That disagreement is the point. Then name the top three accepted tradeoffs out loud. Choose standardization when consistency and governance outweigh team autonomy. Choose managed when you would rather pay a vendor than staff permanent maintainers. Choose central when a single path serves most workloads. Choose the opposite when agility, control, or lock-in risk matters more. Close by defining your proof of concept: scope, named owners, timeline, and a week-zero baseline of deployment frequency, lead time, and ticket volume, so you can prove impact later. Next, we'll look at common failure patterns and how to avoid them.
developers.devsquareops.comblog.localops.co+22 min - 13Common Failure Patterns and How to Avoid ThemLet's look at the failure patterns that stall platforms, and how to avoid each one. First, golden paths nobody uses. If your path needs more YAML than the alternative it replaces, developers will take the alternative. So co-design with developers and measure adoption, because the evidence shows voluntary adoption above eighty percent when a golden path is well designed, and under twenty percent when it is not. Second, the platform team as ticket-takers, which is manual provisioning disguised as DevOps. The fix is genuine self-service: developers trigger automation, your team builds it. Third, day-two drift. Golden paths diverge as Kubernetes versions and security policies change, so without drift detection and a template update process, you fragment into as many shapes as you have tenants. Fourth, shared runners and clusters without isolation widen the blast radius, so enforce workload boundaries before scale forces you to. Fifth, measurement. Nearly thirty percent of platform teams track no adoption metrics at all, which means they cannot tell which side of the adoption cliff they are on. Start with time to first deploy, deploy frequency, rollback rate, and voluntary adoption. Finally, remember that stalls are organizational. Weak sponsorship, unclear ownership, and no developer experience research explain most failures, not tooling. Choose self-service when independence matters more than control, and keep an explicit escape hatch so exceptions stay visible and owned. Let's pull this together in Key Takeaways and Next Steps.
cncf.iocubepath.comsystem-design.space+22 min - 14Key Takeaways and Next StepsLet's close on the decisions worth carrying out of this session. First, selection stays weighted and evidence-based: define pass or fail gates, such as compliance certifications and integration depth, before you sit through any demos, so polished demos can't reframe your priorities. Second, separate the control plane, the data plane, and the integration surfaces, and draw explicit trust boundaries, because that is where security reviewers and platform engineers will negotiate blast radius. Third, name and own every tradeoff out loud: you are trading developer autonomy for throughput, and centralized control for speed. Fourth, sourcing is rarely build versus buy. Buy-and-extend is the pragmatic middle ground: buy the portal, and own the machinery that touches your cloud, identity, and secrets. Fifth, baseline first, instrument your golden path, then track DORA, developer experience, and voluntary adoption. Remember that around seventy percent of platforms stall without a well-designed golden path, and adoption above eighty percent is what that investment buys. Finally, revisit the decision on a defined cadence with named owners, because platforms fail in year three when nobody owns the upgrade. So: gate ruthlessly, bound your trust zones, staff the golden path, and measure honestly. You now have a defensible framework.
developers.devsquareops.comblog.localops.co+22 min
Take the deck with you
Download this course as a file — free, no sign-up needed.
- PDF handoutEvery slide page, ready to print or share.15 pages · 3.3 MBDownload
- Narrated PowerPointThe deck that presents itself — every slide carries the digital human's narration video.15 pages · 15.5 MBDownload
- PowerPoint slidesThe full deck as a .pptx — open it in PowerPoint, Keynote, or Google Slides.15 pages · 3.2 MBDownload
Free to use in your own training — please keep the PersonWise credit page at the end.
Have your own deck? Turn it into a course
Sources consulted
Web sources consulted while building this course.
- Platform Engineering And Internal Developer Platform (IDP) - Market Share Analysis, Industry Trends & Statistics, Growth Forecasts (2026 - 2031) — giiresearch.com
- Announcing the State of Platform Engineering Report Vol 4 — platformengineering.org
- Internal Developer Platform (IDP) Market Size & Share, 2026-35 — snsinsider.com
- Platform Engineering 2026: The Rise of Internal Developer Platforms and the DevOps Evolution | Zylos Research — zylos.ai
- Platform Engineering in 2026: The Definitive Guide to the Discipline, the Role, and the Career — LevStack Blog — levstack.io
- Platform engineering maturity: From toolchain to self-service — cncf.io
- Internal Developer Platform Design Patterns - CubePath Docs | CubePath — cubepath.com
- Internal Developer Platforms: Product Thinking, Self-Service, and Golden Paths — System Design Space — system-design.space
- Internal Developer Platforms | Imrul Sheikh — imrul.tech
- Platform Engineering in 2026: The Internal Developer Platform Maturity Model · Dev Note — devstarsj.github.io
- Cloud Native platform sovereignty through multi-plane architecture | CNCF — cncf.io
- Design control, management, and data planes for resilient infrastructure | Well-Architected Framework | HashiCorp Developer — docs.hashicorp.com
- Solution Architecture Document — Stellar Platform (Internal Developer Platform) — archstandard.org
- Architecture | OpenChoreo — openchoreo.dev
- Internal Developer Platform (IDP) Reference Architectures — devops.com
- The IDP Decision: Build vs. Buy vs. Open Source for Enterprise — developers.dev
- Build vs Buy an Internal Developer Platform (2026) — squareops.com
- Build vs Buy: Internal Developer Platform Costs Compared — blog.localops.co
- Build vs. Buy Guide for Internal Developer Platforms (IDPs) — spacelift.io
- Buy vs Build vs Assemble an Internal Developer Platform | LiveWyer — livewyer.io