Generative AI Platform Selection and Architecture
Begin
14 pages · ~28 min
Interactive digital-human course

Generative AI Platform Selection and Architecture

This training helps technical decision-makers evaluate generative AI platforms, covering selection criteria, architecture patterns, and key tradeoffs to make informed adoption choices.

A digital instructor presents all 14 pages. Hold “Ask” at any point and ask out loud — the answer comes from this course. No sign-up needed.

28 minFree to watchDownloads

What you’ll learn

  1. 01Generative AI Platforms: Selection, Architecture, and TradeoffsWelcome. Over the next hour, we are going to build a defensible method for selecting a generative AI platform. This is not a vendor tour. The market is crowded, and it moves fast. Gartner expects end-user spending on AI platforms and models to reach sixty-four billion dollars in twenty twenty-six, up sixty-three percent from last year. That pace means every shortlist you build has a shelf life. So we will focus on something more durable: a repeatable decision process. Here is our scope. We will look at the twenty twenty-six landscape, sketch a reference architecture, build a weighted evaluation model, cover governance, and, critically, planning for how you exit a platform once you are in it. We will also separate the field into three tiers: frontier APIs, managed application platforms, and self-hosted stacks. Each carries different economics, latency profiles, security boundaries, and lock-in risk. The through-line is build versus buy versus assemble, and how tightly your choice couples to the product roadmap. You will leave with two artifacts: a weighted decision matrix you can run against real vendors, and a reference architecture sketch you can defend in a review. Next, let us define what counts as a platform, and why the twenty twenty-six market looks the way it does.Generative AI Platforms: Selection, Architecture, and Tradeoffsavasant.comgiiresearch.comgartner.com+22 min
  2. 02What Counts as a Platform and Why the 2026 Market Looks This WayLet's start by defining the category, because the term platform is doing a lot of work right now. A generative AI platform is the governed path from a request to a model, to your context, and back to telemetry. Governance is what defines it. It is not an ML platform, and it is not application software as a service. If prompts and context can leave your boundary with no policy, no logging, and no cost attribution, that is a model API, not a platform. Now look at why the market looks this way. Gartner projects worldwide spending on AI models and platforms at sixty-four billion dollars in 2026, up sixty-three percent. Generative AI model spend alone is up one hundred seventeen percent. And the fastest-growing segment, specialized and domain-specific models, is up two hundred ten percent. That tells you the buying question has moved. It is no longer which model is smartest. It is where does our data and our agents live. So if you are architecting this, decide the control plane first. Next, the three vendor camps: frontier APIs, cloud platforms, and open weights.What Counts as a Platform and Why the 2026 Market Looks This Wayavasant.comgiiresearch.comgartner.com+22 min
  3. 03The Three Vendor Camps: Frontier APIs, Cloud Platforms, and Open WeightsLet's sort the field into three vendor camps. First, frontier labs like OpenAI, Anthropic, Google, and Mistral. They push the capability ceiling, sold mainly through an API, and you operate the least. The tradeoff: your prompts and context leave your boundary, and the provider sets price and deprecation cadence. Second, cloud platforms like Azure AI Foundry, Bedrock, and Vertex. These give you a governed environment inside your tenancy, with identity and access management already attached, and they resell frontier models alongside open-weight ones. Third, open weights like Llama, Mistral, and Nemotron. You self-host for sovereignty and predictable GPU cost, but you own serving, scaling, and model updates. Now, the camps blur deliberately. Labs also distribute through the clouds, so a real shortlist usually compares a frontier API against a platform-hosted version of the same model against a self-hosted open alternative. Pick the model for the task; pick the platform for control, governance, and exit. Coming up next, Hyperscaler Platforms Compared: Where the Operational Edges Differ.The Three Vendor Camps: Frontier APIs, Cloud Platforms, and Open Weightsciopages.combearplex.cominternative.net+22 min
  4. 04Hyperscaler Platforms Compared: Where the Operational Edges DifferLet's compare the three hyperscaler platforms and focus on where the operational edges actually differ. All three host the top commercial and open-weight models, with native identity management and observability. So the real question is fit, not features. Vertex AI offers the broadest catalog through Model Garden, aggressive Gemini and Gemma pricing, and its Agent Development Kit shipped in 2026. Bedrock gives you the easiest AWS procurement and compliance path, with near-perfect Claude parity, but a narrower catalog outside Anthropic, Titan, and Llama. Azure AI Foundry gives you OpenAI exclusivity inside your compliance boundary and deep Microsoft 365 integration, though provisioned throughput unit pricing hurts bursty loads. The practice pattern: most serious organizations run two platforms. A primary for roughly eighty percent of workloads, and a secondary for specific needs. So before you commit, ask which tradeoff your workloads can tolerate, then buy a primary and a deliberate secondary. Next, we move into architecture layers, from request ingress to governed response.Hyperscaler Platforms Compared: Where the Operational Edges Differciopages.combearplex.cominternative.net+22 min
  5. 05Architecture Layers: From Request Ingress to Governed ResponseLet's look at how a request actually moves through a production system. Seven layers, in order: ingress, gateway, context assembly, model execution, validation, observability, and a feedback store. The raw model call is only ten to fifteen percent of the work. Everything else is the machinery that makes it reliable. For example, ingress handles authentication and rate limiting before you spend a single token. The gateway resolves a model alias rather than hardcoding a specific provider. Cross-cutting planes, observability, governance, and model lifecycle, span every layer, so they cannot live inside just one. Why separate the layers? So retrieval, guardrails, or telemetry can change independently, without touching inference code. One anti-pattern to avoid: building the agent as a sidecar instead of inside the enterprise. That decision, made early, determines whether retrieval, guardrails, and telemetry stay independently changeable later. Next, the Gateway as the Highest-Leverage Layer.Architecture Layers: From Request Ingress to Governed Responseironclad.academydevfloor9.github.iocompelframework.org+22 min
  6. 06The Gateway as the Highest-Leverage LayerLet's look at the layer that gives you the most leverage: the gateway. Think of it as a thin internal service that every model call passes through. One interface, many providers behind it. Why does that matter? Because four concerns attach at one chokepoint: routing, guardrails, cost caps, and observability. You build them once. Model aliasing is the mechanism that makes this practical. Instead of hardcoding a model name in your application, you call an alias like chat-standard, and the gateway resolves it to a concrete model at runtime. Swapping providers becomes a config change, not a rewrite. Fallback chains, semantic caching, and per-tenant budgets all live here too. Semantic caching is worth calling out, because thirty to sixty percent hit rates are realistic for customer-facing apps, and every hit saves the entire downstream pipeline. Now, the production reality. Teams under-provision fallback capacity, and that is your last line of defense during an outage. And streaming locks you to one provider mid-stream, so you cannot switch once tokens start flowing. The practical implication is straightforward: a clean gateway contract contains vendor lock-in and gives you one place to enforce cost and safety. That is why a thin gateway pays for itself even in a single-model system. Next, we turn to retrieval and context engineering as the real quality lever.The Gateway as the Highest-Leverage Layerironclad.academydevfloor9.github.iocompelframework.org+22 min
  7. 07Retrieval and Context Engineering as the Real Quality LeverNow let's talk about the lever that quietly matters most: context engineering. Teams spend weeks tuning prompts, then wonder why answers miss. In production, context assembly is the layer with the biggest impact on quality, and the one most often treated as an afterthought. Context engineering means designing the entire information payload: what to include, what to exclude, what order, and how to compress when you exceed the token budget. Consider a support assistant. Pure vector search finds semantically similar passages, but it often misses exact identifiers like order numbers or part codes. Hybrid retrieval, combining vector search with keyword search such as BM25, then applying a re-ranker, recovers those exact matches. Then match retrieval mode to corpus behavior. Changing corpora need query-time retrieval. Stable, reused knowledge can be cached or pre-computed, which cuts latency and cost. The practical implication is direct: weak retrieval produces weak output, no matter which frontier model sits behind it. So treat retrieval quality and policy boundaries as first-class architecture decisions, versioned and testable on their own. Next, we'll look at evaluation criteria and a weighted decision matrix that survives review.Retrieval and Context Engineering as the Real Quality Leverironclad.academydevfloor9.github.iocompelframework.org+22 min
  8. 08Evaluation Criteria and a Weighted Decision Matrix That Survives ReviewNow let's talk about how you actually make the platform decision defensible. Start with criteria families: capability, cost, latency, scalability, security, ecosystem, operability, and strategic fit. Then normalize your scores, weight them by use case, and sensitivity-test which weights flip the ranking. That last step is what survives review. Separate hard gates from soft preferences. Data residency and regulatory class are gates. SDK ergonomics and good documentation are preferences. Don't let preferences outweigh gates. For cost, count the full stack: inference, storage, egress, evaluation, human review, and platform engineering headcount. Loaded cost, not list price, is the number that matters. Watch the bias traps. Demo-driven choice, incumbent bias, ignored switching costs, and comparing list prices instead of loaded costs. The practical implication: build your matrix before you hold the bake-off, because vendors will anchor you otherwise. Next, we'll look at cost, latency, and quality, and why that trilemma is managed, not solved.Evaluation Criteria and a Weighted Decision Matrix That Survives Reviewdigitalocean.cominfoq.compromptdojo.dev+21 min
  9. 09Cost, Latency, and Quality: Managing the Trilemma Instead of Solving ItLet's talk about cost, latency, and quality together, because they move as one system. Improve one, and the other two usually shift. So stop trying to solve the trilemma. Manage it. Start by defining service level objectives per budget: quality, p ninety-five latency, cost per success, and capacity. Here is the metric that matters most. Loaded cost per accepted result, not raw token rate. A cheap model that fails validation and retries three times is not cheap. The highest-leverage levers are straightforward: route to the cheapest sufficient tier, cache prompts, cap output length, and defer batchable work out of the interactive path. For latency, stream tokens, parallelize tool calls, shrink context, and precompute answers for common queries. The practical takeaway: pick which axis this feature lives or dies on, then make the other two acceptable at the lowest cost. Next, we look at managed versus self-hosted deployment, and where the break-even points actually fall.Cost, Latency, and Quality: Managing the Trilemma Instead of Solving Itdigitalocean.cominfoq.compromptdojo.dev+22 min
  10. 10Managed Versus Self-Hosted: Break-Even Points and Hybrid ArchitecturesNow let's talk about when managed actually stops being the right answer. Start managed for the first six to eighteen months. The simplicity is dramatic. You ship in hours with no GPU cluster, no capacity planning, no inference engineers. Then watch three signals. One: volume. Self-hosting typically wins around one million requests per month, and by ten million it wins decisively. Two: sovereignty. If data cannot leave your boundary, managed APIs are simply off the table. Three: deep customization, meaning you need to modify weights, not just prompts. Here is the trap. If your GPUs sit below roughly sixty percent utilization, your cost per inference is higher than the API you replaced. You provisioned for peak, and you are paying for idle silicon. And don't forget to price your own engineers. Managed dedicated endpoints often beat self-managed once you count forty or more engineering hours per month. That is the real line item. So what usually wins? Hybrid. Self-host a small model for high-volume routine work, route the hard cases to frontier APIs, and put both behind one router. That gives you economics where they matter and quality where it matters. The implication for your architecture: build the routing layer early, so the decision stays reversible. That flexibility is worth more than any single-vendor discount. Next, let's look at data, security, privacy, and regulatory posture as selection gates.Managed Versus Self-Hosted: Break-Even Points and Hybrid Architecturesciopages.combearplex.cominternative.net+22 min
  11. 11Data, Security, Privacy, and Regulatory Posture as Selection GatesLet's treat data, security, and regulatory posture not as a checklist, but as selection gates. Start by mapping data flow, because prompts, completions, embeddings, logs, and evaluation sets each cross boundaries differently. A log that quietly retains a customer identifier is a very different risk from an embedding stored in a managed vector index. Next, verify vendor terms in writing: training-data use, zero-retention options, and enterprise data-processing addenda. If a vendor won't substantiate a claim, treat that silence as a risk signal and document it. On regulation, Article 50 transparency has been live since the second of August, 2026, and GPAI enforcement is active. Annex III high-risk rules arrive on the second of December, 2027, and Annex One on the second of August, 2028. For agents, disclosure must survive routing: identify as artificial and name the principal represented. Controls to insist on: tenant isolation, secret management, injection and exfiltration defenses, and six-month log retention. Practically, if a platform can't evidence these, it's a downstream liability you own. Next, observability, evaluation, and release engineering for non-deterministic systems.Data, Security, Privacy, and Regulatory Posture as Selection Gates2 min
  12. 12Observability, Evaluation, and Release Engineering for Non-Deterministic SystemsLet's talk about the discipline that keeps non-deterministic systems trustworthy: observability, evaluation, and release engineering. Start with a hard truth. Standard application monitoring will not save you here. You can have zero errors, ninety-nine point nine percent uptime, and still ship confidently wrong answers. Semantics are the failure mode, not availability. So you need two halves of one loop. Offline evals, a fixed test set scored before release, and online observability, traces and quality signals from real traffic. Evals tell you if a version is good. Observability tells you what is actually happening. Reach for the cheapest method first. Code assertions for format and safety. LLM as a judge, calibrated against human labels, for subjective quality. Human review for high-stakes flows. Now, prompts, models, retrieval, and tools ship as one pinned manifest. Pin dated model snapshots, never floating aliases, or the provider's release becomes yours with no gate. And gate on risk slices, not aggregate scores. A change can raise the average and break the ten cases your biggest customer depends on. Finally, canaries should watch behavioural signals: user corrections, escalations, refusals, not just errors. The practical implication? Treat every prompt or model change as a release worth gating. Next, we move to the operating model, team topology, and FinOps governance for token spend.Observability, Evaluation, and Release Engineering for Non-Deterministic Systems2 min
  13. 13Operating Model, Team Topology, and FinOps Governance for Token SpendNow let's talk about the operating model that makes all of this sustainable. The proven pattern is centralized enablement with federated execution. A small central team owns the shared platform contract: the gateway, the evaluation harness, and the guardrails. That contract should ship in days, not quarters, because every team behind it inherits attribution, quotas, and safety for free. Execution stays federated through embedded champions inside product teams. Next, stamp every request at the gateway with team, project, and agent tags. Tokens carry no tags of their own, so attribution has to be built into the request path. On chargeback, resist the naive model. Billing teams by raw token consumption punishes the wrong behavior. Charge back on cost per validated outcome, like cost per resolved ticket. A team at forty cents per ticket and falling deserves more capacity; a team at four dollars and rising needs a different conversation. Finally, staff for retrieval engineering, evaluation design, AI security, inference operations, and FinOps. Next, we turn to migration, exit planning, and the decision workshop.Operating Model, Team Topology, and FinOps Governance for Token Spend2 min
  14. 14Migration, Exit Planning, and the Decision WorkshopLet's close on the decisions that outlive this workshop. Your true integration cost is hidden in the surfaces: identity provider, data platform, API gateway, MLOps registry, SIEM, and your IT service management queue. Each one you must wire, audit, and staff. So build for exit while you build for speed. Put a gateway in front of providers, keep a vendor independent eval harness, version your prompts, and confirm you can export embeddings and vector indexes. Then migrate deliberately: run in parallel, shadow evaluate against live traffic, cut over by segment, and agree the rollback criteria before you flip anything. The anti-patterns are quiet. Scattered provider SDK calls, indexes you cannot export, and no eval suite to re-tune when a model deprecates. Finally, take three workloads into the decision workshop, a regulated assistant, a high volume consumer feature, and an internal copilot, and score them on the same criteria. You have a repeatable method now, not a one model bet. Thanks for staying with this, and go architect your next decision with that exit plan already drawn.Migration, Exit Planning, and the Decision Workshopciopages.combearplex.cominternative.net+22 min

Take the deck with you

Download this course as a file — free, no sign-up needed.

Free to use in your own training — please keep the PersonWise credit page at the end.

Have your own deck? Turn it into a course

Sources consulted

Web sources consulted while building this course.