As businesses roll out agentic solutions across functions, they often face difficult decisions and trade-offs. What to build versus what to buy? Where to focus in the short-term versus long-term? How to avoid getting locked into choices that won’t age well? When implementing agentic systems at scale, the most effective response to these questions is a consistent, future-proof platform architecture that brings together internal and external business and orchestrates them with technology services.
As we’ve built QuantumBlack’s Agents at Scale platform and worked on McKinsey engagements with leading financial institutions, we’ve had to answer many of these questions ourselves, and learned powerful lessons about building effective platform architecture. In this article, we’ll share some of the most broadly relevant.
Press enter or click to view image in full size
As AI transformation accelerates, many organizations find themselves facing the generative AI (gen AI) paradox: the technology is being widely adopted, but still shows little measurable bottom-line impact. In a recent McKinsey article, this is framed as a structural imbalance between horizontal and vertical gen AI applications.
- Horizontal solutions, like employee copilots and chatbots, are easy to deploy and scale but tend to sit loosely on top of core business processes, which limits their value.
- Vertical gen AI applications, embedded into specific business functions and processes, can be genuinely transformative when they are supported by scalable, agentic capabilities that can automate complex, end-to-end workflows.
Why platform architecture matters
As they encounter the gap between gen AI’s promise and its impact, technology leaders are now looking beyond isolated tools and pilots, and toward the foundations that enable agentic systems to work in the real world. In practice, this raises a difficult (though familiar) set of questions:
- How do we capture short-term impact without creating waste or long-term technical debt?
- What architecture best supports a mesh of AI agents and traditional, deterministic capabilities across workflows?
- Which services and capabilities do we need to build now to support what we’ll need in two years?
- Is it even worth building today when a viable market solution may exist in a couple of years?
- How do we balance speed with control, and set standards for security, traceability, and observability in an agentic system that’s non-deterministic?
There are no universal answers to these questions yet. But through our work with three financial institutions deploying agentic systems at scale, we have developed a set of practical approaches that define what matters most, where teams risk over-investment, and what tends to be harder than expected.
The case studies
Throughout the article, we’ll share insights from these three engagements:
Case example 1: A leading European bank is deploying Agentic AI to increase productivity for relationship managers and free up more than 20 percent of their time by automating time-consuming credit application write-up. The plan is for a set of scalable technical foundations and reusable atomic agents to rapidly scale impact across business units, including personal lending, know your customer (KYC), risk, and software development.
Case example 2: A large financial services player built a digital factory of agents which works side-by-side with human developers, accelerating code and architectural requirements by 50 percent. Over 100 developers are using the platform daily with reusable AI agents (across Cobol and Java) to accelerate their modernization workstreams.
Case example 3: A global bank is reinventing the software development lifecycle by introducing end-to-end automated agent workflows to build applications from scratch. They have moved to daily sprint cycles, with the factory of agents working through the ‘night shift’ whilst humans review and plan during the ‘day shift’.
The case for an enterprise-level platform architecture
Agentic applications are taking over processes across the enterprise. This requires organizations to manage a broad range of agentic systems, each with bespoke needs and relying on a variety of capabilities and services. Some of these are system-specific; others, such as standards, evaluation, observability, security, and protocols, are shared across systems.
Press enter or click to view image in full size
An integrated enterprise-level agentic platform architecture will support scalability and maintainability. The architecture should include all the business and technology services needed to support the agentic systems:
- Existing in-house systems, workflows and data repositories
- Services purchased from external providers: agentic applications are broad enough that a single market solution rarely answers every need
- Selected custom-built agentic capabilities to meet bespoke organizational needs
The platform architecture provides the ‘glue’ between agentic capabilities from different platforms. It integrates and orchestrates the components listed above into one coherent and flexible system and supports reuse to accelerate impact and reduce duplication and maintenance cost.
Press enter or click to view image in full size
A range of components come together to form the platform architecture, among which:
- Agentic systems are deployed across the enterprise to augment or automate business processes. We see four archetypes emerging: Enterprise productivity agents, highly custom agents, purpose-fit agents, and workflow automation solutions. Each system will be built based on the most relevant external solution (e.g., MS Copilot, ServiceNow, AgentForce, LangGraph, Google ADK).
- Agentic runtimes such as MS AI Foundry, Google Vertex, AWS Bedrock, Ark, Kagent, etc., are selected based on the need of each specific agentic application, and connected into the platform.
- Interfaces and agentic orchestration provide connectivity across systems and coordination of agents across platforms. The use of open and standard protocols enables wide-ranging integrations, allowing agents to interact with internal and external systems (e.g., different LLMs, Enterprise IT & data systems) and supports scalable multi-vendor workflows.
- Agentic shared services provide shared capabilities to ensure the safe scaling of agentic applications. Each service may rely on a mix of house-built and external solutions (e.g., Phoenix, Argo, Camunda).
This platform architecture lays the foundations for accelerated deployment of new applications by business units. It’s critical, therefore, to have clear governance that balances speed and control by defining which capabilities should be held centrally and consumed consistently across the enterprise.
The European bank in Case example 1 follows a hybrid model, where agentic use cases are owned and executed within each domain. This model promotes reuse and maintainability by gradually expanding the set of central capabilities, suchb as declarative agent definition and translation engines, a common agentic framework, and a central observability platform.
Creating a ‘composable and compostable’ architecture and building selectively
Creating an agentic platform architecture does not mean building from scratch so much as building for reuse today and replacement tomorrow by combining existing services. The ideal architectur is ‘composable and compostable’: Composable means using modular building blocks that can be combined into complex systems; compostable means that components can be replaced without redesign.
This approach is the best way to capture immediate value and establish market leadership while minimizing lock-in and technical debt. It emphasizes assembling capabilities rather than building them.
For each component in this assembly, organizations should follow standard principles for ‘build vs buy’ decisions:
- Where there are clear existing industry/market approaches, or solutions that require only minor tweaks, buying ready-built solutions can accelerate deployment and impact while reducing development risk. They’re also more likely to align to industry direction. Where CIOs decide to ‘buy’, they should look for capabilities that are modular, interoperable, and ideally open-source.
- Partner where buying is not an option. The category of enterprise-grade agentic solutions is still emerging, and many capabilities do not yet have a viable ‘buy’ option. Few of these, though, will remain whitespace for long. Partnering avoids building capabilities that are likely to become commoditized or replaced by a clear market winner in the near-to-medium term. Look to partner with either technological front-runners (to ensure strong alignment to industry progress) or providers who are actively developing and integrating holistic solutions (to share risk, cost, and reward of development).
- Where ‘buy’ and ‘partner’ are not possible, selectively build for differentiation. Differentiation is most often possible in cases where functionality is specific to a given industry, region, or locally competitive landscape. An example might be a specific approach to agentic evaluations that provides more reliable outcomes from agents, giving you a material leg-up on competitors. When buying and partnering are not possible, it’s often because highly custom solutions are needed or specific internal knowledge is required. These areas may not be differentiating, despite organizational specificity. Even so, the teams building the solution should use open-source building blocks where possible.
When partnering or building a solution which is likely to become available in the market in the medium term, CIOs should ensure that the short-term benefit outweighs the temporary technical debt it incurs. ‘Compostability’ takes on greater importance here, where the capability is integrated in such a way that it can be easily removed and replaced when an external solution appears.
Press enter or click to view image in full size
In practice, most agentic platform components are purchased, with building reserved primarily for configuration, integration, and orchestration. Within each component or layer, an organization may still decide to build certain elements, or to build rather than buy even where a fit-for-purpose market solution is available. There are two main reasons for this:
- Strategic differentiation or competitive edge. In some cases, custom-built elements can create technical capability that ensures you outperform competitors. These capabilities can include saleable products/solutions as well as enablement and the ability to execute in unique ways.
- Greater control. Given that an agentic platform could eventually power most, if not all, of an enterprise’s critical processes, CIOs may choose to prioritize retaining control over the platform and the data it uses. This becomes more likely in areas where it’s difficult to predict how the market will progress, prompting CIOs to actively avoid lock-in to a major provider. This approach can sometimes be justified: vendor-owned solutions can be hard to evolve, and their roadmaps difficult to influence when bargaining power is limited. Given how quickly this technology is evolving, keeping the flexibility to innovate can become a strategic imperative.
Key design principles as you create the platform architecture
As we’ve worked to build future-proof platforms in a variety of institutional contexts, three design principles have emerged as exceptionally useful. To create a platform architecture that is ‘composable and compostable’, apply all three whenever possible.
Principle 1: Be relentlessly protocol-focused for interoperability
For any software development effort, the ability to use multivendor workflows is a prerequisite to enabling innovation and minimizing rework. For agentic systems that interact with external agents and tools (e.g., Figma, Confluence), protocols are the primary method of achieving this.
Out of the 20+ that exist today, two main protocols are emergent: Agent2Agent Protocol (A2A) and Model Context Protocol (MCP). A2A enables direct communication between agents, while MCP has the advantage of greater modularity and reusability. There is not yet a consensus on which will become standard.
- MCP: Enables secure access by agents to external tools and data at run time.
- A2A: Enables coordination of agents across systems. A2A can securely connect agents from different vendors, enabling a cross-platform ecosystem. It can also support distributing workflows across partner platforms and thereby reduce reliance on a single vendor.
- APIs: Provide agents access to tools to perform actions — can be
public (e.g., Google Maps), partner (e.g., Salesforce, OpenAI responses), or internal APIs (e.g. to internal databases). APIs can be exposed to agents directly or via MCP.
Both protocols are evolving rapidly, which requires teams to adapt systems and enterprise integrations as they build. This also makes it useful for engineering and operations teams to familiarize themselves with the lower-level details of the most common connectivity protocols, so that security controls and approaches to scalability can be layered on top.
Using these protocols enables agents to easily collaborate and share information in a variety of environments. Successful collaboration, though, depends equally on how agents limit their sharing. Smart limitations can eliminate unnecessary noise and history from context, which can make outputs less predictable. A standardized approach to intercommunication and context isolation makes enterprise workflows more flexible and resilient, in that:
- Software teams retain the ability to move flows to other providers with minimal effort, for instance, if a provider is falling behind.
- As market solutions evolve, parts of the enterprise workflow can be easily replaced with external agentic capabilities, without needing a complete redesign.
- For critical workflows, agents which can be dynamically reprovisioned onto a different platform could help minimize disruption in the case of failure, and maintain business continuity (as shown by recent outages e.g., of AWS and Azure).
Press enter or click to view image in full size
Principle 2: Design for production from the start
A recent McKinsey State of AI report found that, while many organizations in the survey have experimented with AI agents, fewer than 10% are scaling them. Most initiatives get stuck at the PoC stage, or require major rework because their production requirements are treated as an after-thought. A better approach is to consider all aspects of the production-level solution from the outset, enabling a smooth evolution of the platform through continuous improvement.
On day one, teams should focus on building foundational capabilities which create an environment for the platform and its agentic workflows to evolve in the most effective, safest way possible. This is different from a waterfall approach to development; all components of the platform should evolve iteratively following Agile principles.
In addition to standard enterprise tool integration (e.g., cybersecurity, observability), CIOs need to put four key architectural capabilities in place from the beginning:
Press enter or click to view image in full size
1 - Agentic evaluation to maximize reliability and compliance (see details below): Like test-driven development, this supports continued optimization of built solutions, quality drop prevention, reduced risk of unexpected agentic behavior, and sustained performance across enterprise workflows. Evaluations can start as a validation tool for existing solutions, but should evolve to enable an evaluation-driven approach to the design, development and deployment of agentic systems. Evaluation methodologies must adapt to the complex, non-deterministic nature of multi-agent systems.
Example: To support its credit application, the organization in case 1 is implementing a suite of metrics to assess model performance (latency, cost), consistency of outputs and adherence to guidelines (using LLMs as judges), factual accuracy (searching for hallucinations and biases), and overall agent behavior (agent path convergence, task completion, etc.). These are being built on the Phoenix LLM tracing and evaluation platform, supporting both operational evaluation on every query and running scheduled evaluations on user-provided test sets. The bank is complementing this approach with SME-led qualitative assessments of agentic system outputs, across usability of content, productivity improvement, and user experience.
2 - Marketplaces (in platforms such as Brix) as a mechanism for storing tools, agents and workflows for use across the enterprise: This ensures that existing agentic flows are discoverable, and minimizes the risk of duplication as agentic systems scale.
Example: The organization in case 2 built agents to be reusable from the very first phase. A central CoE team was set up to develop a repository of shared reusable components (including standards for documentation, architectural artefacts). The team embedded clear governance for changes to the repository and a process for each new unit set up to always look first at using existing solutions. This significantly accelerated the five subsequent modernization efforts.
3 - A clear approach to memory management to support agentic evolution and learning: This helps optimize agentic flows over time by making better tool selections, improving data flow between agents, collecting data for regulatory purposes, and reducing costs (e.g., using vector databases). This is particularly relevant in use cases that rely on large files (e.g., audio, video) or with agents that require significant context to reach the right performance level.
Example: An American bank built a custom memory framework to support conversational agents which employees can interact with on Teams. The framework, built on LangGraph’s LangMem, delivers intelligent, persistent memory across conversations. It combines short-term and long-term memory for adaptive, context-aware interactions.
Short-term memory captures recent conversation history ensuring smooth state continuity. Long-term memory, stored in a backing database, preserves user preferences, past events, feedback, and semantic context organized by process and scoped to users or groups.
Agents can automatically extract and update memories in the background with low latency, while also supporting active (“hot path”) memory management when needed. Embedding this capability from the start enabled seamless recall, personalization, and continuous learning by agents across sessions. This significantly accelerated the five subsequent modernisation efforts.
4 - Feedback mechanisms to enable continuous improvement and ensure that configurations used in agentic solutions cater to evolving goals
Principle 3: Constantly explore new solutions and integration strategies
Given how nascent the agentic market is, any platform architecture will need to evolve with the needs of enterprise and innovations in the market. Today’s platforms need to pull together many relatively small components, not all of which are future proof or enterprise-grade yet.
Dedicating resources to R&D is a critical strategy for supporting this evolution, and brings two primary benefits:
1 - Ongoing simplification: As the market matures, more capabilities will become available through fewer components, enabling CIOs to consolidate their architecture and reduce technical debt. Staying on top of these capabilities (e.g., model advancement) will help them identify simpler solutions that are more robust, more likely to converge, more predictable and less likely to get confused. Maintenance will be simplified too, although definition and maintenance of a library of prompts and specifications is required.
Example: Models evolve rapidly, with changing engineering and prompting requirements
In the modernization effort of case #2, some agentic systems built last year used several agents and heavy prompting; today many workflows require fewer agents and in some cases one agent is proving to be enough. Staying on top of this changing landscape required maintaining a library of evolving prompts and specifications, with versions linked to each model and a defined process for maintenance and evolution of agentic systems.
2 - Strengthening and expanding the platform capability: Continued experimentation to expand the architecture with new solutions can help solve new problems as they emerge. Two examples show how this can be done effectively:
Integrating Agents with Knowledge Graphs
A key area of innovation is GraphRAG: the convergence of autonomous agents and graph databases, which preserve the semantic relationships between entities (unlike flat data stores). By configuring agents to target specific nodes to retrieve relevant embeddings and their semantic relationships, we can improve retrieval accuracy. This approach offers two distinct advantages for enterprises over standard RAG:
- Contextual Precision: Agents retrieve connected concepts, not just keyword matches, ensuring the LLM reasons with the full context.
- Auditability: Graph traversals are deterministic, improving the transparency of agent behavior, a critical factor for multi-agent orchestration. Leading graph providers are actively building support for this architecture, validating it as the standard for complex data retrieval.
Example: In a recent implementation, we mapped an IT ecosystem onto a knowledge graph comprising thousands of servers, applications, and incident reports. Agents traverse this topology to retrieve critical system state information, identifying root causes, flagging security risks, and detecting applications with low observability. By leveraging LLMs, operators can diagnose these complex systems using natural language, which agents automatically translate into executable Cypher queries. Agents can also run predictive simulations, such as taking a server node offline, to assess the cascading impact of faults on business operations.
Automating agentic workflows using workflow engines
Workflow engines reduce enterprise risk by constraining agent autonomy inside a deterministic process: explicit steps, validations, approvals, and standardized failure handling (timeouts/retries/idempotency). This improves predictability because the agent’s reasoning happens within bounded stages, with clear audit logs and controllable handoffs.
For example, when integrating workflow automation platforms like n8n and Zapier with our own (ARK) agentic runtime platform, we saw significant business potential by removing coordination overhead between operations teams and AI teams. Using the approach of maintaining composability, we showed that fitting agentic work doesn’t require a change to existing operating models. Many agents can be reused across workflows in the same way as non-agentic tools, so product teams can focus on building the flows.
Example: The bank building a new approach to SDLC (case #3) is integrating existing workflow orchestration technology (Argo CD) into its agentic platform, wrapping each agent into a deterministic workflow. This pattern provides better control over end-to-end outputs and reduces the need for intermediate human reviews in each development cycle. Each agent is inserted into a structured flow with defined steps. This includes checks of agents’ outputs, and establishes limits to what the agent can access, resulting in more controlled use of tools. The resulting workflow is highly prescriptive on the format of outputs at each stage, with precise documents and templates to remove ambiguity, enabling direct hand-offs between agents/workflow stages without errors. This is particularly relevant for highly controlled processes (e.g., KYC, software engineering) where the use of agentic squads can impact predictability and replicability of outputs.
Working with open-source solutions can be a key unlock for CIOs, as it enables staying on top of technological evolution without owning the entire maintenance of the solution. At QuantumBlack, we used this ‘composable and compostable’ architecture approach ourselves when we open-sourced our ARK platform.
Example: For the core orchestration layer, the same bank (case #3) chose to go open-source when developing their own agentic platform architecture. This gave them full ownership without commercial lock-in. They contribute back to the platform with their own developments, and the community helps maintain it to the most recent standards. They focus their efforts on building the agentic workflows which are truly differentiating for them.
Summary
Enterprises are moving rapidly from experimentation with generative AI to deploying agentic systems across core functions. But many leaders are encountering the “gen AI paradox”: high adoption with limited measurable impact. The issue is structural. Horizontal solutions, such as copilots and chatbots, scale quickly but rarely rewire end-to-end work. Capturing material value requires vertical, workflow-embedded applications supported by scalable agentic capabilities.
This article set out a future-proof enterprise agentic platform architecture: an integrated “glue layer” that connects in-house systems, data and workflows with multiple external agent ecosystems and runtimes, complemented by selective custom capabilities. The platform accelerates deployment, improves maintainability, and enables consistent governance across an expanding portfolio of agentic solutions.
We proposed a “composable and compostable” approach paired with clear partner or build when you cannot buy choices. Three design principles underpin the architecture:
- Protocol-first interoperability to enable multi-vendor workflows and portability
- Production from day one, including evaluation, marketplaces, memory management and feedback loops
- Continuous integration of emerging capabilities (such as GraphRAG and workflow engines) to increase reliability, transparency and control.
The result is a scalable platform architecture that balances speed with enterprise-grade security, traceability, and predictability.