What a CEO needs to decide before the first agent

An AI-First organization is not the one that uses the most models. It is the one that organizes its operation so that human decisions are leveraged by agents with clear bands of authority. The mid-market CEO who approves a first agent in 2026 is no longer debating whether AI arrived. The debate is how to keep the company off the wrong side of the curve. McKinsey reports that 65% of global companies use generative AI in at least one function, a figure that stood at 33% in 2023 (McKinsey, “The State of AI”, 2024). World Economic Forum projects that 40% of the skills required at work will change between 2025 and 2030 (WEF, “The Future of Jobs Report”, 2025).

This playbook condenses, into six decisions, the roadmap a CEO of a 50 to 500 employee company needs before signing the first service order. Each decision rests on verifiable sources and translates into a concrete action.

Decision 1: define the value thesis before touching technology

The documented root cause of failure in AI initiatives is not the choice of model. It is the absence of an explicit value thesis. Boston Consulting Group reports that 74% of companies fail to scale value with AI and that the strongest predictor of success is clarity about where ROI is measured before the first deployment (BCG, “Where’s the Value in AI?”, 2024). Deloitte confirms the pattern in its annual global survey (Deloitte AI Institute, “State of Generative AI in the Enterprise”, Q4 2024).

The value thesis answers three questions:

  • Which process loses money today? Not the most visible one. The one that costs the most in team hours or in tolerated error.
  • What is it worth to solve it? In business units: tickets, documents, conversations, reconciliations per hour.
  • What changes if it works? Margin, service capacity, team time freed to sell or to serve.

The CEO who cannot answer these three questions on one page is not ready to approve an agent. They are ready to approve a diagnostic.

Decision 2: choose between building and operating

Andreessen Horowitz documents three enterprise adoption models: build, buy-and-operate, partner (a16z, “The Economic Case for Generative AI”, 2023 and subsequent reports). For mid-market, the build model is almost always the expensive and slow path. The company does not need research teams; it needs agents in production. The partner model, designing with a specialist and operating with the internal team, wins against the other two options when the horizon is 8 to 12 weeks to the first productive workflow.

The operational question is who keeps responsibility for the runtime. The answer that is defensible before a regulator and before the board is: the customer’s team keeps control, the provider designs, and both sign an adjustment cadence. The Cootradecun case illustrates the model.

Decision 3: define the three layers, not the verticals

The trap that destroys the ROI of enterprise AI is organizing the roadmap by business area. That multiplies pilots without platform economics. Gartner recommends organizing the portfolio by functional layer, not by department (Gartner, “Hype Cycle for Generative AI”, 2024).

The three layers that make up an operation with agents are:

  • Conversational. Agents on every channel: WhatsApp Cloud API, web, voice, email. Intent routing and handoff with full context to a human.
  • Operations. Workflows that execute. Strict RAG over the corpus, multi-agent orchestration, document validator, traceability.
  • Intelligence. Warehouse and dashboards over conversations, workflows and the audit ledger. Semantic modeling and role-based views.

The central orchestrator decides which layer handles each turn. The CEO does not approve agents by department. They approve capabilities by layer.

Decision 4: governance BEFORE deployment

Governance is not a committee. It is an architecture. The European AI Act and the OCDE principles agree on five concrete requirements that any deployment must meet: traceability, explainability, bias control, personal data protection and operational continuity (UE, Reglamento 2024/1689; OCDE, “AI Principles”, 2019, updated 2024). Each requirement materializes in a technical piece.

  • Traceability. Audit ledger signed per turn. Immutable. Exportable.
  • Explainability. Strict RAG with mandatory citation. No source, no answer.
  • Bias control. Pre-production testing and drift monitoring.
  • Personal data. Tenant isolation. Encryption in transit and at rest. Auditable retention.
  • Continuity. Model redundancy. Explicit human fallback.

For regulated companies in Colombia, the Superintendencia Financiera, the Superintendencia de Industria y Comercio (Habeas Data, Ley 1581 de 2012) and the UGPP impose concrete criteria. A non-financial company faces less regulatory pressure but the same operational requirement. The CEO who defers governance until after the first deployment pays for the retrofit at five times the cost.

Decision 5: metrics the board understands

The right metrics are not the ones the technical team publishes. They are the ones the board reads without a translator. Stripe documents the pattern in its API economy report: companies that report AI ROI to the board in operational metrics (tickets resolved, documents processed, cycle time) sustain the investment; those that report technical metrics (accuracy, F1, BLEU) lose the budget the following quarter (Stripe, “The Hidden Economy of API Calls”, 2024).

Four metrics a mid-market CEO should ask for from the first month:

  • Resolution without escalation. Percentage of turns where the agent closed the case without a human. In financial credit unions, the internal benchmark exceeds 80% in steady state.
  • Cycle time. From request to resolution. Reduction measured against the previous quarter’s baseline.
  • Detected hallucination rate. Cases where the agent answered without a source and QA caught it. This metric must fall month over month.
  • Cost per turn. Compute cost plus the human cost of supervision, divided by turns served. Reported weekly.

Decision 6: the deployment order

Order matters. The expensive mistake is attacking the most visible process first. The order that produces defensible ROI is the following:

  1. AI Readiness diagnostic. Two weeks. Eight dimensions assessed with academic criteria (Nortje and Grobbelaar, 2020; Holmstrom, 2022).
  2. First conversational agent. Three weeks. One channel. One area. Defined decision bands.
  3. Extension to operational back office. One to two weeks. Strict RAG over the document corpus.
  4. Intelligence layer. One week. Warehouse and dashboards over the audit ledger.
  5. Handoff and quarterly cadence. The customer’s team operates. The provider adjusts every quarter.

Eight to ten weeks, not eighteen months. McKinsey validates the cadence: companies that scale AI in short cycles, under 12 weeks per iteration, reach 2.1 times more value than those operating in annual cycles (McKinsey, 2024).

The mistake the CEO cannot afford

The expensive mistake is not choosing the wrong model. It is delegating the technical decision without understanding the framework. The CEO who approves without asking for the audit ledger, without asking for the three-layer architecture, without asking for the five governance answers, inherits a system that in six months cannot be defended before an auditor, a customer or a board. PwC and BCG agree that the CEO’s role in AI has changed: no longer a sponsor. An architect of bands (PwC, “AI Predictions”, 2024; BCG, 2024).

An AI-First company is not measured by how many agents it deployed. It is measured by how many operational decisions are signed, traceable and defensible. That is the difference between adoption and theater.

Risks the CEO must be able to name

Three risks appear in every serious regulatory framework. The CEO who can name them before approving the project can mitigate them. The one who discovers them in production pays for them.

  • Hallucination risk. The model answers with confidence what it does not know. Mitigation: strict RAG with mandatory citation and answers of the “no encontré la información” type.
  • Bias risk. Demographic disparity in automated decisions. Mitigation: pre-production fairness testing and drift monitoring.
  • Continuity risk. The agent goes down and the operation goes down with it. Mitigation: explicit human fallback and model redundancy.

What the CEO signs when approving the first agent

A CEO’s signature on an AI project in 2026 includes five elements the contract must name: an explicit value thesis, an assigned functional layer, a governance architecture, reportable operational metrics and a deployment order with visible cadence. If one is missing, the signature commits the company to an endless pilot. World Economic Forum and OCDE agree that the difference between real adoption and innovation theater lies in the presence or absence of these five elements (WEF, 2025; OCDE, 2023).

The published Xplouse evidence documents the model in financial credit unions, fintechs and legal-tech. The figures are operational, not estimated. Every documented case has an audit ledger, metrics and a described handoff.

The 24-month window

BID and the IMF agree that the 2025-2027 period defines which LATAM companies capture the productivity AI releases and which import technology without integrating it (BID, 2024; IMF, 2024). The CEO who decides late does not enter the market late. They enter a market where the learning curve already belongs to someone else. That is the honest reading of the moment.

The question is not whether AI agents will operate inside your company. The question is who will operate them, under which bands and with whose signature.

Sources cited

  1. McKinsey Global Institute, “The State of AI”, 2024.
  2. BCG, “Where’s the Value in AI?”, 2024.
  3. Deloitte AI Institute, “State of Generative AI in the Enterprise”, Q4 2024.
  4. Gartner, “Hype Cycle for Generative AI”, 2024.
  5. World Economic Forum, “The Future of Jobs Report”, 2025.
  6. OCDE, “AI Principles”, 2019 actualizados 2024.
  7. Andreessen Horowitz, “The Economic Case for Generative AI”, 2023.
  8. Stripe, “The Hidden Economy of API Calls”, 2024.
  9. PwC, “AI Predictions”, 2024.
  10. Unión Europea, Reglamento de Inteligencia Artificial 2024/1689.
  11. BID, “Inteligencia Artificial para todos”, 2024.
  12. IMF, “Global Financial Stability Report”, 2024.