A VP of strategy is deciding whether to enter a new market. A plant manager is trying to work out why one production line keeps drifting out of tolerance. A treasury analyst is stress-testing a hedging position against a rate scenario nobody has modeled before. None of them have a colleague who can answer the question in ten minutes, so all three open a browser tab, paste in the context, and ask a third-party model what to do.

The answer comes back quickly. So does everything they typed to get it: the market entry thesis, the tolerance data, the hedge structure, the customer names, the internal cost assumptions, and the fact that the company is considering the move at all. None of it stays inside the building, and none of it was classified before it left.

In a recent conversation on the Machine Dreams podcast, Shakudo co-founder and CEO Yevgeniy Vahlis walked through why that pattern has become the central problem in enterprise AI, what sovereign AI means as a governance model rather than a product category, and what it actually takes to bring models, data and controls inside your own perimeter.

<figure class="cms-video" style="margin:2rem 0;max-width:100%;"><div style="position:relative;padding-bottom:56.25%;height:0;overflow:hidden;"><iframe src="https://www.youtube.com/embed/MlCiD3VTNOs" title="Sovereign AI and Bringing AI In-House with Yevgeniy Vahlis" style="position:absolute;top:0;left:0;width:100%;height:100%;border:0;" scrolling="no" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="true"></iframe></div></figure>

## What sovereign AI actually means

Vahlis defines the term narrowly and practically. Sovereign AI, in his framing, is:

> AI that is in the control of the decision makers of any entity, whether it is a country, a corporation, a military, whoever is making the decisions should be able to fully put under their control the AI that the members of their organization are using.

That is Yevgeniy Vahlis, co-founder and CEO of Shakudo, on the Machine Dreams podcast, and the definition is a governance one rather than a technology one. It does not name a model, a vendor, or a cloud provider. It asks a single question: can the people who own the consequences of a decision also control the system that produced it?

The reason that question has moved from philosophy seminar to procurement checklist is a change in what enterprise data is for. For most of the last thirty years, data inside a large organization was backward-looking. It fed reports, audits, regulatory filings and quarterly analysis of what had already happened. The warehouse sat beside the business and described it.

With AI on top of that data, the same records move underneath the business and start driving it. A model reading contract terms, claims history, plant telemetry and customer correspondence is not describing decisions after the fact. It is producing recommendations that become decisions, and those decisions move financial results, employees, customers and investors. Once data sits under decision-making rather than beside it, control of the model becomes an operating question rather than an IT preference.

Three consequences follow directly from that shift:

1. **The model joins the control environment.** An organization that cannot say which model produced a recommendation, which version of that model, and what data it saw cannot answer a regulator, an auditor, or its own risk committee. Model identity, prompt, retrieved context and output have to be recorded with the same discipline as any other system of record.
2. **Data classification becomes operational rather than administrative.** When data only fed reports, a misclassification was an inconvenience. When the same data feeds a decision, a misclassification is a control failure. The classification has to travel with the data into the model, enforced by the platform, instead of living in a spreadsheet maintained by a governance team.
3. **Vendor risk becomes business risk.** A provider outage used to mean a delayed report. A provider outage, a model deprecation or a pricing change now means a stalled underwriting queue, a production line that cannot be diagnosed, or a support function that cannot answer customers.

We treat sovereignty as an architectural property rather than a certification, which is why the practical questions in this post are about where inference runs, who holds the keys, and what gets logged. For a deeper look at the layers involved, see our breakdown of [sovereign AI architecture](/blog/sovereign-ai-architecture).

## The data your employees are handing over every day

Most organizations we work with have a written policy about generative AI. Very few have an accurate picture of how it is actually being used. The gap is rarely malicious. It is the predictable result of giving knowledge workers a tool that answers in seconds while the approved tooling, the data classification and the policy catch up months later.

Vahlis describes the behavior without much drama. Everyone these days, when in doubt, goes to ChatGPT or Claude and drops the question in. The questions are not trivia. They are the hard ones: should we invest in this company, here is the profile, tell me what to do.

> It's not just intellectual property that you're giving away. You're giving away your current view of your own business and the market because your employees are going to ChatGPT and Claude and typing their business problems in there to get advice.

That second category gets underestimated because it does not look like a breach. A leaked source file is a static asset with a known value and a known owner. A leaked question reveals what you are worried about this quarter, what you do not know, which assumptions you are unsure enough about to test, and where your own analysis has stalled. That is competitive intelligence about your business, generated by your own people, in near real time, and it is far more current than anything in your data warehouse.

The mechanism that turns this from an individual lapse into a structural exposure is aggregation. No human has to read your prompts for a conclusion to be drawn from them. As Vahlis puts it, a provider aggregates at scale and draws conclusions at scale "without having people look at the data. It's software that consumes all of that, analyzes it, and knows how to infer and understand what's happening."

![How strategic prompts leave the building: employee prompt to third-party API to vendor infrastructure to cross-tenant aggregation](https://cdn.shakudo.io/images/blog/sovereign-ai-bring-ai-in-house/01-shadow-ai-leak-path.png)

Consider what a provider can infer from usage patterns alone, with no human in the loop and no need to attribute anything to a named customer:

- **Sector-level intent.** A cluster of questions about a specific regulatory change, a specific acquisition target, or a specific supply constraint tells the provider where an industry's attention is concentrated before any of it appears in public filings.
- **Timing.** The hour a question is asked, and the day it starts being asked, is a signal. A sudden increase in questions about liquidity, or about a named competitor, is visible in aggregate long before it is visible to analysts.
- **Framing.** How a question is phrased reveals the asker's model of the problem, including which options they have already rejected and which constraints they believe are fixed.
- **Follow-up chains.** Multi-turn conversations are far more revealing than single prompts. A sequence that starts with a market question and ends with a hiring plan describes a strategy in progress.
- **Absence.** What an organization stops asking about is also a signal, and it is one that no internal dashboard will ever show you.

The IBM [Cost of a Data Breach report](https://www.ibm.com/reports/data-breach) has tracked for years how much of the cost of an incident comes from lost business rather than remediation. The shadow AI case is harder to price because nothing is stolen in a way that triggers an alert. The disclosure is continuous, incremental, and entirely within the terms of service the employee accepted on the way in.

## Why enterprise AI programs stall after the demo

The failure mode is not new. Vahlis points out that enterprises and non-technical organizations have been funding innovation labs since roughly the dot-com era, and that the pattern repeats with depressing regularity: the labs produce genuinely impressive work, and the organization never pulls the trigger to take that work to market. The prototypes go on a shelf, the tooling ages out, and the effort is written off.

What is new is the cost of that pattern now that model capability improves on a quarterly cadence. A shelved prototype in 2005 was a shelved prototype. A shelved AI prototype in 2026 is a team that has learned to distrust the platform, a data access pattern that has gone stale, and a competitor that shipped the same idea eleven months ago.

The bottleneck Vahlis identifies is not model quality and not enthusiasm. It is environment:

> One of the bottlenecks is the teams do not have an end-to-end environment where they can do everything from inception, from idea all the way to full production deployments of their solution, and it has to be connected to real data.

That sentence contains the whole problem. Most enterprises can give a team a sandbox with synthetic data. Most can give a team a production deployment process measured in quarters and multiple review gates. Very few can give a team both at once, with real data, on the same platform, so that the thing demonstrated in week two is the thing running in production in month four.

![The innovation lab trap: prototypes that are never connected to real data get shelved until the tooling ages out](https://cdn.shakudo.io/images/blog/sovereign-ai-bring-ai-in-house/02-innovation-lab-trap.png)

The specific gaps we see when a lab stalls are consistent across industries:

1. **Data access is granted per experiment.** Each new project negotiates its own extract, its own copy, and its own refresh schedule. By the third project the team spends more time on access requests than on modeling, and the copies have started to diverge from the systems of record.
2. **There is no path from notebook to service.** The data science environment and the production environment have different identity models, different secrets handling, and different deployment tooling. Crossing that boundary is a project in itself, and it is usually not funded.
3. **Governance is retrofitted.** Controls, retention rules and audit logging are designed after the model works, which means the first version that reaches production is the version with the fewest controls.
4. **Nobody owns the workload after the demo.** The lab is funded to prove a concept. Nobody is funded to run it, monitor it, retrain it, or answer for it at 3am.

SpeedGauge lived in exactly that gap. The fleet safety and risk-management company has to reconcile telemetry from a long list of providers, with no industry standard for how any of it is logged, and its previous workflow ran AWS EMR Spark clusters against JSON-based geospatial and temporal data in jobs that took hours. The same work now starts in a Jupyter notebook and is promoted directly into a scheduled job. Matt Moehr, a Senior Data Scientist at the company, calls that alone "a significant upgrade from our previous developer workflow, which used AWS EMR and PySpark." The speed matters, but the structural point matters more: the notebook the scientist wrote is the job that runs in production, which is the property most platforms do not have. You can read the full account on the [SpeedGauge customer page](/customers/speedgauge).

There is a second-order effect that matters for regulated industries. Banks and other heavily regulated organizations are genuinely good at controlling data flows and preventing leaks. Vahlis notes that this competence does not automatically translate into taking full advantage of the data. The same controls that keep data in also keep value in, unless the platform gives teams a sanctioned way to work with the real thing. That is the gap the [Shakudo Platform](/platform) was built to close, and it is the evaluation question we return to in [how to evaluate sovereign AI platforms](/blog/evaluate-sovereign-ai-platforms).

Two customers show what closing that gap looks like in practice, and in both cases the model was never the hard part.

Ritual, one of the fastest-growing food ordering services in North America, had data scientists who could build models but could not get to them reliably, because every step between a notebook and a running service depended on somebody else's time. In the words of Jeff Zakrzewski, Ritual's VP of Engineering, the answer was "effectively outsourcing our DevOps and operational needs to the Shakudo umbrella as a whole," which is what lets the data team "execute tasks without getting blocked by lack of availability of data and ML engineers." Food is a low-margin business, and the constraint was never modelling talent. It was the distance between a model and production.

FlexiVan arrives at the same conclusion from the opposite direction. The intermodal logistics company has been moving freight for 70 years and runs more than 120,000 chassis, which it spent a decade instrumenting: IoT sensors from 2016, then gate cameras and computer vision that read container, chassis and plate identifiers and reconcile the three objects automatically. That replaced a manual process whose error rate, even at two percent, cascaded into misrouted containers and the customer calls that followed. As FlexiVan's CIO, Sagar Chikkala, describes the shift, "AI used to be experimental at FlexiVan. It is no longer experimental. It is operational." You can read more on the [FlexiVan customer page](/customers/flexivan).

## What bringing AI in-house actually requires

Shakudo was built to close that gap. From the beginning the goal was to build a system that is end to end and fully self-contained inside the customer organization, running on the customer's own infrastructure, whether that is their cloud account or an on-premises private cloud, so that teams can start from the most sensitive data they hold and build on it without the data ever leaving the perimeter.

That is a platform claim, so it is worth being concrete about what has to be in place. Four layers do the work, and each one has a real operational cost:

1. **Inference serving.** The runtime that hosts model weights, batches requests, manages GPU memory, and exposes an API that the rest of the stack can call.
2. **Storage and context.** Vector and relational storage, document pipelines, and the retrieval layer that decides which internal context is relevant to a given request.
3. **Routing and policy.** The control plane that decides which model handles which request, under which data classification, with which budget and which guardrails.
4. **Observability.** Traces, evaluations, cost attribution and audit records that let a risk committee and an engineering team answer the same question with the same evidence.

![Reference architecture for in-house AI: data sources, storage and context, model serving, AI gateway, agents and applications, all inside the customer perimeter](https://cdn.shakudo.io/images/blog/sovereign-ai-bring-ai-in-house/03-in-house-reference-architecture.png)

**Inference serving** is the layer most teams underestimate. Serving open-weights models well is a systems problem, not a download. You need continuous batching, KV cache management, quantization decisions that trade quality against memory, and a scheduler that packs workloads onto expensive accelerators without starving latency-sensitive endpoints. [vLLM](/integrations/vllm) is the reference implementation for high-throughput serving, and its [source repository](https://github.com/vllm-project/vllm) is worth reading before you commit to a hardware profile. For smaller models and internal tooling, [Ollama](https://github.com/ollama/ollama) provides a much simpler runtime, and Shakudo maintains an [Ollama integration](/integrations/ollama) so teams can move between the two without rewriting application code. Underneath both, GPU scheduling on Kubernetes is its own discipline: node affinity, taints, MIG partitioning and topology awareness all affect utilization, and the [Kubernetes scheduling documentation](https://kubernetes.io/docs/concepts/scheduling-eviction/) plus the [NVIDIA GPU Operator](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html) are the practical starting points.

**Storage and context** is where most of the value and most of the risk sits. Retrieval quality determines answer quality more than model choice does in most enterprise workloads, and the retrieval layer is also where data classification has to be enforced. [Qdrant](https://qdrant.tech/documentation/) is a common choice for vector search at this scale and has a [Shakudo integration](/integrations/qdrant); [Chroma](/integrations/chroma) is a reasonable fit for smaller or embedded deployments. The protocol layer matters too. The [Model Context Protocol](https://modelcontextprotocol.io/introduction) has become the standard way to expose internal systems to agents as typed tools rather than as ad hoc prompt scaffolding, which is what makes it possible to govern tool access at all.

**Routing and policy** is the layer that turns a collection of models into a governable service. A single control plane should decide which requests go to self-hosted weights, which may go to an approved external API, what gets redacted on the way out, what gets logged, and what gets blocked. [LiteLLM](https://docs.litellm.ai/) is widely used as the routing substrate, and Shakudo maintains a [LiteLLM integration](/integrations/litellm) for teams that want to keep their existing client code. This is exactly the problem the [Shakudo AI Gateway](/ai-gateway) addresses: one place where model access, rate limits, spend caps, guardrails and immutable audit trails are configured, so that adding a new model or a new team does not mean adding a new uncontrolled egress path.

![AI gateway as a control plane: agents, internal apps and business users route through one gateway to self-hosted and approved external models, with telemetry and cost tracking](https://cdn.shakudo.io/images/blog/sovereign-ai-bring-ai-in-house/04-ai-gateway-control-plane.png)

**Observability** is what makes the whole thing defensible. Traces need to follow a request from the application through retrieval, routing, inference and response, and they need to carry identifiers that a compliance team can use. The [OpenTelemetry generative AI semantic conventions](https://opentelemetry.io/docs/specs/semconv/gen-ai/) define the span attributes for this, and [Langfuse](/integrations/langfuse) is a practical open source backend for storing and reviewing them. Cost attribution belongs in the same system: if the platform cannot tell you which team spent which budget on which model, cost control becomes an argument rather than a measurement.

The honest comparison between buying AI as a service and running it inside your perimeter looks like this:

| Dimension | Third-party AI SaaS | AI inside your perimeter |
|---|---|---|
| Where prompts and outputs live | Vendor infrastructure under vendor policy | Your infrastructure under your retention and residency rules |
| Who can aggregate usage | The vendor, at machine scale | Your team, under your access controls |
| Model lifecycle | Vendor decides when a model is deprecated or replaced | You decide which weights you run and when they change |
| Cost shape | Per token, variable, correlated with adoption | Fixed capacity plus utilization, predictable at steady state |
| Data residency | Region options, still vendor-operated | Your region, your account, your network boundary |
| Audit trail | Vendor logs, vendor retention window | Immutable logs you own, exportable to your SIEM |
| Time to first result | Hours | Weeks, then faster for each subsequent workload |
| Failure domain | Shared with every other tenant | Yours alone |

Two caveats belong in the same breath. First, bringing AI in-house is not automatically safer. Vahlis is explicit that you still need your own controls, the right team in place, and careful guardrails. A self-hosted model with no evaluation harness and no access control is a worse position than a well-governed external API, because it feels safe without being safe. Second, the perimeter is not a substitute for policy. It is the mechanism that makes policy enforceable.

What a gateway actually enforces, in practice:

- **Model allowlists per team and per data classification**, so that a workload handling regulated data cannot silently fall back to an external endpoint.
- **Spend caps and rate limits** at the team, application and key level, which is what turns an unbounded cost curve into a budget line.
- **Prompt and response logging with redaction rules**, so that audit records exist without creating a second copy of the sensitive data.
- **Guardrail checks on input and output**, using frameworks such as [Guardrails AI](/integrations/guardrails-ai) or the control categories in the [OWASP Top 10 for LLM Applications](https://genai.owasp.org/llm-top-10/).
- **Immutable audit records** for every request, including the model version and the policy decision that allowed it.

## The concentration risk of a two-provider world

There is a version of the sovereign AI argument that is entirely about compliance, and there is a version that is about systemic risk. Vahlis makes the second one directly, and it is the part of the conversation we find most executives have not yet sat with:

> Imagine that AI sits with a duopoly. There's only two companies globally that have the strongest largest AI models in the world and everyone's inner thoughts are funneled to those machines.

The phrase he uses is "inner thoughts," and it is not rhetorical. The queries people type into a model sit much closer to a transcript of their reasoning than to a search history. They include the half-formed hypotheses, the options under consideration, and the things the organization does not yet know. When a large share of the global economy routes that material through two sets of machines, the concentration itself becomes the risk.

![Concentration risk: many enterprises routing through two model providers creates two failure domains for the economy](https://cdn.shakudo.io/images/blog/sovereign-ai-bring-ai-in-house/05-model-concentration-risk.png)

Why concentration is structurally worse than distribution, in the terms a risk committee already uses:

- **Blast radius.** A single misconfiguration, a single bad weight update, or a single compromised deployment path affects every dependent organization at once. Diversification is the oldest control in risk management, and a two-provider world has almost none of it.
- **Leverage.** Two vendors set pricing, deprecation schedules and acceptable-use terms for the entire market. Any organization whose critical workflows depend on them has transferred negotiating power it cannot get back quickly.
- **Opaque change.** Frontier models are updated continuously, and the customer is not told what changed. A model that behaved one way in March may behave differently in June, with no version pin available and no changelog that describes the difference.
- **Shared assumptions.** Models trained on similar data with similar objectives will fail in correlated ways. Correlated failure is exactly what risk frameworks try to avoid, and it is the default outcome of a duopoly.
- **Single regulatory surface.** When two companies serve the whole market, regulators have two doors to knock on. That is convenient for regulators and dangerous for everyone else, because it means the terms of the entire economy are negotiated in two rooms.

Vahlis's counterargument is that sovereign AI decentralizes the risk rather than eliminating it. If banks, hospitals, defense agencies and manufacturers each run models on their own infrastructure, tied to their own data and their own objectives, then different AIs are tied to different organizations and different ways of thinking, and they can potentially de-risk each other. Diversity of model provenance and model behavior is a control, in the same way that diversity of suppliers is a control in any other critical dependency.

![Decentralized sovereign AI: banks, hospitals, defense agencies and manufacturers each running models on their own infrastructure with no shared failure domain](https://cdn.shakudo.io/images/blog/sovereign-ai-bring-ai-in-house/06-sovereign-ai-decentralized.png)

This is also why we think the interesting enterprise deployments are not the ones with the largest models. Loblaw Companies Limited, Canada's largest retailer, is the useful comparison here. It operates more than 2,400 corporate and franchise stores under banners including Loblaws, Shoppers Drug Mart and No Frills, employs over 220,000 people, and in May 2026 it announced a partnership with Shakudo to standardize how the company builds, deploys and governs AI. The problem it set out to solve is the one nearly every large organization now has: teams were each assembling their own toolchain, with different ML frameworks, deployment pipelines, data access patterns, and identity and secrets handling. That produces technical duplication at scale and burns engineering time on plumbing nobody gets credit for. In the words of Charu Pujari, Loblaw's SVP of Engineering and AI, the platform "allows our teams to focus on solving real problems rather than reinventing core plumbing." Data sovereignty was part of the calculus as well, which is what makes it a materially different posture from sending the same prompts to a shared endpoint. You can read more on the [Loblaw Digital customer page](/customers/loblaw-digital).

## Regulating how AI is used, not whether it exists

The regulatory conversation usually starts in the wrong place. It starts with whether AI should exist, how fast it should be built, or which lab should be allowed to continue. Vahlis is direct about where he stands:

> I don't think regulating AI in the sense of limiting it is the right way. I don't think we should slow it down.

His reasoning is practical rather than ideological. Open-weights models are published globally, and the organizations publishing them have done a genuinely good job of building models that are competitive with the strongest commercial systems and making them available to anyone. Once that has happened, a regulation that restricts who may build or distribute a model loses most of its force, because the capability is already downloadable. In Vahlis's framing, those releases disintermediate domestic regulators: the regulator can control the entities in its jurisdiction, but it cannot control the weights.

His conclusion is that the enforceable path runs through use, not access. You cannot tell a person on the street not to download an open-weights model and use it, because that instruction is not enforceable. You can, however, regulate the decisions that a licensed institution makes about a person, and require that the institution be able to explain and defend them.

Where use-based regulation is already workable, because the decision point is defined and the institution is already supervised:

- **Loan adjudication and credit decisions**, where adverse action notices and model risk management expectations already exist.
- **Mortgage underwriting**, where the lender is licensed, the decision is documented, and the applicant has appeal rights.
- **Benefits eligibility**, where a public agency makes a determination about an individual and must be able to justify it.
- **Clinical decision support**, where the clinician remains accountable and the evidence standard is already established.
- **Insurance pricing and claims**, where rate filings and claims handling are already subject to review.

![Regulating the use of AI at defined decision points is enforceable, while regulating access to models is not](https://cdn.shakudo.io/images/blog/sovereign-ai-bring-ai-in-house/07-regulate-use-not-technology.png)

The two approaches differ in almost every dimension that matters to an operator:

| Dimension | Regulating access to models | Regulating the use of AI |
|---|---|---|
| What it targets | Who may train, release or distribute a model | Where AI may inform a decision about a person |
| Enforceability | Weak, because open weights can be downloaded anywhere | Strong, because the regulated entity is already licensed and auditable |
| Who is accountable | Model developers, often outside the jurisdiction | The bank, hospital, insurer or agency making the decision |
| Effect on development | Slows capability building across the board | Leaves model development alone and constrains specific decisions |
| Evidence required | Model internals, which are not reliably inspectable | Decision records, model version, evaluation results, human review |
| Existing precedent | Export controls on dual-use technology | Lending, insurance, clinical care and benefits regimes that already exist |

For enterprises, the practical implication is that the evidence trail is the deliverable. If a regulator asks how a decision was reached, the answer has to be a record: which model, which version, which inputs, which policy, which human review. That is the same record set the [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) asks for, and it overlaps substantially with the documentation obligations in the [EU AI Act](https://artificialintelligenceact.eu/) and its [official text](https://eur-lex.europa.eu/eli/reg/2024/1689/oj). We would rather build that trail once, in the platform, than reconstruct it under examination. Organizations in supervised sectors can see how this plays out in practice on our [financial services industry page](/industries/financial-services).

CentralReach shows what that looks like when the stakes are clinical. The company builds software for autism and intellectual and developmental disability care, a field where the population needing support is growing around nine to ten percent a year while the supply of qualified clinicians grows far more slowly. That arithmetic forces automation, but it also forces care: the AI work there turns session data into payor-compliant clinical notes that a clinician reviews and signs. The decision stays with the human, the record of how the draft was produced is complete, and the time saved is real. Chris Sullens, the company's CEO, describes the change in terms of work that used to take months and years now taking weeks or months. That is use-based governance working in practice rather than in a policy document, and you can read about it on the [CentralReach customer page](/customers/centralreach).

## Where your model was trained is now a procurement question

A question that almost never appeared in enterprise procurement reviews two years ago now appears in every serious one: where was this model trained, and under whose rules?

Vahlis is blunt about the state of knowledge. The models you get from the United States are trained differently from the models you get from China, and you do not actually know how they were trained. That is not a criticism of any particular vendor. It is a statement about the limits of what a model card, a license file and a system prompt can tell you about the data, the filtering decisions, the fine-tuning objectives and the evaluation criteria behind a set of weights.

The customers most likely to have already made this decision are defense organizations, which insist on American models for reasons that are easy to understand. Vahlis's observation is that most organizations bringing AI in-house have not started thinking about geographic origins at all. They are choosing models on benchmark scores and price per token, which are the two dimensions least connected to the question of what the model was optimized to do.

![Model provenance decision flow: data sensitivity and acceptable model origin determine whether a workload can use a hosted API or must run self-hosted weights](https://cdn.shakudo.io/images/blog/sovereign-ai-bring-ai-in-house/08-model-provenance-decision.png)

The questions worth adding to a model procurement review:

- **Which jurisdiction trained the model, and under which data protection regime?** This determines your exposure to foreign data handling rules and to future export restrictions.
- **What does the license actually permit?** Commercial use, derivative weights, redistribution and distillation are separate rights, and open weights does not mean unrestricted.
- **Is there a model card, and what does it omit?** The [Hugging Face model card guidance](https://huggingface.co/docs/hub/model-cards) describes what a good card contains. What matters for procurement is what is missing from it.
- **Can the weights be pinned to an immutable digest?** If you cannot pin a version, you cannot reproduce a decision. Provenance and artifact integrity practices from [SLSA](https://slsa.dev/) and [SPDX](https://spdx.dev/) apply directly here.
- **What is the deprecation policy?** A model that disappears on a vendor's schedule is a production incident waiting for a date.
- **Who can be held accountable if the model causes harm?** For regulated workloads, an unaccountable upstream is a control gap, not a legal technicality.

A workable internal policy is usually tiered rather than absolute, because very few organizations can run everything self-hosted on day one:

| Workload class | Data sensitivity | Acceptable model origin | Deployment pattern |
|---|---|---|---|
| Public marketing and documentation drafting | Public | Any approved commercial API | Hosted third party, gateway-mediated |
| Internal knowledge search and summarization | Confidential | Approved commercial API or self-hosted open weights | Gateway-mediated with redaction |
| Customer PII processing and support automation | Restricted | Self-hosted weights inside the perimeter | Self-hosted, no external egress |
| Regulated decisioning and defense workloads | Highest | Self-hosted weights with documented provenance | Self-hosted, air-gapped where required |

Building that tiering is easier when the model layer is pluggable. Shakudo supports open weights including the [Llama family](/integrations/meta-llama) alongside commercial endpoints, which means a workload can move tiers without a rewrite. Huntington Bank is the clearest reference point we have for how a regulated institution works through this decision. It runs more than 100 data scientists and AI practitioners, with AI spread across everything from intelligent document processing to customer-facing agents, and it found that fragmented MLOps, siloed tooling and limited cost transparency were capping what the team could ship. Moving its model portfolio off Amazon SageMaker and onto infrastructure it controlled delivered four things at once: cost visibility deep enough to budget against, decoupling from a single vendor's roadmap, the choice of running entirely in its own cloud account or on premises with full data sovereignty, and model monitoring with drift detection that a risk committee will accept. You can read the full account on the [Huntington Bank customer page](/customers/huntington-bank).

## What changes over the next five years

Asked what the next five years look like, Vahlis gives an answer that is more mundane than the surrounding discourse and considerably more demanding:

> AI will be running a lot of the business processes, the internal processes. Doesn't mean that the people will not be there. I think their jobs will change and evolve and will become more productive.

Two claims sit inside that sentence. The first is that AI becomes operational infrastructure rather than a project category. The second is that the competitive question shifts from whether you use AI to how good your in-house AI is. Once every competitor has access to the same frontier models through the same APIs, the differentiator is the data and context you can bring to bear, the workflows you have instrumented, and the quality of the system you built around the model.

By function, that looks like this:

- **Operations.** Scheduling, maintenance triage and quality analysis move from periodic review to continuous recommendation, with the model reading sensor and work-order history that never left the plant.
- **Finance.** Close, variance analysis, forecasting and contract review get faster, and the interesting shift is that the model can work from the full ledger and the full contract set rather than a summarized extract.
- **HR.** Policy questions, onboarding, and workforce planning become self-service, with the sensitive cases escalated to humans and the reasoning recorded.
- **Sales.** Account research, proposal drafting and pipeline hygiene improve, and the constraint becomes data quality in the CRM rather than model capability.
- **Engineering.** Code review, incident triage and migration work accelerate, and the governance question becomes which repositories and which production systems an agent may touch.

![In-house AI layer feeding operations, finance, HR, sales and engineering, raising output per function](https://cdn.shakudo.io/images/blog/sovereign-ai-bring-ai-in-house/09-five-year-operating-model.png)

The uncomfortable part of his forecast is not automation. It is measurement. Vahlis expects the hardest conversation to be about the use of AI for measuring the performance of the business, of its employees and of its leadership. His argument for why it is coming is that it can be objective: if you agree on the rules in advance and apply them consistently, a model can evaluate performance against them without the inconsistency that human managers bring. His argument for why it is hard is that it removes the human element from judgments that people expect to be human, and that it forces cultural changes around who gets hired and fired, how the company scales, and which strategic decisions get revisited.

We think that is the right way to frame it, and we would add one caution from operating these systems. A model can measure what is instrumented, which is rarely the same as what matters. If the metric becomes the target, the measurement system degrades quickly, and the governance burden lands on the rules rather than on the model. The organizations that handle this well will treat the rules as a versioned artifact with review, appeal and revision, not as a model output.

## A phased path to running AI in house

We would not advise a CIO to move everything at once, and the customers who have done this successfully did not. The path that works has three phases, and each one earns the next by producing evidence.

1. **Pilot on one workload with real data.** Choose a workload where the value is measurable, the data is genuinely sensitive, and the failure mode is tolerable. Run it on your own infrastructure against real systems of record, not a synthetic extract. The goal of this phase is not a demo. It is proving that your team can get from idea to a running service on real data inside your own perimeter.
2. **Move to production connected to systems of record.** Take the pilot workload live, with the identity model, secrets handling, monitoring, rollback path and cost attribution that production requires. This is the phase where most programs fail, because it is the phase that requires platform work rather than modeling work.
3. **Scale across business functions.** Once one workload is running with real governance, the second and third are dramatically cheaper. This is where the platform pays for itself, and where the operating model has to be defined: who approves a new workload, who owns a model version change, who answers for an incident.

![Three-phase rollout: pilot with real data, then production connected to systems of record, then scale across business functions](https://cdn.shakudo.io/images/blog/sovereign-ai-bring-ai-in-house/10-in-house-ai-rollout-phases.png)

The controls that have to be in place before phase two, not after:

- **Identity and access tied to your existing directory**, with per-team model allowlists and no shared API keys.
- **Evaluation harnesses with versioned test sets**, so that a model change is a measurable event rather than a leap of faith.
- **Prompt, response and decision logging with retention rules** that satisfy your regulator and your privacy office at the same time.
- **A documented model inventory** covering self-hosted weights and approved external endpoints, including provenance and license.
- **An incident process for AI failures**, including the ability to roll back to a pinned model version in minutes.

The objections a skeptical CIO will raise are fair, and we do not think they should be waved away:

- **GPU supply.** Accelerator availability is constrained and lead times are real. The honest answer is that you do not need frontier-scale hardware for most enterprise workloads, that quantized open weights run well on far less capacity than teams assume, and that scheduling and utilization discipline matters more than raw node count.
- **Talent.** Serving and operating models is a distinct skill set, and hiring for it is competitive. The answer is partly to buy the platform layer rather than build it, and partly to accept that your existing platform and SRE teams can run inference if the operational surface is small enough.
- **Keeping pace with frontier releases.** This is the objection with the most substance. Self-hosting does not mean falling behind, provided the architecture treats models as interchangeable artifacts behind a gateway. You run the best model you can host for sensitive workloads, and you keep the option to use an approved external endpoint for workloads that do not require the perimeter. The gateway is what makes that a configuration change rather than a migration.
- **Cost of idle capacity.** Reserved GPUs that sit unused are the most common way these programs lose their budget. The mitigations are real: share capacity across teams, scale down non-production environments, use smaller models for classification and routing, and measure utilization as a first-class metric.
- **Model quality gap.** Open weights are not always equal to the strongest commercial models on every task, and pretending otherwise damages credibility. The practical approach is to evaluate on your own tasks rather than on public benchmarks, and to route only the workloads that need frontier capability to frontier models.

Whitecap Resources is the clearest example we have of the phased approach working in an industrial setting. The company has grown from roughly 1,400 barrels of oil equivalent per day sixteen years ago to about 375,000 today, and its combination with Veren roughly doubled it in a single transaction. Growth at that rate turns analytics into a competitive instrument rather than a reporting function, and it means large volumes of sensitive operational data: on the order of a terabyte a month moving into relational systems, before you count the microsecond-resolution drilling, fracturing and subsurface streams that can multiply it again. James Wakelin, who has led Whitecap's 30-person data organization for over a decade, frames the goal as getting from backward-looking reporting to prediction fast enough that an operator can act on it, with a team sized for a lean producer rather than a supermajor. You can read about that deployment on the [Whitecap Resources customer page](/customers/whitecap-resources), and the same pattern applies in [manufacturing](/industries/manufacturing) and [energy](/industries/climate-energy).

## Frequently asked questions

### Q: What is sovereign AI?

Sovereign AI means the organization that owns the consequences of an AI-driven decision also controls the systems that produce it. Yevgeniy Vahlis of Shakudo defines it as AI under the control of the decision makers of an entity, whether that entity is a country, a corporation or a military. In practice that means models run on infrastructure you control, data stays inside your perimeter, and you hold the audit records. It is a governance property, not a specific product or a certification, and it does not require you to build your own foundation model.

### Q: What is the difference between sovereign AI and a private AI deployment?

A private deployment usually means a dedicated tenancy: your data is separated from other customers, but the vendor still operates the models and still sees your usage. Sovereign AI goes further. You control the weights, the serving infrastructure, the routing policy and the logs. That distinction matters when the risk you are managing is aggregation rather than exposure, because a provider that cannot read your prompts individually can still draw conclusions from them in aggregate. Private tenancy addresses confidentiality. Sovereignty addresses control.

### Q: What is shadow AI and why is it a security risk?

Shadow AI is the use of third-party AI tools by employees without organizational approval, governance or visibility. It is a security risk because the disclosure is continuous rather than episodic. Employees paste strategic questions, customer data and internal context into external endpoints, and the organization has no inventory of what left, no retention control, no audit trail and no ability to delete it. The IBM Cost of a Data Breach report has long tracked lost business as one of the largest components of incident cost, alongside detection and escalation, and shadow AI produces exactly that kind of unquantified, ongoing exposure.

### Q: Can we keep using third-party models if we bring AI in house?

Yes, and most of the organizations we work with do. The realistic architecture is a gateway that routes each request according to data classification and workload sensitivity. Public and low-sensitivity work can go to approved commercial endpoints. Regulated data, customer PII and anything that reveals strategy stays on self-hosted weights inside your perimeter. The gateway enforces the boundary, applies redaction where appropriate, and logs every request. That gives you access to frontier capability where it is safe and full control where it is required, instead of choosing one or the other.

### Q: Is open-weights AI actually safe enough for regulated industries?

Open weights are safe enough when they are deployed with the controls that regulated workloads require, and not otherwise. The relevant controls are access management, evaluation against versioned test sets, input and output guardrails, human review at defined decision points, and complete logging. Open weights also give you something commercial APIs cannot: a pinned, reproducible artifact that does not change underneath you. The tradeoff is that you own the operational burden, which is why we recommend starting with one workload and building the harness before scaling.

### Q: What does it cost to run AI on our own infrastructure?

The cost shape is different rather than automatically lower. Self-hosting converts a variable per-token bill into fixed capacity plus utilization, which is cheaper at steady high volume and more expensive at low volume with idle accelerators. The line items are GPU capacity, storage for models and context, the platform layer, and the people who operate it. The two levers that matter most are utilization, because idle reserved capacity is the main way these programs lose budget, and model right-sizing, because routing classification and extraction work to smaller models reduces the load on the expensive ones.

### Q: How long does it take to bring AI in house?

A first workload on real data typically takes weeks rather than quarters when the platform is in place, and the honest constraint is usually governance rather than engineering. The work that takes longer is the second phase: connecting to systems of record, wiring identity and secrets, adding evaluation and monitoring, and getting the audit trail to a standard your risk function accepts. Organizations that plan for a three-phase rollout, pilot then production then scale, consistently reach production faster than those that attempt a single enterprise-wide program.

## Bringing AI in house with Shakudo

Shakudo runs models, agents and governance inside your own perimeter, on your cloud account or an on-premises private cloud, so teams can build on the most sensitive data they hold without it leaving the boundary. [Shakudo Kaji](/kaji) provides the governed agent runtime, the [Shakudo AI Gateway](/ai-gateway) centralizes routing, cost control and audit across every model you allow, and the [Shakudo Platform](/platform) supplies the infrastructure underneath both. If you are moving from shadow AI to a controlled in-house capability, we can walk your team through the architecture and the rollout. [Talk to us about bringing AI in house](/contact-us).