Model Vault
Captured source
source ↗North Mini Code. Cohere's first model for developers.
Jan 28, 2026
6 minutes read
Introducing Model Vault: Your private platform for secure and scalable model inference
Model Vault simplifies serving and scaling Cohere models, so teams can focus on building, not infrastructure.
_Key Contributors: Manoj Govindassamy, Maxime Brunet, Jeremy Pekmez, Inna Shteinbuk, Elliott Choi_
Cohere is launching Model Vault, a dedicated, fully isolated SaaS platform for customers to run Cohere models securely, at scale, and with guaranteed performance.
Model Vault combines the reduced operational overhead of fully managed SaaS with the security and performance advantages of self-hosting. North users deploy their application within a secure VPC while offloading the maintenance-heavy model inference and scaling to their secure, cloud-based Model Vault. The result: lower cost of ownership and faster enterprise adoption.
Model Vault marks a milestone in Cohere’s journey to deliver transformative AI that solves the real-world challenges faced by enterprise customers today. Read on to learn more about what’s on offer, or get started now.
##### The convenience-control tradeoff
At the heart of enterprise AI transformation lies a fundamental constraint: how to balance speed and scale in adoption with the infrastructure control and compliance requirements of a mature, modern-day organization.
This is increasingly a problem of inference. As enterprises work to incorporate models into more workflows, teams, and products, inference is gradually replacing training as the dominant AI workload. McKinsey now projects that inference will account for the majority of AI compute by the end of the decade, even as per-unit costs continue to fall. For businesses, this shifts the core concern from funding episodic training runs to accounting for inference — a recurring cost that scales directly with increased adoption.
For most enterprises, available deployment options clearly reflect this tension. Multi-tenant SaaS platforms optimize for speed and operational simplicity by amortizing infrastructure costs across multiple customers. However, the lack of workload isolation inherent to shared environments introduces problems like noisy-neighbour effects, enforced rate limits during peak usage, and unpredictable latency. These platforms also tend to limit model configurability, provide little visibility into workload-level performance, and often fall short of the compliance standards of highly regulated enterprises.
Self-hosted deployments — whether on-premises or within a customer-managed VPC — address many of these limitations by offering greater control over infrastructure, performance characteristics, and security boundaries. The capital and operational burden underpinning that control, however, can be prohibitively costly, particularly for teams looking to scale.
With North, Cohere’s enterprise platform for building agentic AI applications, this problem of infrastructure becomes even more acute. Model serving still requires provisioning and managing hardware, but agentic workloads by their design are bursty, multifaceted, and unpredictable.

##### A new framework for scalable, secure AI deployment
Model Vault is the latest addition to Cohere’s suite of secure deployment options. It is designed to address this tradeoff between convenience and control by transferring the operational overheads of model inference to a secure, Cohere-managed cloud environment.
The result is a SaaS deployment model that removes the constraints of shared infrastructure, such as resource contention and unpredictable performance, while preserving the speed, elasticity, and ease of use associated with multi-tenant SaaS.
To be clear, there is no single deployment approach that fits every organization. For some enterprises, multi-tenant SaaS or self-hosted deployments will remain the right choice, depending on internal capabilities, data strategy, and regulatory requirements. Industries with highly standardized workflows, for example, may meet compliance needs through application-level controls alone and continue to favor shared SaaS platforms.
But for many enterprises, infrastructure management has become a growing barrier to production-grade agentic AI. They want to scale model inference capacity without scaling operational load. Model Vault is built for these teams.

##### Decouple inference from development
Model Vault is a Cohere-managed solution. That means we assume the full operational burden of production inference: deploying models, managing upgrades and dependencies, provisioning and scaling capacity, and ensuring performance and availability. We deliver Model Vault users 99.9%+ guaranteed availability, backed by production-grade SLOs and latency performance that matches or beats leading managed AI/ML deployment platforms.
This materially lowers the total cost of ownership by removing the need for customers to procure, provision, and operate GPU-backed inference infrastructure. North deployments can now run entirely on CPU-only environments, while the highest overhead ML infrastructure work is done within your Model Vault.
Model Vault lets ML teams scale workloads elastically without pre-allocating capacity or absorbing the risk of idle GPUs. Inference capacity expands and contracts with demand, eliminating the tension between _overprovisioning_ for peak performance and _underprovisioning_ that degrades latency and availability.
In simple terms, Model Vault lets Cohere absorb the operational complexity of model serving, freeing your most valuable asset — your engineers — to focus on moving agentic AI applications from experimentation to production. Inference is not your...
Excerpt shown — open the source for the full document.
Notability
notability 5.0/10Substantive post by Cohere, no traction info.