Business Models

Business Model Breakdown

Vast Data Business Model Breakdown: How AI Data Platforms Make Money

Vast Data built a $30B AI data platform by combining all-flash storage, vector search, analytics, and compute in one codebase. This breakdown explains how Vast makes money through per-gigabyte subscriptions, Neo cloud revenue sharing, and enterprise infrastructure contracts.

2026-08-206 min readEpisode 124
We charge by gigabyte per month. That becomes essentially a revenue share with our cloud computing partners.
Vast Data Business Model Breakdown: How AI Data Platforms Make Money

At a glance

Valuation
$30 billion
Revenue model
Per GB / month
Gross margin
90%+
Rule of 40
228

Vast Data is an AI data platform that helps enterprises, AI clouds, and hyperscalers store, search, and process data at scale. Founded in 2016, Vast grew from an all-flash storage company into a platform spanning file, object, and block storage, vector search, SQL analytics, streaming, and serverless compute. This breakdown explains what Vast does, how it makes money, and why its economics are unusual — based on our conversation with Jeff Denworth, Co-founder of Vast Data.

What is Vast Data?

Vast Data builds the data layer for AI infrastructure. Its original thesis was that deep learning would change the relationship between computing and data, and that traditional storage architectures would struggle with the scale and performance AI workloads require.

Vast serves:

  • Neo cloud providers such as CoreWeave, Lambda, Crusoe, and Nscale
  • Enterprises running AI infrastructure on-premises
  • Hyperscalers including AWS, Google, and Microsoft
  • Life sciences and genomics organizations
  • Quantitative trading and other data-intensive businesses
  • Government and large-scale video infrastructure deployments

The platform gives customers unified access to the same data through file, object, and block protocols, while also supporting vector databases, analytics, streaming, and serverless workloads.

The Product

Traditional distributed storage systems often assign data to individual nodes. As clusters grow, those nodes must communicate with one another, creating East-West traffic and coordination overhead. Customers then add separate systems for storage, vector search, analytics, and streaming.

Vast separates stateless compute nodes from a globally shared pool of SSDs over NVMe-oF fabrics. Every core can see the same global volume without needing to coordinate with every other node.

That architecture enables:

  • Constant-time vector search across billions or trillions of vectors
  • Linear performance scaling as clusters grow
  • Atomic updates across file, object, vector, and permission layers
  • One data layer for training, inference, analytics, and streaming
  • Similarity-based compression that can reduce flash requirements by 50–70%

The strategic product insight is not simply faster storage. It is a unified data platform that replaces a collection of separate services with one codebase.

Data Infrastructure as a Mental Model

Vast is a picks-and-shovels business for the AI economy. During a gold rush, the most durable suppliers often sell the infrastructure every participant needs rather than betting on which individual miner wins.

Vast does not need to predict which model, application, or AI cloud will dominate. It supplies the data infrastructure underneath many of them. As AI workloads grow, the amount of data that must be stored, searched, and moved grows with them.

Business Model

How the business works
Vast Data business model diagram
From customer demand to recurring revenue — and the structural advantages that compound over time.

Vast makes money through software subscriptions priced by gigabytes per month. It does not primarily sell the storage hardware itself. Instead, it provides the software platform that runs across customer infrastructure and charges according to the capacity consumed.

How Vast charges:

  • Per gigabyte per month — capacity-based subscription pricing
  • Multi-year software contracts — enterprise customers typically sign longer-term contracts and prepay upfront
  • Pay-as-you-go — short-term, multi-tenant cloud deployments can pay based on infrastructure consumed
  • Neo cloud revenue sharing — cloud partners resell Vast capacity to their customers at a higher rate and share the economics with Vast
  • Hyperscaler deployments — partnerships with AWS, Google, and Microsoft extend the platform into managed cloud environments

This model combines recurring software revenue with infrastructure-scale consumption. Prepayment improves cash flow, while usage expansion creates a natural growth path as customers add more data and workloads.

Distribution & Growth

Vast reaches customers through several channels:

  1. Neo cloud partners — approximately 60 AI cloud providers distribute Vast to GPU-cloud customers.
  2. Direct enterprise sales — large organizations deploy Vast on-premises or across hybrid environments.
  3. Hyperscaler partnerships — AWS, Google, and Microsoft partnerships expand access to cloud buyers.
  4. Infrastructure ecosystem partnerships — OEMs, ODMs, flash manufacturers, and AI hardware vendors help Vast reach large deployments.

The distribution advantage compounds when a customer uses Vast across multiple environments. Jane Street, for example, can run Vast in its own data center and in a Neo cloud while stitching those environments together through a unified platform.

Moat & Differentiation

Vast's moat is not the SSD hardware. The defensibility sits in the software and architecture:

  • Distributed architecture: stateless compute plus a shared SSD pool avoids traditional coordination bottlenecks.
  • Unified data model: file, object, block, vectors, and permissions operate on the same underlying data.
  • Full-stack platform: storage, vector search, SQL analytics, streaming, serverless compute, and workflow orchestration run from one codebase.
  • Similarity compression: fuzzy matching reduces flash consumption across large clusters.
  • Enterprise features: multi-tenancy, governance, atomicity, and security make the platform usable in production environments.
  • Cross-environment federation: customers can connect on-premises, Neo cloud, and hyperscaler deployments.

The longer customers run production data and AI workloads on the platform, the higher the switching cost becomes. Data gravity, operational familiarity, and cross-cloud federation reinforce the relationship.

Competition

Vast competes across several categories:

  • Legacy enterprise storage companies such as NetApp and Dell EMC
  • HPC and parallel file-system companies such as DDN and Hammerspace
  • Cloud-native storage and analytics services from AWS, Google, and Microsoft
  • Vector database specialists
  • Analytics platforms such as Snowflake and Databricks
  • Streaming platforms such as Kafka

Vast's positioning is that a cloud provider may need 10–20 separate services to assemble the capabilities Vast provides in one integrated platform. Its advantage is strongest when customers value unified data, consistency, performance, and deployment flexibility over selecting separate best-of-breed tools.

Why It Works (and the Risks)

The model works because AI makes data infrastructure mission-critical. Customers are willing to pay for storage software when it improves GPU utilization, reduces flash requirements, simplifies operations, and keeps data consistent across training and inference workloads. Vast's high gross margins and prepaid contracts create unusually strong software economics for an infrastructure company.

The main risks are concentrated in the same areas that create the opportunity:

  • Flash supply constraints: The market may remain constrained through 2028.
  • AI spending concentration: Growth depends heavily on AI infrastructure and Neo cloud expansion.
  • Hyperscaler execution: AWS, Google, and Microsoft partnerships are still developing.
  • Platform breadth: Expanding across storage, vector search, analytics, streaming, and compute increases execution demands.
  • Competition: Large cloud and storage companies have resources to build overlapping capabilities.

Key Questions

How does Vast Data make money?

Vast charges for software capacity on a per-gigabyte-per-month basis. Its revenue comes from multi-year prepaid enterprise contracts, pay-as-you-go cloud deployments, and revenue-sharing arrangements with Neo cloud partners that resell Vast capacity to their customers.

What makes Vast Data different from traditional storage companies?

Vast separates stateless compute from a globally shared SSD pool, allowing its cores to access data in parallel without the East-West coordination traffic common in older architectures. It also combines storage, vector search, analytics, streaming, and serverless compute in one platform.

Why is Vast Data's Rule of 40 score of 228 significant?

The Rule of 40 adds growth rate and free-cash-flow margin. A score above 40 is generally considered strong for a software company. Vast's score of 228 reflects the unusual combination of rapid growth, more than 90% gross margins, positive cash flow, and accounting profitability.

Takeaways

  • AI turned storage from a back-office utility into a strategic layer of the computing stack.
  • Usage-based infrastructure software can combine recurring revenue with expansion as customer data grows.
  • A unified platform can create more value than a collection of disconnected point products when consistency matters.
  • The strongest infrastructure moats often come from architecture, data, and workflow integration rather than commodity hardware.
  • Vast's unusual economics show why software can capture significant value even when it runs on physical infrastructure.

This breakdown is based on our conversation with Jeff Denworth on The Startup Project.

Continue reading

More business model breakdowns

View all breakdowns →