Nebius Clickhouse Analysis
How Nebius’ stake in ClickHouse could become one of the most valuable assets in the AI infrastructure stack and a hidden engine behind the NBIS investment thesis
1. What ClickHouse Is and Why It Matters to Nebius
ClickHouse is a high-performance analytical database built for one kind of job above all else: answering large, complex questions against very large datasets with very little delay. If traditional databases are built more like filing cabinets, optimized to retrieve or update one record at a time, ClickHouse is built more like a scanning engine that can rip through entire libraries of information and return the pattern hidden inside. That difference sounds technical on paper. In practice, it is one of the reasons ClickHouse has become increasingly relevant in the modern AI and infrastructure stack.
The database began inside Yandex, where the problem was not theoretical. Yandex Metrica needed to generate analytical reports in real time from non-aggregated data that was growing rapidly, and conventional row-based systems were too slow and too expensive for that workload. ClickHouse was designed to solve that bottleneck.
Over time, it developed into a broader analytical engine capable of processing billions of rows per second on commodity hardware, with compression advantages that place in the range of roughly 12x to 20x versus uncompressed data. Those two traits, speed and efficiency, are the center of the product. They are also the reason ClickHouse has moved far beyond its origin as an internal tool.
ClickHouse is no longer just a database engineers respect. It is now a scaled infrastructure asset with real commercial weight. The company was spun out in 2021, raised capital at an initial $2B valuation, and by early 2026 had reportedly reached a $15B valuation after a $400M Series D.
The business had also surpassed 3,000 ClickHouse Cloud customers and growing ARR by more than 250% year over year. That is the kind of jump that changes how an asset should be viewed. It moves from “interesting” to “strategically consequential.”
For Nebius, this is where the story becomes especially important. Nebius is already building a vertically integrated AI cloud business with more than $3B in starting cash, active infrastructure across multiple geographies, utilization near peak levels by the end of Q2, and hyperscaler contracts that include a $17.4B to $19.4B Microsoft agreement and a ~$30B Meta contract. ARR guidance for the core business has already been framed at $7B to $9B by the end of 2026.
That is a serious infrastructure build. It is also a capital-heavy one. In that context, ClickHouse matters as both strategic ballast and financial optionality. Nebius is not simply carrying a passive minority investment. It is holding a stake in a business that sits much higher in the software margin stack and much closer to the analytics layer where AI systems are monitored, measured, and optimized.
That distinction is worth slowing down on. Nebius is building the roads, the power, and the compute clusters. ClickHouse sits closer to the control tower. One company is spending heavily to expand physical capacity. The other helps customers make sense of what happens inside that capacity in real time. When these two layers coexist under the same broader ecosystem, the relationship becomes more interesting than a simple investment mark. It starts to resemble a stack. Nebius gives customers access to infrastructure. ClickHouse helps them interrogate the data exhaust that infrastructure produces.

There is also a valuation dimension that investors should not ignore. Nebius retains an approximately 28% stake in ClickHouse. At a $15B valuation, that implies a stake worth roughly $4.2B. That is large enough to matter in a very real way. It provides a hard asset inside the Nebius story that is easier for the market to understand than future AI infrastructure utilization curves.
It also introduces a second path to value creation. Nebius shareholders are not only exposed to the execution of a capital-intensive neocloud. They also have embedded exposure to a scaled, fast-growing software infrastructure company that could eventually monetize through secondary transactions, strategic sales, or a public listing.
A clean way to frame the relationship is this:
Asset | Role in the stack | Why it matters |
|---|---|---|
Nebius core infrastructure | Compute, power, deployment, cloud orchestration | Drives large-scale AI capacity and long-duration revenue ramp |
ClickHouse | Real-time analytics, observability, AI data workflows | Adds software-layer value, margin optionality, and strategic asset support |
Combined relevance | Infrastructure plus analytical control | Creates a broader ecosystem story than raw GPU rental alone |
The market often values AI infrastructure businesses by staring at the shovels and forgetting to look at the measuring instruments. ClickHouse is one of those instruments. It helps explain why the Nebius story should not be reduced to capex, megawatts, and GPU throughput alone. There is a software intelligence layer here as well, and it carries very different economics.
That is why ClickHouse deserves its own deep dive. It is not a side note inside the Nebius story. It is one of the clearest reasons the Nebius equity can be more durable, more flexible, and more valuable than a straightforward infrastructure multiple would suggest.
Table of Contents
- What ClickHouse Is and Why It Matters to Nebius
1.1 What ClickHouse actually does
1.2 Why this asset matters inside the Nebius story - How ClickHouse Was Built Inside Yandex
2.1 The Yandex Metrica origin story
2.2 Why real-time analytics required a different database architecture
2.3 The decision to open source ClickHouse - Why ClickHouse Is Technically Different
3.1 Columnar design and vectorized execution
3.2 Why ClickHouse is fast at large-scale analytical workloads
3.3 Compression, efficiency, and cost-performance advantages - The Shift From Open Source Project to Independent Company
4.1 The 2021 spin-off
4.2 Early funding rounds and adoption curve
4.3 Why Nebius kept the stake - Leadership and Commercial Scaling
5.1 Aaron Katz and the enterprise go-to-market buildout
5.2 Alexey Milovidov and technical continuity
5.3 Yury Izrailevsky and the cloud platform layer - ClickHouse Cloud and the Move Up the Stack
6.1 Why ClickHouse Cloud changed the business model
6.2 Customer growth and enterprise adoption
6.3 How cloud monetization changed the valuation profile - ClickHouse in the AI Data Stack
7.1 Real-time analytics for AI workloads
7.2 Observability, telemetry, logs, and traces
7.3 Why this matters more in the agentic AI era - The Langfuse Acquisition and LLM Observability
8.1 What Langfuse adds
8.2 Why observability is becoming part of the core AI stack
8.3 How this deepens ClickHouse’s strategic relevance - The Financial Value of ClickHouse to Nebius
9.1 The $15 billion valuation event
9.2 What Nebius’ stake may be worth
9.3 Why ClickHouse acts as a valuation floor inside NBIS - Strategic Optionality for Nebius Shareholders
10.1 Sum-of-the-parts support
10.2 Monetization pathways and capital flexibility
10.3 Why ClickHouse matters even if Nebius remains capital intensive - The IPO Path and What It Could Mean for NBIS
11.1 Why ClickHouse looks like a credible IPO candidate
11.2 How an IPO could re-rate Nebius
11.3 What investors should watch - Final Assessment
12.1 ClickHouse as software leverage inside the Nebius ecosystem
12.2 Why this asset is still underappreciated - Downside Support, Stack Integration, and the Full-Stack Nebius Thesis
13.1 How ClickHouse caps part of the Nebius downside
13.2 How ClickHouse could integrate with Token Factory and Aether
13.3 Why this strengthens the thesis that Nebius is building an AI-native hyperscaler alternative
2. How Nebius ClickHouse Was Built Inside Yandex
The Yandex Metrica origin story
ClickHouse did not begin as a venture-backed startup searching for a problem. It began inside Yandex, under the pressure of a very specific operational need. Yandex Metrica, the company’s web analytics platform, had to generate reports in real time from enormous volumes of raw event data. That sounds routine today, largely due to the modern data stack making this kind of capability feel normal. In the late 2000s, it was a much harder engineering problem.
Yandex Metrica was already operating at meaningful scale, and the system behind it had to answer questions that were both broad and immediate. Users wanted to know where traffic came from, what pages were performing, how people moved through a site, where they dropped off, and how behavior changed over time. Each one of those questions required scanning large datasets quickly. A delay of minutes or hours would have weakened the product. The utility of analytics fades when the dashboard feels like a rearview mirror instead of a windshield.
That environment gave ClickHouse a different starting point than many infrastructure companies. It was forged in production, under live traffic, with performance requirements that could not be faked. Alexey Milovidov and the engineering team started work on the underlying project in 2009, and after roughly three years of development it was launched internally in 2012 to power Metrica. The engine was built for one clear purpose: allow analytical queries on massive raw datasets without forcing the company to pre-aggregate everything into static summaries.
That distinction is important. A lot of older analytics systems worked by heavily summarizing data in advance. That helps speed up common reports, but it narrows flexibility. Once the data has been pre-cut into a certain shape, asking a new question becomes harder. ClickHouse was built for a world where the question might change after the data already arrived. That design choice gave it a much longer runway than an ordinary internal tool.

Why real-time analytics required a different database architecture
The core issue was architectural. Traditional row-oriented databases were built for transactional workloads. They are good at recording individual events, updating single records, and handling operational tasks where you care about one customer, one payment, or one row at a time. Analytical workloads behave differently. They usually ask for patterns across huge populations of data. How many users came from a given channel. Which pages converted best. What happened by geography, device, time, source, and cohort. Those questions do not need the whole row every time. They need selected columns across an enormous number of rows.
That is where ClickHouse’s columnar design became foundational. Instead of storing data row by row, it stores it by column. In simple terms, the system keeps similar types of information grouped together. If a query only needs timestamps, URLs, and referral sources, the engine reads those fields instead of dragging along every other attribute attached to each event. That reduces wasted work and sharply improves performance when datasets become very large.
The effect is easier to picture through a physical metaphor. Imagine trying to study weather patterns by reading entire notebooks page by page, even when you only care about temperature readings. A row-based system behaves more like that notebook approach. A columnar system is closer to having all temperature readings stacked together, all wind readings stacked together, and all humidity readings stacked together. The answer becomes easier to reach since the engine touches far less irrelevant material.
ClickHouse pushed that idea further with vectorized query execution, which allowed the system to process blocks of data in batches rather than handling one value at a time. That gave modern CPUs more room to do what they are good at doing. The result was a database capable of processing billions of rows per second on standard hardware, while also compressing data aggressively enough to reduce storage demands by roughly 12x to 20x in many use cases. Speed and efficiency were connected from the beginning. One helped make the other economically viable.
This architecture was not just about raw benchmark appeal. It made a different kind of product possible. Yandex no longer had to choose between depth and responsiveness. It could keep large amounts of raw information and still return answers quickly. That combination is one of the reasons ClickHouse eventually escaped the narrow lane of web analytics and became useful across observability, security, product analytics, financial telemetry, and AI workloads.
The decision to open source ClickHouse
By 2016, Yandex made the decision to release ClickHouse as open source under the Apache 2.0 license. That move changed the future of the project. Internal infrastructure can be technically excellent and still remain boxed inside the company that built it. Open sourcing gave ClickHouse a path into the wider market, the wider developer community, and eventually the wider enterprise stack.
The timing made sense. The broader software world was becoming more comfortable with open infrastructure projects, and engineers were actively looking for systems that could handle modern analytical workloads without forcing them into expensive or inflexible legacy architectures.
ClickHouse entered that environment with real advantages. It had already been tested at scale. It had a strong performance profile. It was lightweight enough that a developer could run the same engine locally that might later scale into a far larger production environment. That kind of continuity tends to win loyalty from technical users. It lowers friction, shortens experimentation cycles, and makes adoption feel less ceremonial.
Once open sourced, ClickHouse spread beyond Yandex and began building community traction. Major technology companies including Microsoft, Meta, Uber, and Spotify became part of the broader adoption story. That widened the project’s credibility and created a bridge from developer enthusiasm to commercial demand. Open source, in this case, worked less like a charitable release and more like a carefully placed seed. It allowed the product to travel faster than a closed internal system ever could.
That decision also shaped the company that later emerged. When ClickHouse became an independent business in 2021, it was not starting from zero. It already had years of technical validation, external adoption, and developer awareness behind it. The spinout looked sudden on paper. In reality, the groundwork had been laid much earlier. The open-source release gave ClickHouse distribution before it had a salesforce, mindshare before it had formal enterprise scale, and relevance before it had a standalone cap table.
For Nebius, that history is part of why the asset is so compelling. ClickHouse was not built through financial engineering or market storytelling. It was built under load, refined through necessity, and expanded through developer adoption before becoming a high-value commercial company. That tends to produce sturdier infrastructure businesses than products invented in a pitch deck and optimized backward from a revenue target.
3. Why ClickHouse Is Technically Different
Columnar design and vectorized execution
The technical edge of ClickHouse begins with a structural choice that sounds simple and ends up changing almost everything: it stores data by columns rather than by rows.
In a row-oriented system, each full record is stored together. That works well for transactional tasks where the goal is to retrieve or update one customer, one payment, or one event at a time. Analytical workloads behave differently. They usually ask broad questions across massive datasets, and those questions often need only a handful of fields. A system optimized for rows keeps hauling entire records through memory even when the query only wants a few pieces of each one. At large scale, that becomes expensive in both time and compute.
ClickHouse approaches the problem from the opposite direction. It stores similar values together, which means the engine can pull only the fields relevant to the question being asked. If a query needs timestamps, geographic sources, and conversion values across billions of events, the system can focus its work on those columns rather than dragging the full record behind it like a cart full of unused tools. The effect is mechanical, almost industrial. Less unnecessary data moves through the machine, so the machine answers faster.
That storage model is paired with vectorized execution. Rather than processing values one by one, ClickHouse processes them in batches that fit efficiently into CPU cache and modern instruction sets. Think of the difference between moving sand grain by grain versus lifting it by the shovel. The underlying work is the same. The throughput changes completely. Vectorization allows ClickHouse to turn modern CPU hardware into something closer to an assembly line for analytical queries. That is one of the reasons the platform developed a reputation for unusually high performance on commodity infrastructure.
This combination of columnar storage and vectorized execution gave ClickHouse a technical profile that fit the world it entered. Logs, telemetry, product analytics, event streams, observability data, AI system traces, and security signals all produce huge volumes of information where the goal is usually pattern recognition across scale rather than point retrieval. ClickHouse was shaped for that terrain from the beginning.

Why ClickHouse is fast at large-scale analytical workloads
Speed in analytical systems comes from more than one trick. ClickHouse performs well at scale due to several design choices working together.
First, the engine reduces read amplification. It reads only the columns needed for a query. That seems obvious once stated plainly, though it becomes unusually powerful when the underlying table contains dozens or hundreds of attributes and the query only touches a small fraction of them.
Second, the execution engine is optimized around large scans, aggregations, and filters. Analytical workloads often involve grouping, counting, summing, ordering, and slicing over huge populations of records. ClickHouse was built with those motions in mind, which means it behaves more naturally under that workload than systems originally designed for transaction processing.
Third, the system scales well across parallel hardware. Large analytical jobs can be distributed across cores and nodes in ways that preserve the advantages of the underlying columnar model. That helps the database maintain strong performance even as the dataset expands from millions of rows to billions and beyond.
The result is a platform that has been described as capable of processing billions of rows per second on standard hardware. That kind of figure should not be treated like a magic trick or marketing slogan. It is the consequence of engineering choices that reduce wasted reads, improve CPU utilization, and align the database with the actual shape of modern analytical work.
This profile is especially relevant in environments where data arrives continuously and answers are needed quickly. Product teams want to understand usage behavior while campaigns are still running. Infrastructure teams want to inspect logs and traces while incidents are unfolding. AI teams want to monitor outputs, latency, hallucination patterns, and token behavior without waiting for delayed reporting pipelines. In those cases, speed changes the value of the system itself. A slow answer can still be technically accurate and commercially useless. ClickHouse has done well by shrinking that gap.
A useful comparison is to think of large-scale analytics as air traffic control rather than archival bookkeeping. The challenge is not simply storing what happened. The challenge is seeing the pattern clearly enough, and early enough, to respond while the system is still moving. ClickHouse fits that role well, which helps explain why it has become increasingly relevant in observability, telemetry, AI data operations, and real-time analytical infrastructure.
Compression, efficiency, and cost-performance advantages
Performance alone would not have made ClickHouse strategically interesting. The second half of the equation is efficiency.
ClickHouse uses advanced compression techniques that can reduce data footprint by roughly 12x to 20x relative to uncompressed data, depending on the workload and data type mix. That is a material advantage. Lower storage footprint reduces infrastructure cost, improves memory behavior, and often makes queries faster since less data needs to be moved and scanned. In analytical systems, those effects stack on top of one another. Efficiency at the storage layer feeds efficiency at the execution layer.
This is one reason ClickHouse has built a strong reputation for cost-performance. It does not merely answer quickly. It often answers quickly without requiring the same level of spending associated with heavier or more rigid platforms.
The benchmark comparisons included in the source material illustrate that point clearly:
System (100B rows) | Runtime | Estimated Cost | Position |
|---|---|---|---|
ClickHouse Cloud (20 nodes) | 275 seconds | $17.62 | Fast and low-cost |
Databricks (4X-Large) | 1,049 seconds | $107.69 | Slower and higher-cost |
Snowflake (4X-Large) | 1,212 seconds | $135.58 | Slower and higher-cost |
No benchmark should be treated as universal law. Workloads differ, configurations differ, and vendors tune for different environments. Even so, the directional signal is hard to miss. ClickHouse has developed a reputation for delivering strong analytical throughput at a materially lower cost profile. That combination tends to travel well in real enterprises, especially when cloud bills are already under scrutiny and AI workloads are starting to place additional pressure on budgets.
There is also a philosophical difference in how ClickHouse fits into the engineering workflow. One of the more distinctive traits described in the material is the so-called laptop test. A developer can run the same engine locally that powers much larger production environments. That continuity supports experimentation, shortens feedback loops, and reduces the friction between prototype and deployment. Many platforms talk about flexibility. Fewer preserve that feeling across both local development and large-scale production use.
Taken together, these traits explain why ClickHouse has become more than a fast niche database. It combines architectural efficiency, analytical speed, and attractive economics in a way that maps well to the needs of modern data-heavy systems. For Nebius, that is a meaningful feature of the asset. The company holds a stake in a business whose value is tied to real engineering leverage rather than narrative excess. ClickHouse became important the old-fashioned way. It solved a hard problem well enough that more and more of the stack started to route through it.
4. The Shift From Open Source Project to Independent Company
The 2021 spin-off
By the time ClickHouse became an independent company in 2021, the hard part had already been done. The engine had been built inside Yandex, tested under real production demands, opened to the outside world in 2016, and adopted by a growing set of developers and enterprises who needed a faster way to handle analytical workloads. The spin-off was not a leap into the unknown. It was more like taking a machine that had already proven itself on the factory floor and giving it its own building.
ClickHouse, Inc. was incorporated in Delaware in September 2021. That formal separation did two things at once. It gave the project a dedicated commercial structure, and it made it easier to attract outside capital, outside talent, and outside customers at a much larger scale. Open-source infrastructure can spread widely without becoming a great business. Turning technical admiration into a durable company requires a very different set of muscles: go-to-market discipline, cloud packaging, enterprise support, pricing architecture, and long-term product governance. The spin-off created the room to build those muscles.
The timing was also favorable. The broader market had entered a period where modern data infrastructure was becoming a strategic priority rather than a niche engineering preference. Enterprises were dealing with rising data volumes, more demanding observability requirements, and an increasingly obvious gap between legacy analytical systems and modern workloads. ClickHouse was well positioned for that moment. It already had credibility with technical users. What it needed next was a structure capable of commercializing that credibility without crushing the product under corporate weight.
That is a delicate transition for any infrastructure company. Some open-source projects lose their soul when they become businesses. Others never learn how to sell. ClickHouse had a better starting point than most. It had a product people already wanted, a technical reputation that had been earned rather than marketed, and a clear problem set that was getting larger over time.
Early funding rounds and adoption curve
The first major outside validation came quickly. ClickHouse launched as an independent company with a $50M investment from Benchmark and Index Ventures. That was followed by a $250M Series B that established an initial valuation of roughly $2B. For a company coming out of an infrastructure and database lineage, that was a serious signal. Venture markets do not hand out that kind of pricing for a hobby project with a cool GitHub page. It reflected a view that ClickHouse could become one of the key analytical layers in the modern data stack.
The adoption curve helped justify that optimism. Open source had already given the project reach. What the company needed to prove was that this reach could convert into enterprise and cloud monetization. Over the following years, that conversion accelerated. ClickHouse spread into prominent technology organizations including Microsoft, Meta, Uber, and Spotify. That list is important less for brand glamour than for what it implies operationally. Large-scale technical organizations usually do not adopt infrastructure tools out of curiosity. They adopt them because the performance characteristics are good enough to solve a real pain point.
The company’s cloud business became especially important in this phase. Many infrastructure companies are admired in self-managed environments and then struggle to build a scalable commercial layer on top. ClickHouse avoided getting trapped in that pattern. ClickHouse Cloud gave enterprises a managed path into the product across AWS, Azure, and Google Cloud, allowing the company to monetize beyond community enthusiasm and into recurring enterprise spend.
By early 2026, the company had reportedly surpassed 3,000 ClickHouse Cloud customers and was growing ARR by more than 250% year over year. The $400M Series D pushed the valuation to $15B. That is a very different company than the one valued at $2B in late 2021.
A simple way to frame the progression:
Stage | Milestone | What it signaled |
|---|---|---|
Internal product phase | Built inside Yandex for Metrica | Technical product-market fit under real load |
Open-source phase | Apache 2.0 release in 2016 | Developer adoption and external credibility |
Spin-off phase | Independent company formed in 2021 | Commercialization begins in earnest |
Early venture phase | $50M initial funding, then $250M Series B | Investor conviction around enterprise scale |
Expansion phase | 3,000+ cloud customers, 250%+ ARR growth | Conversion from admired tool to major infrastructure company |
Scale validation | $400M Series D at $15B valuation | Strong private market confidence in long-term relevance |
That curve is unusually important for Nebius investors. It shows that ClickHouse did not become valuable by association with AI hype alone. It moved through the harder sequence: internal utility, open-source adoption, enterprise packaging, cloud monetization, and then scale valuation. That order tends to produce sturdier businesses than stories built on thematic excitement first and commercial proof later.
Why Nebius kept the stake
Nebius retained roughly 28% of ClickHouse, and that decision now looks highly rational.
At the simplest level, the stake is valuable. At a $15B valuation, Nebius’ holding implies a mark of roughly $4.2B. That number is too large to be treated as background noise. It is one of the most concrete assets inside the broader Nebius story and one of the easiest for outside investors to anchor on when they try to assess the company beyond the moving pieces of infrastructure expansion.
The more interesting reason, though, is strategic adjacency. Nebius is building capital-intensive AI infrastructure: power, data centers, GPU clusters, orchestration, regulated cloud environments, and developer tooling. ClickHouse sits higher in the value stack, closer to the analytical and observability layer where customers inspect, optimize, and monitor what those systems are doing. One company is exposed to the physics of compute. The other benefits from the logic and telemetry that compute creates. Holding both layers gives Nebius broader exposure than a narrow infrastructure operator would normally have.
There is also a portfolio construction advantage inside the company itself. Nebius is pursuing a business that requires large capex, rapid execution, and tolerance for operational intensity. ClickHouse introduces a different kind of asset alongside that effort: software-heavy, high-growth, and structurally less burdened by the same physical buildout cycle. The relationship does not erase risk in the core Nebius business, though it does change the shape of the overall enterprise. It gives the company a valuable software stake that can support valuation, create financing flexibility, and potentially serve as a future monetization source if Nebius ever needs to raise capital without leaning as heavily on equity issuance.
That optionality comes in several forms. Nebius could eventually sell part of the stake in a secondary transaction. It could use the stake as collateral in financing structures. It could hold it through a future ClickHouse IPO and let the market provide a more transparent public valuation. Each path has different implications, though all of them are better than having no such asset at all.
In practice, the stake acts like a strategic reserve sitting beside a buildout that is otherwise very capital hungry. For shareholders, that reserve changes the quality of the Nebius story. They are not only underwriting a neocloud ramp. They are also underwriting ownership in a scaled analytical software company that already has external validation.
That is likely why the stake survived the restructuring and remained with the international entity that became Nebius. It was too important to lose, too valuable to ignore, and too connected to the future architecture of AI infrastructure to treat as a disposable legacy holding.

5. Leadership and Commercial Scaling
A database can be technically elegant and still fail commercially. That gap has buried plenty of infrastructure products over the years. Engineers build something powerful, developers admire it, and the business never quite learns how to translate performance into durable enterprise revenue. ClickHouse avoided that trap in part because its leadership structure covers the three jobs that matter most at this stage: preserve technical integrity, build a real commercial engine, and package the product into a scalable cloud platform.
That combination is one of the reasons the company’s trajectory looks more durable than a typical open-source success story. The leadership bench is not built around one charismatic founder trying to do everything. It is closer to a relay team, with each executive carrying a different leg of the race.
Aaron Katz and the enterprise go-to-market buildout
Aaron Katz brought the commercial discipline that infrastructure companies usually need once the product is already proving itself. His background includes more than two decades in enterprise software, with important roles at Salesforce and Elastic. That matters for a company like ClickHouse. Open-source adoption creates awareness. It does not automatically create quota-carrying sales teams, enterprise procurement fluency, or a repeatable global expansion model.
Katz’s presence helped move ClickHouse from admired tool to sellable platform. His time at Salesforce gave him experience inside one of the most important enterprise software growth machines in the modern era, including international expansion work that required more than just transplanting a U.S. sales motion into new markets. That kind of experience tends to matter more in infrastructure than people expect. Selling to large enterprises is less like flipping a switch and more like building a bridge while traffic is already trying to cross.
His work at Elastic is also relevant. Elastic faced a similar strategic tension: how to translate strong developer and open-source adoption into real commercial scale without breaking the product’s appeal. Katz came into ClickHouse with that pattern recognition already developed. He understood that technical popularity is useful, but only if it is converted into enterprise contracts, managed cloud adoption, and customer expansion over time.
That commercial layer became increasingly visible as ClickHouse Cloud scaled. The company moved from technical credibility to real business traction, with more than 3,000 cloud customers and ARR growth exceeding 250% year over year by early 2026. A number like that does not appear from product quality alone. It usually means the go-to-market machine has stopped being experimental and started behaving like a system.
Katz’s role is especially important because ClickHouse does not sell a simple, top-down software category. It sells performance, efficiency, and architectural leverage. Those are attractive traits, though they still need to be translated into language a budget holder understands. His job has been to turn engineering advantage into procurement logic, and the company’s valuation path suggests that translation has worked.
Alexey Milovidov and technical continuity
If Katz represents commercial scale, Alexey Milovidov represents continuity of the product’s original logic. That continuity is more important than it sounds. Infrastructure companies often lose their edge when the founding technical vision gets diluted by commercial pressure. Features multiply, priorities scatter, and the product starts behaving like a committee instead of an engine.
Milovidov is the opposite of that drift. As the original architect of ClickHouse and now its CTO, he carries forward the design philosophy that made the database useful in the first place. That includes the obsession with performance, the focus on analytical workloads rather than generic sprawl, and the refusal to let the system become bloated just because more markets can be served by becoming less opinionated.
That kind of technical continuity is one of the least flashy and most valuable traits in infrastructure. It gives customers confidence that the product will keep improving along the axis they care about, rather than wandering into a mush of adjacent features that look good in a sales deck and weaken the actual machine.
Milovidov’s continued presence also helps preserve credibility with technical users. Developers and data engineers are usually good at sensing when a product is still being run by people who understand why it exists. ClickHouse has benefited from that. The company can scale commercially without looking like it has handed the keys to people who treat the database as just another SKU.
In practical terms, his role helps keep the company anchored to the workloads where it is strongest: large-scale analytical queries, high-ingest telemetry, observability, AI traces, and real-time operational analytics. That anchor matters even more now that ClickHouse is expanding into adjacent areas like LLM observability and cloud monetization. A strong technical center helps prevent adjacency from turning into drift.
Yury Izrailevsky and the cloud platform layer
Yury Izrailevsky fills the third major leadership function: turning a powerful engine into a scalable cloud platform enterprises can actually live on.
His background includes senior roles at Google and Netflix, which is about as relevant a résumé as ClickHouse could ask for at this stage. Both environments operate at enormous technical scale, and both require deep expertise in distributed systems, reliability, and production-grade platform engineering. Those are not decorative credentials here. ClickHouse Cloud only works if the product can move from admired database to managed service without losing the performance and flexibility that made users care in the first place.
That conversion is harder than it sounds. Plenty of infrastructure products are strong in self-managed mode and stumble in the cloud. Running a production service across AWS, Azure, and Google Cloud requires a different layer of competence: orchestration, reliability engineering, multi-tenant isolation, performance tuning, deployment automation, and customer experience under real enterprise conditions. Izrailevsky’s role sits directly in that terrain.
His presence helps explain why ClickHouse Cloud has become such an important part of the company’s scaling story. The cloud layer is where adoption becomes monetization at much larger scale. It reduces operational friction for customers, expands the market beyond teams that want to self-manage infrastructure, and gives ClickHouse a cleaner recurring revenue engine than open-source goodwill alone could provide.
It also raises the strategic ceiling of the company. A fast database is valuable. A fast database delivered as a reliable, enterprise-grade, multi-cloud service is much more valuable. That is where software infrastructure starts to command larger valuations and broader adoption across mainstream enterprises rather than only technically aggressive teams.
A simple way to think about the leadership structure is this:
Leader | Core function | Strategic effect |
|---|---|---|
Aaron Katz | Enterprise sales and go-to-market scaling | Converts technical adoption into commercial growth |
Alexey Milovidov | Product architecture and technical continuity | Preserves performance edge and engineering credibility |
Yury Izrailevsky | Cloud platform and production scalability | Turns the engine into a scalable multi-cloud service |
Together, these three roles create a much sturdier company than ClickHouse would have been under a narrower leadership model. Katz builds the revenue engine. Milovidov protects the soul of the product. Izrailevsky builds the cloud layer that allows the product to scale without collapsing under its own success.
That balance helps explain why ClickHouse feels less like a great open-source project that happened to get funded and more like a serious infrastructure company now moving through the next phase of scale.

6. ClickHouse Cloud and the Move Up the Stack
Why ClickHouse Cloud changed the business model
ClickHouse became valuable first as a database engine. It became far more valuable once it began to operate as a managed cloud platform.
That distinction is more important than it looks. A strong open-source product can spread widely, though open-source adoption alone often produces uneven economics. Plenty of teams experiment, many deploy internally, and only a smaller subset ever become large commercial accounts. The revenue base can stay fragmented, support-heavy, and hard to scale. ClickHouse Cloud changed that equation. It gave the company a way to package performance into a service that enterprises could buy without carrying the operational burden themselves.
That is a move up the stack in a very real sense. Instead of only offering the engine, ClickHouse began offering the engine plus provisioning, management, scaling, upgrades, reliability, and multi-cloud accessibility. Customers no longer had to assemble the whole machine on their own. They could consume it more like a utility. For a product built around speed and efficiency, that is a powerful commercial unlock. It expands the buyer universe from infrastructure-heavy engineering teams to enterprises that want the outcome without becoming experts in the plumbing.
It also changes the company’s relationship with its users. In a self-managed model, adoption can be broad but shallow. In a managed cloud model, usage becomes easier to expand over time. More workloads can be added, more teams can be onboarded, and consumption can deepen with less friction. That leads to a better business even before it leads to a bigger one.
A good way to frame the shift is this:
Model | What the customer buys | Economic profile |
|---|---|---|
Open-source/self-managed | The database engine | Broad reach, weaker monetization, more operational burden on user |
Managed cloud | The engine plus operation, scaling, and reliability | Recurring revenue, smoother expansion, higher monetization per account |
ClickHouse Cloud did not replace the value of open source. It sat on top of it. Open source remained the distribution layer. Cloud became the monetization layer. That pairing is one of the strongest patterns in modern infrastructure software when it works well.
Customer growth and enterprise adoption
The company’s customer growth gives a sense of how well that pairing has been working. By early 2026, ClickHouse Cloud had reportedly surpassed 3,000 customers. That is a meaningful number for an infrastructure company operating in a technical category rather than a broad seat-based productivity market. More important than the raw count, though, is what it suggests about the range of adoption. ClickHouse is no longer confined to early enthusiasts or highly specialized teams. It has moved into a phase where more enterprises are comfortable putting serious workloads on the platform.
That adoption curve also fits the product’s strengths. ClickHouse works well in environments where data volumes are large, query demands are heavy, and latency still matters. Observability, security analytics, product telemetry, application logs, AI traces, event pipelines, and real-time reporting all live in that zone. These are not toy workloads. They sit close to operational decision making. Once a platform wins in those areas, expansion paths tend to open naturally.

The company’s broader adoption story reinforces that point. The ecosystem already included major names such as Microsoft, Meta, Uber, and Spotify. Those references are useful partly because they are recognizable and partly because they suggest the product has survived the scrutiny of technically demanding organizations. Enterprises rarely place critical analytical workloads onto a platform unless the performance profile is strong enough to justify the migration and the operating experience is stable enough to keep them there.
Cloud delivery makes that second part easier. A fast engine is the invitation. A reliable service is what gets the contract renewed.
There is a more subtle commercial effect here as well. Cloud products usually improve the company’s ability to observe usage patterns and shape expansion. That allows sales and product teams to understand where customers are deepening usage, where workloads are sticking, and where the next layer of monetization may come from. Over time, the relationship starts to look less like a one-time product decision and more like an expanding operational dependency. For an infrastructure company, that is where the economics start to compound.
How cloud monetization changed the valuation profile
The clearest evidence that ClickHouse Cloud changed the company is the valuation path. The spin-off began with strong early funding and an initial valuation of roughly $2B. By early 2026, after a $400M Series D, ClickHouse was valued at $15B. That kind of re-rating does not happen simply because a database is fast. Plenty of fast systems never receive that treatment. Investors began valuing ClickHouse differently once it showed signs of becoming a scaled cloud business rather than only an admired piece of infrastructure software.
The cloud model improves the valuation profile in several ways.
First, it increases revenue visibility. Managed cloud services usually produce cleaner recurring revenue than a mixed model driven mostly by support contracts and fragmented enterprise licensing.
Second, it improves expansion dynamics. Consumption can grow as customers add workloads, data volume, teams, and use cases. That creates a longer runway inside existing accounts.
Third, it raises the monetization ceiling. The company is no longer selling only software access. It is selling an operating environment around a critical analytical function.
Fourth, it broadens the strategic narrative. ClickHouse stops looking like a niche technical tool and starts looking like a core layer in the modern data and AI stack. That tends to support stronger multiples, especially when paired with rapid growth.
The growth figures help explain why investors leaned in. As of that Series D period, the company had surpassed 3,000 cloud customers and was reportedly growing ARR by more than 250% year over year. Those are the kinds of numbers that push a company into a different valuation category. The market begins to see platform behavior rather than isolated product adoption.
A simplified way to think about the valuation shift:
Phase | Business perception | Likely investor lens |
|---|---|---|
Open-source project | Strong technical product | High potential, unclear monetization |
Early commercial company | Database with enterprise traction | Promising infrastructure business |
Scaled cloud platform | Managed analytics layer with rapid recurring growth | Strategic platform asset with premium valuation potential |
For Nebius, this rise in valuation has direct implications. A roughly 28% stake in a $15B company implies around $4.2B of value. That number becomes more credible and more durable when the asset is supported by a cloud monetization engine rather than only private enthusiasm around a developer-loved product. In other words, the quality of the ClickHouse stake improved as the company moved up the stack. It is not just worth more on paper. The business under that paper mark became more legible, more scalable, and more institutionally understandable.
That is part of why ClickHouse deserves to be viewed as more than a fortunate legacy holding inside Nebius. The cloud transition gave it a stronger business model, a broader customer base, and a valuation profile that now looks much more like a strategic software platform than a specialized database vendor.
7. ClickHouse in the AI Data Stack
Real-time analytics for AI workloads
ClickHouse has become more relevant as the center of gravity in software has shifted from static applications toward systems that generate constant streams of machine-produced data. Traditional analytics was already demanding. AI workloads make that demand heavier, faster, and more continuous.
A modern AI system does not only produce an output. It produces prompts, completions, token flows, latency signatures, routing decisions, fallback events, retrieval activity, confidence signals, safety checks, infrastructure metrics, and cost data. A single interaction can create a trail of telemetry far denser than a normal application request. Once that activity scales across thousands or millions of users, the data layer starts to resemble a river delta rather than a clean pipeline.
That is where ClickHouse fits naturally. It was designed for large-scale analytical workloads where the value comes from scanning, filtering, aggregating, and interrogating huge volumes of information quickly. AI systems create exactly that kind of environment. Teams need to ask questions against fresh data, often while models are still running in production. They want to know which prompts are failing, which routes are getting slower, which models are spiking in cost, where quality is degrading, and which customer cohorts are seeing worse outcomes. Those are analytical questions at machine speed.
This is one reason ClickHouse is increasingly tied to the AI stack rather than just the older world of dashboards and reporting. It gives teams a way to inspect model behavior at scale without treating every answer like a delayed postmortem. In AI operations, delayed understanding carries a very real cost. A system can burn money, erode trust, or quietly degrade long before a traditional reporting cadence catches up. ClickHouse helps narrow that gap.

Observability, telemetry, logs, and traces
The more useful way to think about ClickHouse in AI may be less as a database in the old sense and more as a high-speed analytical surface for machine activity.
AI systems generate several overlapping forms of information:
Data type | What it captures | Why it matters |
|---|---|---|
Logs | Discrete events and system messages | Shows what happened and when |
Metrics | Numerical performance indicators | Tracks latency, throughput, error rates, token usage, cost |
Traces | Step-by-step execution paths | Reveals how requests move through multi-stage systems |
Telemetry | Broader operational signals across infrastructure and applications | Helps teams monitor behavior across the full stack |
In simpler software environments, these categories are already important. In AI systems, they become harder to manage and more valuable to understand. A large model application often includes orchestration layers, retrieval systems, model routing, guardrails, external APIs, and multiple infrastructure dependencies. A problem can emerge in any one of those layers, and the visible symptom at the front end may tell only a small part of the story.
That makes observability a first-order requirement rather than a nice extra. Teams need to inspect not just whether a request failed, but how it moved, what it touched, how long each stage took, how much it cost, and whether the output quality shifted as the path changed. ClickHouse is well suited to this environment because the underlying workload looks like what it already handles well: very large analytical datasets with frequent queries across recent and historical activity.
There is also a financial dimension here. AI systems can consume real money very quickly. Tokens, inference calls, retrieval layers, and routing complexity all have economic weight. Observability is partly about system health and partly about economic discipline. A business running production AI without strong telemetry is a bit like flying through weather with half the cockpit dark. The plane may stay in the air for a while. The confidence around where it is headed becomes a different question.
This helps explain why ClickHouse has gained traction in categories like real-time analytics, observability, and operational telemetry rather than staying confined to narrow business intelligence use cases. The database is increasingly useful in places where speed, scale, and analytical flexibility intersect, and AI happens to amplify all three.
Why this matters more in the agentic AI era
The relevance of ClickHouse rises again once AI moves from passive assistance toward more agentic behavior.
A normal software tool waits for a human to click. An agentic system can initiate, route, retry, decide, evaluate, and trigger follow-on actions with much less human supervision. That shift increases both the volume and the importance of machine-generated activity. The system is no longer producing a single answer and stopping there. It is often taking a path through a chain of decisions. Each decision leaves a trail. Each trail becomes analytically important.
That has two consequences.
First, the quantity of useful data expands sharply. Agentic systems can query more often, run more continuously, and generate more intermediate states than human-driven software. The marginal customer in the system starts to look less like an employee checking a dashboard twice a day and more like a machine process operating all the time. In practical terms, the data stack has to serve a much denser stream of reads and writes around live operational activity.
Second, trust becomes harder to maintain without deep visibility. The more autonomy a system has, the less acceptable it becomes to treat failures as black boxes. Teams need to know why an agent chose one route over another, why a workflow slowed down, where hallucinations increased, when costs spiked, and which tools or models introduced the problem. That is not only a governance issue. It is an operating issue. The moment AI systems begin acting with more initiative, observability shifts from being a support function to being part of the control system itself.
This is where ClickHouse’s position becomes more interesting. It sits at the point where machine behavior can be turned back into readable operational intelligence. That role grows more important as the software world becomes less human-paced and more system-paced.
For Nebius, this has broader strategic relevance. Nebius is building the compute-heavy side of the AI stack. ClickHouse sits closer to the analytical feedback layer that helps customers understand what their AI systems are doing once they are deployed.
One company helps run the workload. The other helps inspect the exhaust. In the agentic era, that exhaust becomes much richer, much noisier, and much more valuable. That is one reason ClickHouse looks increasingly like a core AI data infrastructure asset rather than a specialized analytical database with a fortunate timing tailwind.
8. The Langfuse Acquisition and LLM Observability
What Langfuse adds
The Langfuse acquisition pushes ClickHouse one layer deeper into the operational center of AI systems.
Before the acquisition, ClickHouse was already well positioned to serve workloads tied to logs, traces, metrics, and large-scale analytical telemetry. Langfuse adds a more specialized lens aimed directly at LLM applications. It gives teams tools to monitor prompt behavior, track model outputs, evaluate quality, compare responses across models, inspect cost and latency, and manage prompt iterations over time. In other words, it helps make model behavior inspectable rather than mystical.
That is a useful addition because LLM systems do not behave like normal software. A standard application usually produces deterministic outputs from structured logic. A language model is far more slippery. The same prompt can produce different answers. Response quality can drift. Cost can move around based on routing, context length, and model choice. Safety issues can surface unevenly. The system is powerful, though it is rarely tidy.
Langfuse helps instrument that mess.
A simple way to frame the addition is this:
Capability | What Langfuse adds |
|---|---|
Prompt tracking | Visibility into prompt versions, changes, and behavior over time |
Output evaluation | Tools to inspect response quality, consistency, and failure patterns |
Cost and latency analysis | Better understanding of inference economics and performance tradeoffs |
Tracing for LLM workflows | Visibility across multi-step model pipelines and agent chains |
Prompt management | A cleaner operational layer for iterating and testing prompts in production |
That set of capabilities pushes ClickHouse closer to the actual workflow of teams building LLM-based products. It is no longer only the place where machine telemetry can be stored and queried at speed. It becomes part of the operational interface for understanding how AI applications behave in the wild.
There is also a timing advantage here. LLM tooling is still early. The category has not fully hardened. That gives companies room to establish useful positions before the stack settles into more durable patterns. By acquiring Langfuse, ClickHouse moved early into a layer that is likely to become increasingly important as AI applications mature from demos into production systems.

Why observability is becoming part of the core AI stack
Observability used to be treated as support infrastructure, something teams layered on after the product was already built. In AI, that view ages poorly.
The reason is simple. AI systems are probabilistic, multi-stage, and expensive enough to punish blind operation. A business can tolerate weak observability in some traditional software environments for longer than it should. In AI, the cost of flying blind rises quickly. Teams need to understand not only whether a request completed, but why the model answered as it did, which prompt version was active, which retrieval path was used, how long it took, how much it cost, and whether output quality is improving or degrading.
That turns observability into part of the operating core.
In practical terms, AI systems increasingly need a stack that looks something like this:
Layer | Function |
|---|---|
Compute infrastructure | Runs training and inference workloads |
Model layer | Produces reasoning, generation, and prediction |
Application layer | Connects model output to users and business logic |
Observability layer | Tracks quality, latency, cost, routing, and behavior across the system |
Governance and control | Enforces evaluation, safety, and accountability |
The observability layer has become more central as systems become more autonomous. An AI application that only answers occasional user questions is one thing. An AI system that routes tasks, triggers actions, manages workflows, or chains tools together creates a much denser operational surface. Each layer introduces more places where quality can break, latency can balloon, or cost can leak.
That is why LLM observability is starting to look less like monitoring and more like instrumentation for decision systems. Teams are not only trying to keep systems online. They are trying to keep them legible. In the AI era, legibility is part of reliability.
This trend should strengthen over time rather than weaken. As enterprises move beyond experimentation and into broader production deployment, the questions become more practical and less romantic. Which workflows are worth the spend. Which model actually performs better on a given task. Where are hallucinations clustering. Which prompt update improved conversion and which one quietly hurt it. Those are the questions of a real operating environment. Observability is how those questions get answered.
How this deepens ClickHouse’s strategic relevance
The Langfuse deal deepens ClickHouse’s relevance by giving it a more direct role in the day-to-day management of AI systems.
Before, ClickHouse was already useful as a high-speed analytical backend for telemetry-heavy environments. After adding Langfuse, it moves closer to becoming part of the control surface for AI applications themselves. That shift is strategically important. Infrastructure companies usually gain value when they move from being useful to being hard to remove. The closer a product gets to operational visibility, evaluation loops, and workflow understanding, the more deeply it embeds into the customer’s system of record.
That makes ClickHouse more than a fast analytical engine. It starts to look like a broader intelligence layer for production AI.
This also sharpens the fit between ClickHouse and Nebius. Nebius is building out the compute-heavy side of the stack: data centers, GPU clusters, orchestration, regulated AI cloud environments, and the infrastructure needed to run large-scale workloads. ClickHouse, especially with Langfuse inside the fold, moves further into the observational and evaluative layer where customers inspect what those workloads are producing.
The relationship starts to look increasingly complementary:
Nebius layer | ClickHouse plus Langfuse layer |
|---|---|
Provides compute capacity | Provides analytical visibility into AI system behavior |
Supports training and inference at scale | Supports prompt, output, latency, and cost analysis |
Builds the engine room | Helps operate the dashboard and instrumentation |
Monetizes physical AI infrastructure | Monetizes the data exhaust and control layer around AI usage |
That makes the ClickHouse stake more strategically interesting than a simple financial holding. It gives Nebius exposure to a category likely to gain importance as AI systems become more operational, more autonomous, and more expensive to run incorrectly.
The deeper point is that observability in AI is becoming less like a rearview mirror and more like a navigation system. Langfuse helps ClickHouse step into that role. And once a company becomes part of the navigation system, it usually earns a more durable place in the stack.
9. The Financial Value of ClickHouse to Nebius
The $15 billion valuation event
The financial importance of ClickHouse to Nebius became much harder to ignore once ClickHouse was valued at $15B in its January 2026 Series D round. That round reportedly raised $400M and included participation from Dragoneer, Bessemer, GIC, and T. Rowe Price. A valuation at that level changes the conversation. ClickHouse stops looking like a promising private asset and starts looking like one of the more substantial software holdings attached to any public AI infrastructure story.
That pricing was not assigned in a vacuum. By that stage, ClickHouse had already moved well beyond the stage of open-source enthusiasm and early enterprise curiosity. The company had surpassed 3,000 ClickHouse Cloud customers and was described as growing ARR by more than 250% year over year. Those are scaling numbers, not concept numbers. They suggest a business that has found real commercial traction in a category where performance, efficiency, and cloud monetization reinforce one another.
For Nebius, the significance of the $15B event is straightforward. It provided a more concrete market marker for an asset that had previously been easier to appreciate conceptually than financially. Investors could now point to a recent institutional pricing event rather than relying only on qualitative arguments about strategic relevance. In a company like Nebius, where much of the core story is tied to future infrastructure deployment, future utilization, and future revenue ramp, a marked private asset of this size offers something unusually useful: a visible anchor.
What Nebius’ stake may be worth
Nebius retains roughly 28% of ClickHouse. At a $15B valuation, that implies a stake worth about $4.2B.
That is large enough to materially shape how the broader Nebius equity should be thought about. This is not a side investment tucked into the footnotes. It is a major asset sitting alongside a capital-intensive neocloud buildout. A useful way to think about it is that Nebius contains both an infrastructure operating story and a valuable software ownership story. The market may still spend most of its time staring at the former, though the latter is too large to treat as background decoration.
A simple stake sensitivity table makes the point clearer:
Choose your Northwise access
Keep reading with the path that fits you
Join Northwise Premium
Unlock valuation outputs, downloadable models, portfolios, action frameworks, and complete Premium research.
Join Northwise PremiumReader discussion
Discuss the research
Premium access is required to join this report's discussion.
Join Northwise Premium

No comments yet. Start a thoughtful discussion.