Northwise
Company AnalysisPremiumApril 5, 2026

Nebius Toloka Analysis

By Northwise Research TeamNebius Group N.V.
Nebius Toloka Featured Image

How Nebius Toloka provides a human verification, alignment, and reliability layer that most AI infrastructure peers do not control.


1. The role of Nebius Toloka

Nebius is often framed through the loudest part of its story. Power contracts, GPU clusters, hyperscaler agreements, and the visible expansion of a global AI cloud naturally draw the eye first. That is understandable. Nebius has been building around a capital-rich balance sheet, deep Yandex engineering roots, and a rapidly expanding infrastructure footprint designed to serve the next wave of AI demand.

By late 2025, management was guiding toward roughly 100 MW of active power by year end, with 800 MW to 1 GW targeted by 2026 and more than 2.5 GW of contracted power behind that buildout. Year end-2026 ARR guidance reached $7B to $9B.

That surface-level picture is real, but it is incomplete. A factory is only as useful as the quality of what moves through it. In AI, the value chain does not end at compute. It bends upward into data curation, expert feedback, safety testing, evaluation, and increasingly into the reliability layer that determines whether a model can be trusted in production. That is where Toloka enters the frame.

Toloka should be understood as a strategic extension of Nebius rather than a side asset carried over from corporate history. Nebius provides the industrial base of the stack. Toloka provides part of the judgment layer. If Nebius is building the power plant and the assembly line, Toloka helps decide whether the output is safe enough, accurate enough, and reliable enough to actually ship. That distinction becomes more important as AI systems move from internal experimentation into regulated enterprise workflows and agentic systems where a wrong answer carries real cost.

Nebius Toloka The Agentic AI Platform Stack Northwise

The market has become comfortable valuing compute in visible units such as megawatts, utilization, contracted backlog, and revenue ramp. It is far less comfortable valuing the layers that make compute economically durable.

Nebius is already building beyond raw infrastructure economics through internal tooling, workflow integration, compliance positioning, and adjacent software assets that can widen margins and deepen customer relationships. Toloka fits directly into that logic. It is one of the clearest examples of Nebius trying to own more of the AI stack than the market may currently credit.

1.1 Why Toloka matters to the Nebius ecosystem

Toloka matters inside Nebius because it occupies a part of the AI value chain that is becoming more scarce as model capability rises. The first phase of the AI buildout rewarded access to GPUs, power, and capital. The next phase increasingly rewards access to high-quality human judgment at scale. Frontier models need more than training clusters. They need supervised fine-tuning data, RLHF pipelines, model evaluation, red teaming, multilingual ranking, domain experts, and structured feedback loops that keep improving outputs after pretraining is done. That is Toloka’s ground.

The strategic logic becomes clearer when Nebius is viewed as a full-stack AI platform rather than a cloud landlord. Nebius already operates a vertically integrated model that includes infrastructure, orchestration, compliance tooling, developer interfaces, and inference-facing products such as AI Studio, which had more than 60,000 registered users by the end of Q1 2025. Nebius Cloud 3.0, or Aether, was also positioned for regulated enterprise AI through certifications and compliance alignment including SOC 2 Type II, HIPAA, ISO 27001, NIS2, and DORA. In other words, Nebius has already been building toward customers that care about governance and production readiness, not just access to compute.

That is exactly the customer set for which Toloka becomes more valuable.

A regulated enterprise asks whether the model can be tuned, verified, monitored, audited, and trusted. Toloka’s capabilities in expert data generation, human feedback, evaluation, and red teaming sit much closer to those questions than to the older, simpler task of generic labeling. This is why the relationship is strategically coherent. Nebius is building the roads, the power grid, and the transport network. Toloka helps determine whether the cargo is safe enough to ship.

The relationship is also no longer theoretical. In February 2026, Nebius and Toloka announced plans to integrate Tendem into the Nebius ecosystem to bring human experts on demand directly into AI agent workflows. Nebius framed the combined offering around three layers: intelligence through Token Factory, autonomy through Tavily’s agentic search, and reliability through Tendem’s human verification layer. It shows Nebius itself describing Toloka as part of the product architecture rather than treating it as a passive equity holding.

There is also a portfolio logic here. Nebius has internal assets that introduce optionality beyond raw infrastructure economics, and Toloka fits that pattern unusually well. A pure compute business is powerful during shortage periods, but margin structures can compress over time as supply expands and infrastructure becomes more competitive.

A data and reliability business behaves differently. It tends to sit closer to customer workflow, closer to model outcomes, and closer to the part of the stack where switching costs can become sticky. That means Toloka does not just add another line item in theory. It potentially changes the quality of Nebius revenue over time. Compute gets a customer in the door. Reliability and expert verification can keep that customer in the building.

Layer

Nebius role

Toloka role

Strategic effect

Infrastructure

Power, data centers, GPU clusters

None directly

Enables model training and inference at scale

Enterprise AI stack

Compliance tooling, orchestration, AI Studio, Aether

Human feedback, evaluation, expert verification

Makes enterprise deployment more credible

Agentic workflows

Managed inference, Token Factory, Tavily search

Tendem human verification and expert escalation

Improves trust in higher-stakes automation

Margin optionality

Software and platform integration

Higher-value data and evaluation services

Reduces dependence on raw compute economics alone

1.2 The strategic case for data, alignment, and reliability

The simplest way to frame Toloka’s importance is this: compute makes models possible, but data and alignment determine whether they become useful in the real world.

That sounds obvious, yet markets still tend to price the first layer more aggressively than the second. Part of the reason is visibility. A 1 GW buildout is easy to visualize. A 10,000-person expert network improving model reliability across dozens of domains is harder to fit neatly into a spreadsheet. However, the second layer can become the bottleneck once the first begins to scale.

Toloka’s strategic relevance comes from operating across three increasingly important parts of that bottleneck.

First, there is fine-tuning and expert data generation. As models move into specialized domains such as law, science, medicine, finance, and software engineering, generic internet data becomes less sufficient. What matters more is curated expert-generated data, domain-correct reasoning, and subtle judgment calls that cannot be scraped from the public web without quality degrading. Toloka’s move from mass-market crowdsourcing toward an expert network of more than 10,000 verified specialists across 20+ domains and 40+ languages is central here. That shift changes the business from labor supply to judgment infrastructure.

Second, there is alignment. RLHF, DPO, preference ranking, and post-training feedback systems are now core to how major models become commercially usable. This is part of the assembly process. A powerful model without alignment is a race car with loose steering. Fast, impressive, and one bad turn away from the wall. Toloka’s role in human feedback and model preference ranking places it directly in that stage of the lifecycle.

Third, there is reliability in production. Tendem is built as a hybrid human-AI system where verified experts can be called directly inside agent workflows through MCP. The practical implication is straightforward: when an agent hits uncertainty, ambiguity, or a low-confidence edge case, it can escalate to a vetted human expert without the workflow breaking apart into manual project management.

Nebius Toloka Post Training Human Feedback Layer Northwise

Nebius and Toloka now plan to integrate that capability directly into the Nebius ecosystem, turning human escalation into part of the operating logic of the stack rather than an external service bolted on later.

Nebius is explicitly trying to combine managed inference, agentic search, and human verification in one environment. The commercial logic is clear. Enterprise customers want AI systems that can act, retrieve, reason, and still know when to ask for help. In that sense, Tendem addresses the reliability ceiling that often appears when an impressive demo meets a messy production setting.

The performance claims attached to Tendem are also notable. The hybrid workflow is described as 53% faster than a human-only baseline, with quality improving by 21.3% and task completeness rising by 22.3 percentage points, while also generating audit-ready decision trails. Those are not cosmetic improvements. They speak directly to whether agentic systems can move from curiosity to economic utility.

Nebius is already building for regulated enterprise AI and for a broader AI factory model. Toloka gives Nebius a way to reach into the layer above compute and offer something closer to operational trust. In a world where AI agents begin touching procurement, legal review, research workflows, customer operations, financial processes, or healthcare support, that trust layer may become the toll bridge rather than the highway. Many companies can rent compute. Far fewer can help make the outcome dependable.

This is also where the broader strategic pattern comes into view. The most durable winners in infrastructure tend to control choke points rather than simply participating in throughput. With Nebius, the visible choke point today is power and deployment. The less visible one may increasingly be data quality, alignment, and production reliability. Toloka gives Nebius exposure to that second choke point.

AI lifecycle stage

What is required

Toloka contribution

Why it matters to Nebius

Post-training improvement

Domain-specific supervised fine-tuning data

Expert-generated prompts, dialogues, and reasoning data

Increases the value of Nebius-hosted training and fine-tuning workflows

Alignment

RLHF, DPO, preference ranking

Human rankings across quality, safety, and helpfulness

Supports more commercially viable models

Evaluation

Benchmarking, truthfulness, naturalness, multilingual quality

Advanced human evaluation frameworks

Helps enterprise customers trust deployment outcomes

Safety and governance

Red teaming, edge-case testing, auditability

Structured adversarial testing and human review

Aligns with Aether and Nebius’ regulated enterprise positioning

Agentic production

Escalation paths for uncertain tasks

Tendem human-in-the-loop workflow via MCP

Makes AI agents more usable in real environments

This is the architectural opening to the Toloka story. It places the asset where it belongs, inside the operating logic of Nebius rather than floating beside it as an inherited side business.

Table of Contents

  1. The role of Toloka inside Nebius
    1.1 Why Toloka matters to the Nebius ecosystem
    1.2 The strategic case for data, alignment, and reliability
  2. From Yandex asset to Nebius pillar
    2.1 Toloka’s origin as a human-in-the-loop platform
    2.2 The restructuring that separated Toloka from Russian assets
    2.3 Why independence changed Toloka’s strategic relevance
  3. Toloka’s shift from crowdsourcing to frontier AI
    3.1 The move from mass annotation to expert networks
    3.2 The pivot toward fine-tuning, RLHF, and evaluation
    3.3 Why this matters as AI systems become more agentic
  4. Leadership, governance, and strategic control
    4.1 Olga Megorskaya and the operating direction of Toloka
    4.2 The Bezos round and the reset of Toloka’s governance
    4.3 Why Nebius reduced control but kept economic upside
  5. What Toloka actually does today
    5.1 Fine-tuning data for large language models
    5.2 Human feedback, alignment, and preference optimization
    5.3 Model evaluation, benchmarking, and red teaming
  6. Tendem and the reliability layer for AI agents
    6.1 What Tendem is and why it matters
    6.2 How MCP brings human judgment into agent workflows
    6.3 Where hybrid workflows outperform AI-only systems
  7. The integration of Toloka within the Nebius AI Factory
    7.1 Compute, autonomy, and reliability in one stack
    7.2 How Toloka strengthens the Nebius enterprise offering
    7.3 Why this creates differentiation beyond raw GPU access
  8. Customers, partnerships, and competitive positioning
    8.1 Toloka’s client base and market validation
    8.2 The strategic value of the expert network
    8.3 How Toloka compares with Scale AI and adjacent peers
  9. Financial profile and valuation context
    9.1 What the Bezos round signaled
    9.2 The difficulty of valuing Toloka from the outside
    9.3 The range of plausible benchmark frameworks
  10. Why Toloka matters to the Nebius equity story
    10.1 The hidden asset case inside NBIS
    10.2 Why Toloka changes the shape of Nebius upside
    10.3 How investors should think about strategic relevance versus reported financials
  11. Valuing Toloka within the Nebius sum of the parts
    11.1 Minority stake math versus strategic control value
    11.2 Private market comps and the problem of clean comparables
    11.3 What Toloka could be worth across conservative, base, and upside cases
  12. Future optionality and paths to value realization
    12.1 IPO, spinout, or further private funding
    12.2 Capital recycling and balance sheet flexibility for Nebius
    12.3 How Toloka could help fund or de-risk the broader AI infrastructure buildout
  13. Northwise view: why Toloka may be more important than the market assumes
    13.1 The shift from compute scarcity to data and reliability scarcity
    13.2 Why Toloka gives Nebius a more durable position in the AI stack
    13.3 What would need to happen for the market to fully recognize that value

2. From Yandex asset to Nebius pillar

Toloka did not begin as a polished standalone AI infrastructure asset with Western governance, blue-chip customers, and a clear role inside a new AI stack. It began inside Yandex, built for a more practical need. Machine learning systems required labeled data, search relevance needed human judgment, and algorithms needed a way to compare what they predicted against how people actually interpreted the world. Toloka was created to solve that problem at scale.

Many companies are now trying to build expert data and model evaluation capabilities in response to the generative AI boom. Toloka arrived through a different path. It spent years building the plumbing before the market fully appreciated what that plumbing would eventually be worth. In that sense, Toloka resembles an old rail line that suddenly becomes strategic once the city grows around it. The tracks were already there. The value changed when the surrounding system changed.

The company’s later relevance inside Nebius follows directly from that history. Nebius did not acquire a random annotation platform and try to force a narrative around it. It inherited a business that had already spent roughly a decade building marketplace mechanics, quality control systems, and global human-in-the-loop infrastructure. Once AI shifted from simple prediction toward reasoning, alignment, and autonomous execution, those earlier capabilities became more valuable, not less.

2.1 Toloka’s origin as a human-in-the-loop platform

Toloka was founded in 2014 by Olga Megorskaya as a crowdsourcing and human-in-the-loop platform inside the Yandex ecosystem. The original use case was simple. Search engines, recommendation systems, computer vision models, and early natural language systems all needed ground-truth data. Machines could sort and infer at speed, but they still needed humans to judge whether an image was labeled correctly, whether a search result was relevant, or whether a language output matched the intent of a prompt.

At that stage, the business was built around volume and workflow efficiency. Toloka developed a global pool of freelance contributors, often referred to as Tolokers, who could complete microtasks across large distributed datasets. The model looked simple on the surface, but the real intellectual property sat beneath it. Toloka had to solve for task routing, worker quality, redundancy, disagreement resolution, fraud control, throughput management, and statistical confidence. A mass labor pool without those systems is noise. With those systems, it becomes structured human judgment.

Toloka’s original edge was never just labor access. Labor is easy to describe and hard to defend. The harder thing to build was a system that could consistently transform fragmented human input into reliable training data at industrial scale. That required years of iteration in incentives, task design, and mathematical quality control. By the time generative AI began demanding richer forms of post-training data, Toloka already had much of the operating muscle needed to serve that demand.

Over time, the nature of the work evolved. Early machine learning pipelines often needed binary judgments or narrow annotations. Generative AI required more complex human contributions. Large language models needed multi-turn dialogue generation, ranking tasks, reasoning validation, multilingual review, and nuanced comparisons across outputs. As the market moved in that direction, Toloka shifted with it.

The company’s transition toward a curated expert network was therefore less of a reinvention than it may appear from the outside. The core business remained the same in one important sense: converting human judgment into machine-usable data. The difference was that the market had moved from needing many hands to needing more specialized minds.

Period

Primary use case

Workforce model

Strategic significance

2014 to early 2020s

Search relevance, image labeling, ML annotation

Global mass crowdsourcing

Built the marketplace and quality-control engine

Transition period

More complex NLP and model training workflows

Hybrid workforce

Expanded beyond simple microtasks

Generative AI era

SFT, RLHF, evaluation, safety

Curated expert network

Shifted from annotation supply to judgment infrastructure

2.2 The restructuring that separated Toloka from Russian assets

Toloka’s strategic identity changed sharply after the geopolitical shock of 2022. Following Russia’s invasion of Ukraine, Yandex faced mounting pressure to separate its international and domestic assets. That pressure was commercial, legal, political, and reputational all at once. Businesses with global customers, Western counterparties, and long-duration growth ambitions could not remain comfortably attached to a Russian-controlled parent structure. The corporate architecture had become a constraint.

That process culminated in the broader Yandex restructuring completed in 2024. Russian assets were sold to a domestic consortium in a deal valued at roughly $5.4B. The international businesses that remained outside that perimeter were reorganized into what became Nebius Group, headquartered in Amsterdam and positioned as a Western-facing AI infrastructure company. Toloka was retained as one of the major pillars of that new structure, alongside the AI Cloud business, Avride, and TripleTen.

This was more than a paperwork exercise. For Toloka, the separation changed who could invest, who could buy, who could partner, and how credible the company could be in sensitive enterprise settings. Before the split, Toloka carried a structural overhang. Even if the underlying technology was strong, the governance wrapper made it harder to attract top-tier Western capital and harder to become a trusted vendor to companies building frontier AI systems. The business was effectively trying to run a world-class race with a parachute attached.

Once separated, that drag started to lift.

The timing also mattered. The restructuring occurred just as the AI market was beginning to assign much higher value to data infrastructure, model evaluation, and post-training systems. A few years earlier, Toloka might have emerged as a niche software-services story. By 2024 and 2025, it was re-entering the market during a phase when companies like Scale AI were commanding massive valuations and when frontier labs were spending aggressively on alignment, safety, and expert-generated data.

That combination gave Toloka a very different launch window than the one it would have had under the old structure.

Nebius Toloka Evolution Timeline Northwise

Restructuring element

What changed

Why it mattered for Toloka

2024 Yandex split

Russian assets sold for about $5.4B

Cleared the path for Western-facing governance

New parent structure

Nebius Group formed in Amsterdam

Repositioned Toloka within a European AI infrastructure group

Counterparty perception

Reduced geopolitical and sanctions overhang

Improved enterprise trust and capital access

Strategic timing

Separation aligned with the GenAI data boom

Allowed Toloka to re-emerge into a hotter market category

2.3 Why independence changed Toloka’s strategic relevance

Independence changed Toloka in two ways at once. It improved the business on its own terms, and it improved the way the market could value that business.

On its own terms, the company gained room to operate with greater strategic autonomy. That includes capital raising, board formation, customer acquisition, hiring, and product positioning. The clearest public signal arrived in May 2025, when Jeff Bezos’s investment firm led a $72M funding round into Toloka.

That round did more than provide cash. It served as institutional validation that Toloka had become investable to elite Western capital after the separation. Mikhail Parakhin, Shopify’s CTO and a former senior Microsoft executive, also became Executive Chairman of Toloka’s board. That is the kind of governance signal that materially changes the room a company can enter.

Nebius also made a deliberate governance trade. To attract outside investors and support Toloka’s independence, it gave up majority voting control while maintaining a significant majority economic stake. That structure is strategically elegant. Toloka can operate more like a venture-backed AI company, while Nebius shareholders still retain much of the financial upside. In effect, Nebius loosened its grip on the steering wheel in order to keep a larger claim on the destination.

That shift affects operations, but it matters just as much for valuation. Markets routinely discount conglomerate assets when ownership, governance, and monetization pathways are muddy. A semi-independent Toloka with external investors, an independent board, cleaner governance, and a plausible long-term IPO path is easier to benchmark against AI data peers. It becomes more legible. And legibility is often the first step toward value recognition.

This is also where Toloka stops looking like a legacy inheritance and starts looking like a pillar. Under the old framing, it could have been dismissed as a leftover business from a prior corporate era. Under the new framing, it looks more like a strategic node within Nebius’s broader AI factory architecture. Nebius provides compute, cloud orchestration, and increasingly broader agentic infrastructure. Toloka provides the human verification, expert judgment, and post-training systems that make those layers more useful in production. The fit is tighter after independence than before, partly because the market category itself became more coherent.

The February 2026 product integration announcement pushed that shift further. Nebius and Toloka described a plan to integrate Tendem into the Nebius ecosystem so that AI agents can call human experts on demand inside live workflows. That announcement shows independence did not mean separation in the strategic sense. Toloka became more autonomous as a company while becoming more operationally relevant to Nebius as a platform. That is a powerful combination. One side increases financial optionality. The other side increases product-level importance.

Independence lever

Result

Strategic impact

$72M funding round in May 2025

External capital and validation

Confirmed Toloka could attract top-tier Western investors

New board structure

Mikhail Parakhin as Executive Chairman

Strengthened credibility with U.S. and enterprise customers

Voting versus economic control split

Nebius gave up majority voting control but kept significant majority economics

Improved investability without surrendering long-term upside

Product integration with Nebius

Tendem planned inside the Nebius ecosystem

Reinforced Toloka as an operating pillar, not just an owned asset

Before the split, Toloka was strategically useful but structurally constrained. After the split, it became easier to fund, easier to govern, easier to trust, and easier to integrate into a Western AI platform with real ambitions. That is a meaningful transition. Assets do not become pillars just because management says so. They become pillars when the surrounding system begins to depend on them. Toloka is moving into that category.

3. Toloka’s shift from crowdsourcing to frontier AI

Toloka’s evolution is one of the more important parts of the Nebius story, partly due to how easy it is to misread from the outside. A legacy crowdsourcing platform can sound operationally useful but strategically ordinary. That framing no longer fits. The business Toloka is building today sits much closer to the post-training, evaluation, and reliability layers that increasingly determine whether advanced AI systems can move from technical possibility to commercial usefulness.

Earlier machine learning cycles rewarded scale in raw labeling throughput. The core challenge was often gathering enough human judgments to classify images, rank search results, moderate content, or support early natural language tasks. Those workflows were large, repetitive, and operationally demanding, but they were still relatively narrow in what they asked from the human worker.

Frontier AI raised the bar. Large language models, multimodal systems, and now agentic workflows require something more expensive and more difficult to organize. They require domain knowledge, contextual judgment, chain-of-thought assessment, safety review, multilingual reasoning, and the ability to compare outputs that may both sound plausible while only one is actually correct. That is a different labor model, a different quality-control problem, and ultimately a different business.

Toloka has been moving along that curve for years. Human annotation has been reworked into a platform for structured expert judgment. The shift looks subtle in a headline and substantial in economic terms. A company supplying generic labels participates in the AI supply chain. A company helping frontier models learn, align, and self-correct begins to shape the quality of the final product.

3.1 The move from mass annotation to expert networks

Toloka’s first operating model was built around scale. It developed a broad global workforce of contributors who could complete millions of discrete tasks across search, computer vision, recommendation systems, and language workflows. This model was well suited to the earlier AI era, where value often came from taking a very large dataset and making it machine-readable through repeated human judgments.

That approach gave Toloka a durable operational base. It learned how to break complex tasks into manageable units, route work efficiently, monitor output quality, suppress fraud, and aggregate fragmented human responses into datasets that could actually train models. Those systems matter more than the simple phrase "crowdsourcing platform" suggests. Large distributed labor pools are messy by default. Turning them into reliable data engines requires statistical control, feedback systems, task design, and quality scoring that most newer entrants still need years to build.

But the nature of demand changed. As AI models became more capable, the question shifted from whether a task could be labeled to whether it could be judged correctly in context. A model answering a physics question, reviewing a legal clause, writing software, or ranking multiple reasoning paths does not benefit much from generic low-cost annotation. It benefits from people who actually know what they are looking at.

The company now operates with more than 10,000 verified experts across 90+ domains and 40+ languages. That is a meaningful change in workforce composition and strategic positioning. The economic value is no longer tied mainly to labor volume. It is tied to the ability to source scarce judgment, verify expertise, and deploy that expertise in structured workflows.

Nebius Toloka Domain and Expertise Northwise

This shift also raises the ceiling of what Toloka can help customers do. A mass annotation platform helps improve datasets. An expert network can help shape model behavior itself. That difference is the gap between stocking a warehouse and calibrating the instruments in a control room. Both matter, though only one sits close to the final decision-making layer.

Workforce model

Earlier Toloka model

Current Toloka direction

Strategic implication

Labor base

Global mass crowd

10,000+ verified experts

Higher-value judgment supply

Task type

Microtasks and basic labeling

Domain-specific reasoning, review, ranking

Moves closer to model behavior shaping

Core value

Throughput and scale

Expertise, trust, contextual accuracy

More defensible in frontier AI workflows

Economic position

Data-processing utility

Judgment infrastructure provider

Better fit for premium AI budgets

3.2 The pivot toward fine-tuning, RLHF, and evaluation

The second major shift was in where Toloka placed itself inside the AI lifecycle. Rather than remaining centered on generic training data, it moved toward three of the highest-value layers in modern model development: supervised fine-tuning, alignment through RLHF and related methods, and advanced evaluation.

Supervised fine-tuning is where a base model begins to absorb more structured behavior. At this stage, the goal is not simply to expose the model to more tokens. The goal is to teach it how better answers look in context. That requires curated prompts, ideal responses, multi-turn dialogues, and increasingly nuanced reasoning examples in fields where correctness is not obvious from surface fluency alone. Toloka’s expert network is directly relevant here. A physicist, lawyer, or software engineer can produce training data with a very different quality profile than a generic annotator.

RLHF and related methods such as Direct Preference Optimization push further. Here the model is no longer just learning from examples. It is learning from comparative judgments about output quality. Human evaluators rank responses across dimensions such as helpfulness, truthfulness, safety, reasoning quality, and task completion. Those rankings then feed into reward models or preference-based optimization methods that shape how the final system behaves. In simple terms, a model begins to learn not only what to say, but how to behave when several plausible answers are available.

That is a crucial layer in the current AI stack. Pretraining gives scale. Post-training gives personality, safety, usability, and commercial readiness. Toloka’s relevance rises due to where that work sits. The closer a company moves to post-training and evaluation, the closer it moves to the part of the model lifecycle that decides whether the system is tolerable, useful, and trustworthy in production.

Evaluation may be the least flashy and most underrated part of the group. As model outputs become longer, more agentic, and more context-dependent, measuring quality gets harder. Old benchmark logic begins to break down. A model can sound natural while being wrong. It can complete part of a workflow while skipping the step that actually mattered. It can answer safely in one context and fail on the edge cases that a real enterprise environment will eventually produce. Toloka’s work in evaluation, including preference prediction and multidimensional judgment frameworks, pushes directly into this problem.

The business starts to resemble quality assurance for an industrial process. The model may be powerful, the interface may be elegant, and the inference cost may be falling, though none of that guarantees dependable output. Fine-tuning improves the machinery. RLHF adjusts the steering. Evaluation tells you whether the whole system still drifts when the road curves.

AI development layer

What it does

Toloka role

Why it is strategically important

Supervised fine-tuning

Teaches better task-specific behavior

Expert-generated prompts, answers, dialogues

Improves domain quality and model usefulness

RLHF / DPO

Aligns outputs with human preferences

Response ranking, preference collection, reward-model support

Shapes model behavior and commercial viability

Evaluation

Measures truthfulness, safety, reasoning, completeness

Human scoring, benchmarking, preference prediction

Helps determine whether models are ready for deployment

Red teaming

Stress-tests failures and edge cases

Adversarial prompts and human review

Supports safety, auditability, and enterprise confidence

3.3 Why this matters as AI systems become more agentic

This transition matters even more in the agentic phase of AI. A chatbot that produces a weak answer is annoying. An agent that takes action on weak judgment is a different problem entirely. Once systems begin retrieving information, executing multi-step tasks, making recommendations, and interacting with external tools, the value of expert oversight rises quickly.

That is the context in which Toloka’s newer direction becomes strategically sharper. Agentic systems create a larger surface area for failure. They do not merely respond. They interpret, sequence, decide, escalate, and sometimes act across environments that contain ambiguity, missing context, and real-world consequences. As that complexity rises, the demand for reliable post-training data, rigorous evaluation, and human escalation pathways rises with it.

This is why Toloka’s evolution from annotation platform to judgment infrastructure is so relevant to Nebius. Nebius is building toward a broader AI factory architecture that increasingly includes inference, enterprise orchestration, agentic tooling, and now human verification through Tendem. A more agentic AI system places more value on the ability to detect uncertainty, defer intelligently, and route edge cases to verified experts. That makes Toloka’s capabilities more important in the next phase of the stack than they were in the last one.

There is also a broader economic implication here. As raw compute scales, some parts of AI infrastructure may drift toward greater competition and lower differentiation. The reliability layer may not. High-quality expert data, preference judgment, and human-in-the-loop escalation are harder to commoditize quickly due to the bottleneck is not only capital, but organization. You need talent supply, workflow design, trust systems, quality controls, domain coverage, and feedback loops that improve with use. Those assets take time to assemble.

Toloka begins to look less like a service layer and more like an enabling system. In an agentic world, the question is no longer whether the model can generate language at low cost. The question is whether the broader system can act with enough accuracy, restraint, and recoverability to be trusted in production. Toloka helps address that question from several angles at once.

The useful metaphor here is aviation. Compute gives the aircraft thrust. Foundation models provide lift. Agent frameworks supply navigation. Toloka sits closer to the instruments, the warning systems, and the trained humans in the loop when weather conditions change. Nobody notices those components on a clear day. Once the turbulence begins, they become central.

AI phase

Main constraint

Why Toloka becomes more valuable

Early ML

Data volume

Basic annotation and structured labeling support

Generative AI

Post-training quality

Fine-tuning, RLHF, evaluation, red teaming

Agentic AI

Reliability under action and uncertainty

Expert escalation, human verification, decision quality in live workflows

This is the core of the transition. Toloka followed the market upward, from labeling what models see to helping determine how they behave. That shift changes the strategic category it belongs to, the customers it can serve, and the role it may ultimately play inside the Nebius platform.

Northwise Research

Get Northwise Research in Your Inbox

New research, model releases, portfolio updates, and major Northwise developments delivered directly to your inbox.

4. Leadership, governance, and strategic control

Toloka’s strategic value does not rest only on what it does. It also rests on who is steering it, how it is governed, and what kind of ownership structure sits behind the company as it tries to serve frontier AI customers in the U.S. and Europe. In businesses tied to model safety, enterprise trust, and expert judgment, governance is not decorative. It is part of the product. Customers need confidence that the company is stable, independent enough to scale, and structured in a way that supports long-duration relationships.

That is especially true for Toloka due to its path out of Yandex and into the Nebius ecosystem. Once the business moved from being a useful internal platform to becoming a potentially strategic AI infrastructure asset, its governance model had to mature along with its market role. The result is a structure that blends founder continuity, outside institutional validation, and a deliberate compromise between operational independence and shareholder economics.

4.1 Olga Megorskaya and the operating direction of Toloka

Toloka’s operating direction still begins with Olga Megorskaya. That continuity is significant. In many restructurings, the underlying business survives while the operating soul gets diluted through executive turnover, consultant language, and board-level abstraction. Toloka appears to have avoided that trap. Megorskaya has been part of the company’s evolution since its founding in 2014, which means she has seen the platform across multiple AI eras, from early search relevance and labeling workflows to post-training, RLHF, and now agentic reliability systems.

That kind of continuity is strategically useful in a business like this. Toloka is not a clean-sheet startup chasing a hot category. It is a company that spent years building internal systems for quality control, labor orchestration, and human judgment aggregation before the market fully rewarded those capabilities. Founders who have lived through that full transition tend to understand where the bones are buried. They know what parts of the system are robust, what parts were built for a prior generation, and what needs to be reworked as customer requirements change.

Megorskaya’s background also fits the nature of the product. Toloka sits at the intersection of machine learning operations, human computation, and data quality science. This is a strange corner of the market. It is operationally heavy, technically subtle, and full of failure modes that are easy to miss from a distance. The value of the company is not just in having a network of workers or experts. It is in translating fragmented human input into something that can reliably improve model behavior. That takes marketplace discipline, research literacy, and a deep understanding of how quality degrades when incentives, task design, or reviewer calibration drift.

Under her leadership, Toloka has moved from a broad crowdsourcing platform toward a more specialized frontier AI data and evaluation company. That shift was not cosmetic. It required repositioning the labor model, raising the expertise bar, focusing on higher-value workflows, and aligning the company more tightly with the economics of the generative AI era. The move toward a verified expert network of more than 10,000 specialists across 20+ domains and 40+ languages reflects a business that understands where future demand is likely to sit.

A weaker operator might have tried to maximize short-term volume by leaning on generic labeling demand. Megorskaya appears to have pushed the company toward the harder and more durable category, where expert judgment, evaluation quality, and reliability systems command better strategic value. That is not the easy road. It is closer to rebuilding the ship while already at sea.

4.2 The Bezos round and the reset of Toloka’s governance

The May 2025 funding round led by Bezos Expeditions was a turning point for Toloka. On the surface, it brought in $72M of capital. Underneath that headline, it did something more important. It re-set the institutional perception of the company.

Capital from a high-profile U.S. investor carries signaling value in almost any case, though here the significance was sharper. Toloka had emerged from a corporate structure that previously carried geopolitical complexity and reputational drag. For a company trying to serve frontier labs, enterprise buyers, and Western capital markets, that overhang had to be decisively addressed. The Bezos round helped do that. It functioned as external validation that Toloka could stand on its own as a credible Western-facing AI infrastructure company.

The governance implications ran deeper than the capital itself. As part of that shift, Mikhail Parakhin became Executive Chairman of Toloka’s board. Parakhin is not simply a recognizable executive. His background at Microsoft and Shopify links Toloka more directly to the North American enterprise and model ecosystem. It tells investors and customers that the company is building a board capable of interfacing with the commercial and technical demands of the current AI market, not merely preserving continuity from an older corporate structure.

This is the sort of governance reset that changes what rooms a company can credibly enter. Enterprise customers evaluating vendors for model alignment, expert data, or agentic reliability do not only ask whether the product works. They ask whether the company behind the product will still be around, whether it can meet compliance expectations, whether its board looks serious, and whether its ownership structure is clean enough to support long-term contracts. High-stakes AI vendors are judged partly like software companies and partly like infrastructure counterparties. Toloka needed governance that matched that reality.

The round also sharpened the company’s future strategic pathways. Once a company attracts outside capital of that profile and begins to operate with an independent board structure, it becomes easier to imagine additional private rounds, a clearer valuation benchmark, or eventually some form of public market event. None of that guarantees an IPO, though it does improve the probability that Toloka will eventually be valued in a way that is easier for outsiders to understand.

In that sense, the Bezos round did not simply strengthen the balance sheet. It gave Toloka a new passport.

4.3 Why Nebius reduced control but kept economic upside

Nebius’s decision to reduce voting control while retaining a significant majority economic stake was one of the more interesting moves in the Toloka story. It looks counterintuitive at first glance. Why loosen control over an asset that may become more important over time?

The answer is that tighter control is not always the best way to maximize value. Sometimes it depresses it.

For Toloka to attract top-tier outside investors, deepen its governance credibility, and eventually stand on its own as a higher-value AI asset, it needed more visible independence. A business that appears too tightly held by a parent can struggle to attract fresh capital on favorable terms, struggle to build a clean peer set, and struggle to convince the market that minority shareholders or outside stakeholders will matter over time. That can limit both valuation and strategic flexibility.

Nebius appears to have recognized that. By giving up majority voting control, it made Toloka more investable. By keeping significant majority economic ownership, it preserved most of the upside if Toloka compounds in value from here. This is a trade many companies talk about and fewer execute well. It requires management to accept less direct control in exchange for a stronger long-term value creation path.

There is a second layer to the logic. Toloka’s value to Nebius now runs through two channels at once.

The first is operational. Toloka enhances Nebius’s full-stack AI platform through expert data, evaluation, and now the planned Tendem integration for human-on-demand agent workflows. That gives Nebius product-level differentiation.

The second is financial. If Toloka grows into a larger standalone business, raises more capital at higher valuations, or eventually reaches a liquidity event, Nebius benefits as a major economic owner. That creates optionality around capital recycling, balance sheet support, or simple mark-to-market value recognition over time.

This structure lets Toloka behave more like a venture-backed frontier AI company while still allowing Nebius shareholders to participate in much of the upside. It is a bit like owning a controlling economic interest in a satellite that has been allowed to find its own orbit. The parent loses some ability to micromanage the flight path, though it gains a better chance of seeing the asset travel farther.

When an asset is buried inside a larger structure with unclear governance, uncertain monetization pathways, and no obvious route to external validation, investors often apply a discount. When governance is clearer, capital is external, and the asset begins to look independently legible, that discount can narrow. Toloka is moving from the first category toward the second.

For Nebius, this was a disciplined compromise. It did not fully spin Toloka out. It did not keep Toloka tightly boxed inside the parent either. Instead, it chose a middle path that supports strategic relevance, outside validation, and future monetization flexibility. That is usually what intelligent capital structures look like. Less chest thumping, more geometry.

The key point is simple. Nebius did not reduce control due to Toloka mattered less. It reduced control due to Toloka had the potential to matter more if the structure around it improved.

5. What Toloka actually does today

Toloka’s current business sits in the part of the AI stack that begins after the model is pretrained and before it can be trusted in production. Generic labeling is only a small and increasingly dated way to think about the category. Toloka’s present role is closer to post-training infrastructure for modern AI systems.

In practical terms, Toloka helps model developers improve outputs, align behavior, measure quality, and stress test failure modes before those systems are exposed to real users or embedded inside enterprise workflows. That work spans supervised fine-tuning, reinforcement learning from human feedback, preference optimization, evaluation, benchmarking, and red teaming. The common thread across all of it is that Toloka converts human judgment into structured signal that models can learn from or be judged against.

This is a subtle business on the surface and a very consequential one under the hood. A foundation model can absorb trillions of tokens during pretraining and still fail where it matters most. It can reason inconsistently, hallucinate with confidence, miss domain nuance, or break under ambiguous edge cases. Post-training work is where those rough edges are sanded down, where the model is taught what a better answer looks like, and where teams begin to understand how the system behaves when exposed to the mess of the real world.

That is the lane Toloka occupies today.

5.1 Fine-tuning data for large language models

The first major function Toloka serves is supervised fine-tuning. This is the stage where a large language model is trained on curated examples to improve behavior on specific tasks, domains, or response formats. The goal is not just to give the model more data. The goal is to give it better examples.

That distinction is easy to miss and central to why Toloka adds immense value to Nebius. A large pretrained model already has broad exposure to public information. What it often lacks is precise guidance around how a correct answer should be structured, how much reasoning is appropriate, how a domain expert would respond, or how to behave across specialized use cases that matter in production. Fine-tuning data fills that gap.

Toloka supports that process through expert-generated datasets. They are building systems that must operate across law, finance, medicine, software engineering, science, customer operations, and multilingual enterprise environments. Those systems need data produced or reviewed by people who actually understand the subject matter.

The work itself can take several forms. Experts may write ideal responses to prompts, generate complex multi-turn dialogues, compare answers, validate reasoning, or create structured examples that teach the model how to behave in a narrow context. In more advanced settings, this extends into chain-of-thought style reasoning tasks, tool-use scenarios, code review, and domain-specific judgment calls where surface-level fluency is not enough.

Toloka’s business begins to resemble instructional design for machines. The model already has language. What it needs now is tutoring, correction, and examples from people who know where the subtle mistakes live. A physics model does not improve much from prettier prose. It improves from being shown where the logic broke. A coding model does not get better due to it saw more GitHub text. It gets better when someone who can actually read code shows it what good code looks like and why.

There is also a supply-side reason this function has become more important. High-quality public data is becoming less abundant relative to the needs of frontier AI labs. Once the easy web-scale inputs are absorbed, marginal improvements come more from curated proprietary datasets than from scraping one more pile of internet debris. That pushes value toward companies that can coordinate expert judgment at scale. Toloka is positioned directly in that lane.

5.2 Human feedback, alignment, and preference optimization

The second major function is alignment. Models learn how humans prefer them to behave once several plausible answers are on the table. A system may already be capable of answering a question, though that does not mean it will answer in the most helpful, safe, honest, or contextually appropriate way. Alignment is the layer that shapes those behaviors.

Toloka’s work here centers on reinforcement learning from human feedback, or RLHF, along with adjacent methods such as Direct Preference Optimization. In these workflows, human reviewers compare outputs, rank them, and provide judgments across dimensions such as helpfulness, harmlessness, factuality, reasoning quality, and task completion. Those judgments are then used to train reward models or preference-based optimization systems that guide future behavior.

This is one of the most economically important parts of the modern AI pipeline because it influences whether a model becomes commercially usable. A model can have strong raw capability and still be awkward, brittle, unsafe, or unreliable in ways that limit adoption. Alignment is where the behavior is tuned so that the system becomes more tolerable in everyday use and more dependable in higher-stakes settings.

Toloka’s relevance here stems from scale and nuance. That is a real advantage. Preference data is not one-dimensional. What sounds clear, respectful, safe, or helpful can vary across domains, geographies, and application settings. A global AI product cannot simply optimize for one narrow interpretation of good behavior and assume that translates cleanly everywhere else.

There is a deeper point here as well. Preference optimization is often framed like a quality upgrade, though in many settings it functions more like behavioral governance. These workflows influence what the model does when instructions are ambiguous, when trade-offs emerge between completeness and safety, or when the user is asking for something that falls near a boundary. That is part of why alignment vendors can become strategically important. They are helping shape the conduct of the system, not just polishing the wording.

Nebius Toloka from Physical Compute to Human Alignment Northwise

For Nebius, this layer fits neatly with the broader enterprise and compliance direction of the platform. A company trying to serve regulated or governance-sensitive customers gains value from owning exposure to the part of the stack that helps make model behavior more controllable. Compute may power the engine. Human feedback helps teach the machine how to stay in its lane.

5.3 Model evaluation, benchmarking, and red teaming

The third major function is evaluation. Toloka helps customers answer a simple question that becomes painfully difficult as models improve: how do you know whether the system is actually good enough to deploy?

That question gets harder as outputs become longer, more polished, and more agentic. A model can sound coherent while being wrong. It can complete 80% of a workflow while quietly failing on the one step that carried the highest consequence. It can answer safely in obvious cases and fail under pressure when prompts become adversarial, ambiguous, or domain-specific. Standard benchmarks help, though they do not solve the whole problem. Eventually someone has to judge the behavior in context.

Toloka plays directly into that need. The company provides model evaluation and benchmarking services that go beyond simple legacy metrics. Toloka works in preference prediction and multidimensional evaluation, including criteria such as truthfulness, naturalness, safety, and overall output quality. This is a more demanding form of quality control than checking whether a prediction matched a single reference answer. It is closer to structured adjudication.

That makes sense for where AI has gone. Once models operate in open-ended environments, the test is rarely whether they can produce an answer at all. The test is whether they can produce the right answer, in the right format, with the right level of caution, across enough edge cases that a real business can trust them.

Red teaming extends that logic into failure discovery. Before a model is put in front of users, developers need to know how it behaves when pushed into uncomfortable territory. Toloka helps source and structure those tests by generating malicious prompts, adversarial scenarios, and edge-case inputs that probe for toxic output, policy failure, hallucination, hidden bias, or brittle reasoning. This is especially important for enterprise customers operating in regulated sectors or customer-facing environments where one public failure can create legal, reputational, or operational fallout.

This part of the business may be less glamorous than model launches or benchmark headlines, though it often becomes more important once money is actually on the line. Evaluation and red teaming are where labs and enterprises learn whether their system is merely impressive or genuinely usable. It is the difference between a prototype car that looks beautiful under studio lights and a vehicle that has actually passed crash testing.

Toloka’s role today therefore sits across three linked functions. It helps teach models through fine-tuning, shape them through human feedback, and challenge them through evaluation and red teaming. Each function reinforces the others. Fine-tuning improves capability in specific contexts. Alignment improves behavior. Evaluation reveals where confidence is earned and where it is still fiction.

That is why Toloka’s present-day business belongs in the post-training and reliability layer of the AI stack rather than in the older mental bucket of annotation software. The company is operating where frontier models are made more specialized, more governable, and more deployable. As AI systems move further into enterprise operations and agentic workflows, that layer should only become more important.

6. Tendem and the reliability layer for AI agents

Tendem is where Toloka’s relevance becomes easier to see in product form rather than only in abstract strategic language. Earlier sections explain why expert data, alignment, and evaluation matter. Tendem shows what that looks like when those capabilities are pushed directly into live AI workflows.

This is significant due to the AI market now moving beyond static chat interfaces toward systems that retrieve information, use tools, complete tasks, and make real-time decisions inside production environments. At that point, the challenge is no longer just whether a model can generate a plausible response. The challenge becomes whether the broader system can keep operating when uncertainty appears. Most agents look competent right up until the moment the path gets messy. Then the whole illusion can fall apart very quickly.

Tendem is Toloka’s answer to that problem. It is designed as a hybrid human-AI platform that allows verified human experts to step into agent workflows when the model hits ambiguity, low confidence, or a high-stakes edge case. In plain terms, it gives the agent a way to ask for help without collapsing the whole process into a manual support queue. That is a meaningful shift. Instead of forcing companies to choose between full automation and full human review, Tendem creates a middle layer where automation can continue while judgment is injected where it is actually needed.

Nebius Toloka Human Escalation Layer Northwise

That moves Toloka from being a post-training vendor into something closer to live reliability infrastructure. The distinction is important. Post-training improves the model before deployment. Tendem helps stabilize the system during deployment. One shapes the aircraft before takeoff. The other helps it stay on course once it hits turbulence.

6.1 What Tendem is and why it matters

At a basic level, Tendem is a human-on-demand reliability layer for AI agents. It is meant to sit inside workflows where models are already acting with some autonomy but still encounter situations that require interpretation, caution, or domain expertise that the model alone cannot safely provide.

That category should grow in importance as agentic AI expands. A simple chatbot can fail and still leave only mild damage. An agent that is handling research workflows, procurement support, customer operations, legal review, or financial process automation carries a different error profile. The issue is not just whether it answers. The issue is whether it knows when its own answer may be weak, incomplete, or risky.

Tendem allows human experts to be called into the loop at the moment an agent encounters a task that should not be handled blindly. The broader stack is being positioned around three linked layers: intelligence through Token Factory, autonomy through Tavily’s agentic search capabilities, and reliability through Tendem’s human verification layer.

That framing tells you a lot about how Nebius sees the problem. The company is not trying to build agents that simply move faster. It is trying to build agents that can move with some mechanism for judgment when the environment becomes more complex.

This illustrates significant commercial impact. Enterprise customers rarely reject AI due to the model is weak in a demo. They reject it due to the failure path is too ugly once the model reaches production. Tendem addresses that precise issue. It gives builders a way to create escalation logic inside the workflow rather than bolting on manual review afterward.

The concept is simple enough to explain without flattening it. A good agent can execute. A better agent can defer intelligently. Tendem helps make that second category possible.

6.2 How MCP brings human judgment into agent workflows

The architecture behind Tendem becomes more interesting once MCP enters the picture. MCP, or Model Context Protocol, is emerging as a standard way for AI systems to interact with tools, services, and structured external capabilities. In most discussions, that means databases, APIs, search tools, or software actions. Tendem extends that idea into human judgment.

This is the clever part.

Instead of treating human review as an external interruption, MCP allows the workflow to treat expert input more like a callable function inside the system. When an agent hits uncertainty, it can trigger a Tendem call that routes the task to an appropriate human expert. That expert reviews the issue, provides structured output, and sends a machine-readable result back into the workflow so the process can continue.

That sounds technical, though the operational implication is simple. Human judgment becomes part of the workflow architecture rather than a separate managerial process. The agent does not need to stop functioning entirely. It simply escalates the narrow part of the task that requires higher-confidence reasoning.

That has several advantages. It preserves workflow continuity. It keeps human involvement targeted rather than wasteful. It also creates cleaner audit trails, which matters for regulated use cases and enterprise governance. If a company needs to understand where a decision came from, when human review was triggered, and what the resolution was, a structured escalation path is far more useful than an improvised manual intervention.

This is one of the reasons Tendem deserves to be viewed as more than a service wrapper around contractors. The value is not only in having humans available. The value is in embedding them into the logic of the system in a way that is programmable, selective, and scalable.

For Nebius, that integration is strategically coherent. The company is already building cloud, inference, orchestration, and agentic layers. MCP-linked human escalation gives Nebius a way to extend the stack into the reliability problem that many agent builders will eventually run into. Compute lets the agent run. Search lets it gather context. Tendem helps it avoid driving straight into a ditch when the map gets blurry.

6.3 Where hybrid workflows outperform AI-only systems

Hybrid workflows tend to outperform AI-only systems in the places where confidence and correctness diverge. That problem is becoming more common as models become more fluent. A model can sound certain, move quickly, and still miss the exact nuance that mattered. In low-stakes settings, that may be tolerable. In enterprise workflows, it can turn expensive very fast.

The Tendem hybrid workflow is described as 53% faster than a human-only baseline, while improving quality by 21.3% and task completeness by 22.3 percentage points. It also produces audit-ready decision trails, which pure AI systems do not naturally provide and purely manual systems often handle inconsistently. Those are strong numbers, partly due to they suggest the value is not merely philosophical. It appears operational.

The logic behind the improvement is intuitive. Human-only systems are usually too slow and expensive to scale cleanly. AI-only systems are fast but brittle once ambiguity rises. A hybrid system can let the machine handle the obvious path while reserving human attention for the moments that actually justify it. That raises throughput without forcing the company to accept every failure mode of full autonomy.

The best use cases are usually the ones with some mixture of scale, complexity, and consequence. Legal review, healthcare support, compliance operations, enterprise research, customer workflows with edge-case decision trees, and internal knowledge tasks all fit this pattern. These are environments where full manual handling is too costly and full AI autonomy still feels reckless. Hybrid architecture gives companies a narrower bridge to cross.

A small table helps here.

Workflow type

Strength

Limitation

Human-only

High judgment quality

Slow, expensive, hard to scale

AI-only

Fast, cheap, scalable

Weak on ambiguity, edge cases, and accountability

Hybrid human-AI

Better speed-quality balance

Requires good escalation logic and workflow design

The real strategic point is that hybrid systems often become more valuable as AI gets better, not less. That sounds backwards at first. Many people assume stronger models should reduce the need for humans. In some narrow tasks that will happen. In more consequential systems, stronger models often expand the range of actions companies want to automate, which increases the number of situations where selective human oversight becomes useful. More autonomy can create more need for intelligent deferral, not less.

That is why Tendem fits the direction of travel so well. It is built for the awkward middle ground where AI is already useful, though not yet trustworthy enough to operate alone across every condition. That middle ground may end up being one of the largest commercial zones in the agentic market.

For Toloka, Tendem pushes the company further up the stack. For Nebius, it strengthens the case that the broader platform is trying to solve more than compute supply. It is trying to solve the last mile problem that often blocks deployment. And in AI, the last mile has a habit of becoming the hardest mile.

7. The integration of Toloka within the Nebius AI Factory

The Toloka story becomes more strategically interesting once it is viewed inside the Nebius AI Factory rather than beside it. On a standalone basis, Toloka is already a serious post-training, evaluation, and human-in-the-loop asset. Inside Nebius, it becomes part of a broader attempt to assemble a more complete AI production stack, one that reaches from data center capacity and inference infrastructure up through agentic tooling and into the reliability layer where real-world deployment often succeeds or fails.

That framing matters due to the market still tending to separate these functions too cleanly. Compute is often valued in one bucket. Data and evaluation sit in another. Agents are discussed as a separate software category. Reliability is usually treated as a downstream problem someone else will solve later. Nebius appears to be building with a different logic. The company is positioning itself around the idea that enterprise AI adoption will require these layers to work together in a more integrated way.

7.1 Compute, autonomy, and reliability in one stack

Nebius’s recent product framing makes this integration more explicit. The company has described the combined stack around three layers: intelligence through Token Factory, autonomy through Tavily’s agentic search capabilities, and reliability through Tendem’s human verification layer. That is a useful way to think about the architecture due to each layer addresses a different constraint in practical AI deployment.

The first layer is compute and managed intelligence. Nebius has been building aggressively here through its AI cloud footprint, high-density GPU infrastructure, and model-serving capabilities. This is the most visible part of the story and the one the market is most comfortable valuing. Capacity, power, utilization, and revenue all fit into standard infrastructure logic.

The second layer is autonomy. Intelligence on its own does not create much operational leverage unless the system can retrieve information, use tools, move through workflows, and act across tasks with some continuity. Tavily’s place in the stack points toward that next step. It expands the system from answering to doing.

The third layer is reliability. Once an agent can act with greater autonomy, the failure surface expands. The model may retrieve bad information, misread ambiguity, mishandle an edge case, or continue confidently through a task it should have deferred. Reliability is the layer that deals with those moments. It is the point in the stack where the system gains a way to escalate, verify, and recover rather than simply pushing forward with polished uncertainty.

That three-part structure is strategically coherent. Intelligence without autonomy remains underutilized. Autonomy without reliability becomes dangerous in enterprise settings. Reliability without the first two does not have much to stabilize. Nebius is trying to link all three.

The result is a more complete operating model for AI deployment. A customer can train or run models on Nebius infrastructure, build more agentic behavior on top of that stack, and then introduce human-on-demand escalation when the workflow reaches a boundary condition. In other words, the platform is being built to support action with fallback rather than action with blind confidence.

That is a meaningful distinction. Plenty of AI products can impress in clean conditions. Fewer are being designed around what happens when the conditions stop being clean.

7.2 How Toloka strengthens the Nebius enterprise offering

Toloka strengthens the Nebius enterprise offering due to enterprise buyers rarely care only about raw capability. They care about whether the system can be governed, monitored, explained, and trusted once it enters a real operating environment. That is especially true in regulated sectors and in business functions where errors carry legal, financial, or reputational cost.

Nebius has already been building toward those customers. Its broader cloud and platform positioning has included enterprise compliance alignment, infrastructure designed for serious workloads, and software layers meant to support production AI rather than casual experimentation. Toloka deepens that positioning by giving Nebius a way to address the part of the deployment conversation that begins after the demo ends.

That part of the conversation usually sounds less glamorous and much more commercial.

How is the model being improved after deployment?
How are outputs being evaluated across changing use cases?
What happens when the system encounters ambiguity?
How is the workflow audited?
Where does human oversight enter the process?
How does the company reduce the probability of a costly failure without throwing away the economic benefits of automation?

Toloka gives Nebius better answers to those questions.

Its fine-tuning, alignment, evaluation, and red teaming capabilities already help customers improve models before deployment. Tendem extends that value into deployment itself. Instead of offering only infrastructure and inference, Nebius can increasingly position itself as a provider of production AI systems with built-in escalation logic. That is a stronger enterprise proposition than simply saying the cluster is fast and the GPUs are available.

It also helps Nebius move closer to the customer workflow. Infrastructure businesses often risk sitting one layer too far from the actual point of value creation. Customers buy capacity, though their long-term loyalty is often driven by what the platform helps them accomplish. Toloka shifts Nebius upward into the part of the stack where outcomes become more tangible.

When a customer uses expert-generated fine-tuning data, human preference feedback, structured evaluation, and live human escalation inside the same broader ecosystem, the relationship becomes harder to replace with a commodity compute provider.

Raw infrastructure revenue can grow quickly during shortage cycles, though it can also become more exposed to pricing competition over time. Higher-layer workflow value tends to be stickier. It sits closer to internal processes, governance requirements, and the actual task logic the customer is trying to solve. Toloka helps Nebius reach toward that stickier layer.

Nebius Toloka Trust and Judgement Layer Northwise

7.3 Why this creates differentiation beyond raw GPU access

This integration creates differentiation due to raw GPU access becomes less differentiated over time than the market often assumes during the early stages of a compute boom. In shortage conditions, access itself can look like a moat. As more capital enters the category and more infrastructure comes online, the durable advantage often shifts toward orchestration, workflow integration, customer trust, and control over the layers that sit above the hardware.

Toloka gives Nebius exposure to exactly those layers.

A compute provider can rent capacity. A more integrated platform can help improve model quality, shape model behavior, evaluate readiness, and inject human judgment into agent workflows when failure risk rises. Those are not interchangeable offerings. One sells horsepower. The other helps determine whether the vehicle can finish the route without veering off the road.

As buyers become more sophisticated, they are likely to ask harder questions. Who helps us move from proof of concept to production? Who helps us manage the ugly edge cases? Who gives us a framework for trust when the model is operating inside workflows that actually matter? Nebius’s integration with Toloka is aimed directly at those questions.

There is a second source of differentiation here as well. Toloka gives Nebius an internal asset tied to a different economic profile than infrastructure alone. Compute is capital intensive, capacity driven, and exposed to the timing of deployment and utilization. Human-in-the-loop data, alignment, evaluation, and reliability services sit closer to software and specialized services economics. That does not automatically make them easy or high margin, though it does diversify the strategic mix. Nebius is not only trying to scale more megawatts. It is building exposure to the judgment layer that may become more valuable as AI systems get more capable and more consequential.

This is also where the AI Factory concept begins to feel less like branding and more like architecture. A real factory is not defined only by how much machinery it owns. It is defined by how well the machinery, quality controls, and workflow systems fit together. Nebius appears to be working toward that kind of integration. Compute generates the force. Agentic tooling expands the scope of action. Toloka helps keep the system from confusing speed with dependability.

That is a sharper strategic position than simple GPU access.

And over time, it may prove to be the more durable one.

8. Customers, partnerships, and competitive positioning

A company like Toloka can make a compelling theoretical case on paper. It can talk about expert networks, alignment workflows, agentic reliability, and the future of post-training infrastructure. That all sounds promising. The harder test is whether serious customers are willing to trust it with real work. In Toloka’s case, the answer appears to be yes, and that is impactful far more than polished category language.

Customer quality is one of the cleanest validation signals in this part of the AI stack. Frontier labs, hyperscale platforms, and enterprise software companies do not casually hand over sensitive post-training, evaluation, and data workflows to weak vendors. These tasks sit close to model behavior, safety, and commercial performance. A company working in that layer needs technical credibility, operational discipline, and enough trust to be brought into workflows that customers often consider strategically important.

Toloka’s client and partner base suggests it has crossed that threshold. The names attached to the business are not random logos placed on a sales slide. They are the kinds of counterparties that indicate Toloka is operating inside the real infrastructure of the AI race rather than orbiting around it.

8.1 Toloka’s client base and market validation

The client list includes Amazon, Microsoft, Anthropic, Poolside, and Shopify.

Anthropic is likely the most strategically revealing. A company centered on safety, alignment, and high-quality model behavior does not choose data partners casually. If Toloka is contributing to workflows tied to a lab like Anthropic, that supports the idea that its capabilities in expert feedback, evaluation, and reliability are credible at the frontier edge of the market.

Amazon and Microsoft matter differently. These relationships place Toloka closer to the infrastructure and enterprise side of the AI ecosystem. The significance here is not only revenue potential. It is ecosystem placement. Working with companies of that scale suggests Toloka is legible to customers that care about procurement standards, reliability, and institutional credibility. In this part of the market, trust compounds. Once a vendor proves it can operate competently inside one demanding environment, it becomes easier to win the next one.

Shopify adds another layer. On one hand, the company appears through governance due to Mikhail Parakhin’s role as Shopify CTO and Executive Chairman of Toloka’s board. On the other hand, it also hints at a broader commercial bridge into enterprise AI use cases tied to real business workflows. That combination of governance proximity and commercial relevance strengthens the impression that Toloka is becoming more tightly connected to the North American technology ecosystem.

Poolside is interesting for a different reason. It represents exposure to newer frontier AI builders, not just incumbent platforms. It suggests Toloka can serve both established enterprise giants and high-growth labs that are trying to build next-generation systems. A vendor that can straddle both groups has a wider strategic surface area.

The broader point is simple. Toloka’s validation is showing up where it should. It is visible through customer quality, not just internal claims.

Customer or partner

Why it matters

Anthropic

Strong validation in alignment, safety, and high-quality frontier model workflows

Amazon

Signals relevance inside large-scale AI infrastructure and enterprise environments

Microsoft

Supports Toloka’s credibility with major model and software ecosystems

Shopify

Connects governance credibility with future enterprise AI workflow relevance

Poolside

Shows Toloka can serve emerging frontier model builders as well as large incumbents

That customer mix also supports the Nebius investment case indirectly. Toloka is not being validated in isolation by low-tier buyers. It is being validated by participants already shaping where AI infrastructure, model development, and enterprise adoption are heading. That raises the probability that Toloka is operating in a durable category rather than a temporary service niche.

8.2 The strategic value of the expert network

Toloka’s expert network is one of its most important assets, partly due to it is harder to replicate than it first appears. In software markets, people often assume a strong feature can simply be copied by a competitor with enough budget. That assumption works better for interfaces than for human systems. Expert networks are not just lists of people. They are supply chains of judgment.

Toloka has built a network of more than 10,000 verified experts across 90+ domains and 40+ languages. Those numbers matter on their own, though the strategic value sits in what they imply operationally.

First, the company has to recruit this talent. Second, it has to verify actual expertise rather than resume theater. Third, it has to route the right task to the right expert quickly enough that the workflow remains commercially useful. Fourth, it has to maintain quality control across outputs that are often nuanced, domain-specific, and difficult to adjudicate with simple rule-based checks. Fifth, it has to do all of that across languages and cultural contexts where the meaning of a strong answer can vary in important ways.

That is a difficult system to build. It requires labor-market access, screening systems, workflow design, incentive calibration, data architecture, and a long feedback loop between customer requirements and operational execution. A competitor can throw money at the problem and still end up with a messy roster of experts that does not perform reliably inside production environments.

This is why the expert network should be viewed less as a headcount statistic and more as infrastructure. It behaves more like a specialized logistics network than a marketplace widget. The value comes from reliability, routing, and quality assurance, not just from raw supply.

There is also a strategic timing element here. As frontier AI models exhaust more of the easy public-data gains, incremental improvement increasingly comes from proprietary post-training data, high-quality evaluation, and specialized human feedback. That should raise the value of expert networks over time, especially in domains where correctness is expensive and mistakes are hard to detect automatically.

A generic crowd can help label images. A verified network of physicists, lawyers, coders, and domain specialists can help shape how advanced models reason, how they are evaluated, and when they should defer. Those are different categories of economic value.

That network also matters inside the Tendem and agentic reliability story. Human-on-demand workflows only work if the human side is actually available, qualified, and structured enough to respond cleanly. Tendem without a real expert network would be a nice diagram. Tendem with Toloka’s network becomes much more credible.

8.3 How Toloka compares with Scale AI and adjacent peers

The most obvious peer in the category is Scale AI. That comparison is useful, though it should be handled carefully. Scale is the larger and more visible benchmark, especially in U.S. markets, and it has already achieved a much higher valuation profile. Toloka is smaller, less fully disclosed from a financial perspective, and still carries some complexity due to its history and current relationship with Nebius. At the same time, the comparison is still worth making due to it gives investors a practical reference point for how the market values similar capabilities.

Scale AI was valued at $13.8B in May 2024 and has reportedly been associated with valuation discussions in the $25B to $29B range in 2026. Research points to 2024 revenue around $870M and projected 2026 revenue near $2B. Those numbers matter less as precise anchors than as evidence of what the market is willing to pay for a company that sits inside the post-training, evaluation, and AI data infrastructure layer.

Toloka is not at that scale today. Its revenue is not disclosed as clearly, and valuation estimates range widely from roughly $1B to as high as $9.6B in more aggressive external frameworks. That is a very wide band, which tells you the market still lacks clean visibility. Even so, the strategic comparison still helps. It shows that this category can command significant value when customers believe the company sits close to the bottlenecks of model quality, safety, and deployment.

The more interesting comparison is qualitative rather than purely financial.

Scale AI has built a dominant U.S.-centered position with strong exposure to government and enterprise AI work. Toloka’s relative edge appears to be the combination of frontier post-training services, multilingual expert infrastructure, and increasingly its integration into the Nebius AI Factory through Tendem and the broader reliability layer.

In other words, Scale is a benchmark for what the category can be worth. Toloka may have a different path to value creation due to it sits inside a larger full-stack AI architecture rather than existing purely as a standalone data infrastructure company.

Category

Scale AI

Toloka

Market position

Larger, more visible category leader

Smaller, less visible, strategically embedded player

2024 valuation signal

$13.8B

Not publicly clear

2026 valuation discussion

$25B to $29B range reported

Estimate range roughly $1B to $9.6B

Revenue disclosure

More visible, with 2024 revenue around $870M

Limited public visibility

Core strength

Broad AI data infrastructure and large enterprise relevance

Expert network, multilingual judgment, post-training, evaluation, agentic reliability

Strategic context

Standalone AI infrastructure company

Integrated within the Nebius AI Factory and broader AI stack

Adjacent peers also matter, even if none maps perfectly. Some companies specialize in narrower model evaluation workflows. Others focus on contract research, annotation, or safety testing. Still others provide agent tooling without deep human-in-the-loop infrastructure. Toloka sits in the overlap of several categories at once. That can make it harder to classify quickly, though it can also be strategically valuable if those categories continue converging.

That classification problem is worth paying attention to. Markets often underprice assets that do not fit neatly into one existing bucket. Toloka has characteristics of an expert network, an evaluation platform, an alignment vendor, and now a live reliability layer for agentic systems. That is more complicated than a simple "AI data company" label. It may also be more valuable if the market eventually decides those layers belong closer together than apart.

From a Northwise perspective, the competitive takeaway is fairly clean. Toloka does not need to be Scale AI to matter. It needs to be important enough in a category that already commands premium valuations, credible enough with serious customers to keep compounding, and strategically integrated enough with Nebius that the parent benefits through both operating differentiation and asset optionality. On those dimensions, the setup looks increasingly credible.

9. Financial profile and valuation context

Toloka is easy to understand strategically and much harder to value cleanly from the outside. That gap is important. The company sits in a category the market increasingly treats as valuable, though it does so through a structure that still limits visibility. Revenue is not disclosed with the same clarity as a standalone public company.

Ownership has become more independent, though Nebius still retains a significant majority economic interest. The operating role is becoming more central to the Nebius platform, while the financial reporting is becoming less directly visible at the parent level. That creates a familiar situation in investing. The asset is real, the category is attractive, and the market still struggles to decide which yardstick belongs in its hands.

This is where discipline becomes important. It is very easy to slap a hot multiple on a strategic AI asset and call it insight. It is also easy to dismiss the asset entirely due to the valuation work is messy. Neither approach is serious. Toloka deserves a narrower, more structured treatment. The right question is not whether we can produce a perfect number. The right question is what range of values is plausible given the available signals, the peer set, and Toloka’s strategic role inside a fast-growing part of the AI stack.

9.1 What the Bezos round signaled

The clearest recent valuation signal is the May 2025 funding round led by Bezos Expeditions. On the surface, $72M is a financing event. In context, it is a much broader institutional signal.

First, it validated Toloka as investable to elite Western capital after the Yandex separation. Market access changes valuation even before any spreadsheet is built. A business that can raise clean external capital from a credible U.S. investor moves into a different category than one still carrying structural uncertainty around governance, counterparties, and geopolitical baggage.

Second, the round strengthened the board and institutional profile of the company through the governance reset that followed. This improves more than reputation. It improves future financing flexibility, makes outside benchmarking easier, and increases the probability that Toloka can eventually be valued as a serious standalone AI infrastructure asset rather than as a muddy internal subsidiary.

Third, the round implies that outside investors saw enough strategic value in the business to justify backing it during a period when the AI market had already become more selective. By 2025, investors were no longer funding anything with AI in the pitch deck and a gradient logo on the homepage. Capital was flowing hardest toward the choke points: compute, models, data, software infrastructure, and agentic tooling. Toloka sits near several of those choke points at once.

That still does not give us a disclosed post-money valuation. It gives us something slightly different and still useful. It tells us that sophisticated capital believed Toloka had crossed the threshold from interesting spinout to fundable strategic asset. In valuation work, that is not a finish line. It is the point where the game begins.

9.2 The difficulty of valuing Toloka from the outside

The challenge with Toloka is that several common methods each fail in a slightly different way.

A simple revenue multiple approach is tempting, though Toloka’s standalone revenue is not disclosed with enough precision to make that method clean. That forces outsiders to back into estimates, which can quickly become fragile.

A pure peer comp approach is useful directionally, though the peer set is imperfect. Scale AI is the obvious benchmark, but it is larger, more visible, more deeply embedded in the U.S. market, and shaped by a different ownership structure. Other adjacent companies may be relevant on narrower functions such as evaluation, annotation, safety, or enterprise AI services, though very few sit across expert data, alignment, evaluation, and live human-in-the-loop agent reliability in quite the same way.

A sum-of-the-parts approach inside Nebius is probably the most intellectually honest framework, though even that requires judgment. Toloka is not just a passive financial holding. Its strategic value inside the Nebius AI Factory likely exceeds what a cold external minority-stake math exercise would capture. At the same time, internal strategic relevance does not automatically convert into market-marked value unless the company becomes independently legible to outside investors.

That leaves us with a practical problem. Toloka should probably be valued using a blend of methods rather than any single clean formula.

The first layer is category valuation. What are sophisticated markets willing to pay for businesses operating near the post-training, alignment, evaluation, and reliability layers of AI?

The second layer is execution and scale. How large is Toloka today relative to those peers, and how credible is its path to becoming much larger?

The third layer is structural discount. How much should be shaved off due to limited public disclosure, ownership complexity, and the fact that the asset still sits inside a broader parent ecosystem?

The fourth layer is strategic premium. How much should be added back due to Toloka now functioning as more than a standalone vendor, particularly through Tendem and its integration into the Nebius stack?

This is why the valuation range is so wide. More aggressive external estimates stretch toward roughly $9.6B based on ambitious revenue assumptions and peer-style multiple logic. More conservative analyst-style comparisons place the implied value much lower, often in the $420M to $1B area depending on what share of future growth is considered credible. That is a large spread, though the spread itself tells a story. The market has not settled the category definition yet.

Toloka today is part hidden asset, part strategic operating layer, and part emerging standalone AI infrastructure company. Assets in that kind of transition phase rarely come with neat labels.

9.3 The range of plausible benchmark frameworks

A good framework here is to think in bands rather than precision. Toloka does not need a heroic valuation to matter for Nebius. It needs a credible enough valuation that the market eventually starts treating it as a real source of optionality rather than a background detail.

Framework

What it assumes

Implied reading of Toloka

Conservative comp framework

Modest scale, disclosure discount, lower confidence in near-term monetization

Valuable asset, though still treated cautiously by the market

Mid-range strategic framework

Meaningful growth in post-training and evaluation, stronger peer relevance, governance credibility improving

Serious AI infrastructure asset with room for standalone rerating

Aggressive category leader framework

High revenue scale, premium multiples, market begins valuing Toloka more like a top-tier AI data platform

Major hidden asset with substantial upside to current implied expectations

We can push that one step further conceptually.

Under a conservative lens, Toloka is worth enough to matter, though not enough to dominate the Nebius story. In that world, it acts as a supporting asset that reinforces the broader platform and provides some financial cushion.

Under a mid-range framework, Toloka starts to become more meaningful inside a Nebius sum-of-the-parts view. Here the market begins to recognize that Nebius owns exposure to two different layers of the AI stack: capital-intensive compute infrastructure and higher-value judgment infrastructure. That combination can change how investors think about durability and upside.

Under an aggressive framework, Toloka begins to resemble one of those assets that looked peripheral right up until the market suddenly decided it was core. This is where comparisons to Scale AI become more useful as category proof rather than as direct one-for-one equivalents. The point is not that Toloka must be Scale. The point is that the market has already shown a willingness to place very large values on companies operating near AI data, post-training, and deployment bottlenecks.

Toloka’s valuation today is better approached as an interval of plausibility rather than a precise mark. That may feel unsatisfying, though it is usually the correct posture when disclosure is limited and structure is still evolving. Good investing often begins with that kind of discomfort. The numbers are not clean enough for consensus, which is precisely why the opportunity can still exist.

For Nebius shareholders, the key financial point is simpler than the spreadsheet gymnastics. Toloka gives the parent exposure to a category that can command very high strategic valuations under the right conditions. Even if we stay conservative in the near term, the presence of that optionality changes the shape of the Nebius story. The market may still be focused on racks, megawatts, and cloud revenue. Toloka introduces another dimension, one tied to expert data, model reliability, and the layers of AI that often become more valuable once raw compute stops being the only scarce thing in the room.

Northwise Research

Get Northwise Research in Your Inbox

New research, model releases, portfolio updates, and major Northwise developments delivered directly to your inbox.

10. Why Toloka matters to the Nebius equity story

Toloka matters to the Nebius equity story due to it changes what kind of company investors are actually underwriting. A surface-level read of Nebius can still reduce the business to AI infrastructure ramp, capital deployment, GPU utilization, and the familiar math of turning megawatts into revenue. That part of the story is real and will remain central. It is also only part of the picture.

Once Toloka is taken seriously, Nebius begins to look less like a pure compute buildout and more like a company with exposure to multiple bottlenecks across the AI stack. One bottleneck is physical and obvious. It sits in data center capacity, energy access, cluster deployment, and the industrial work required to serve large-scale AI demand. The other bottleneck is subtler. It sits in expert data, post-training, evaluation, reliability, and the judgment layer increasingly required to make advanced systems commercially useful. Toloka is Nebius’s exposure to that second bottleneck.

That shift affects how markets often value businesses differently when they believe the company owns more than one valuable choke point. A single-asset story can grow quickly, though it usually trades within a narrower mental box. A multi-layer story, especially one spanning both infrastructure and higher-level reliability or data assets, gives investors more ways to be right over time.

10.1 The hidden asset case inside NBIS

The hidden asset case begins with a simple observation. Toloka is strategically important, externally validated, and increasingly relevant to Nebius’s broader platform, though it is still not the part of the business most investors focus on first. That creates the conditions for underappreciation.

The market’s attention naturally gravitates toward what is easiest to model. In Nebius’s case, that means capex, deployments, active power, contracted capacity, cloud revenue, ARR targets, customer announcements, and the operating leverage that could emerge if utilization scales cleanly. Toloka fits awkwardly into that framework. It is neither a neatly disclosed standalone public company nor a trivial side business that can be ignored. It lives in the middle, which is often where mispricing begins.

That middle ground is what makes the hidden asset framing reasonable. Toloka sits inside a category where comparable businesses can command substantial valuations. It has external validation through the $72M Bezos-led round. It has governance credibility through the addition of Mikhail Parakhin as Executive Chairman. It has operational relevance through the Tendem integration and the broader reliability layer inside the Nebius AI Factory. And it still appears to receive less attention than the compute story that surrounds it.

Hidden assets are rarely hidden in the literal sense. They are usually visible, though treated as secondary, hard to value, or too structurally messy for most investors to spend real time on. Toloka checks all three boxes.

The equity implication is not that Toloka must suddenly be valued like the largest AI data peers for the thesis to work. It is that Nebius may be carrying an asset with more strategic and financial significance than the current market framing fully captures. If the market is primarily underwriting a compute ramp while Toloka continues to gain external validation, product relevance, and eventual valuation clarity, then investors may eventually have to adjust what they believe they own.

10.2 Why Toloka changes the shape of Nebius upside

Toloka changes the shape of Nebius upside due to it introduces a second route to value creation beyond the obvious one. Nebius scales infrastructure, fills clusters, converts power into revenue, and eventually demonstrates that the AI cloud can produce large and durable cash flows. Investors already understand that path, even if they debate execution.

Toloka adds a different path. It gives Nebius exposure to a category tied to post-training, alignment, evaluation, and agentic reliability, all of which may become more valuable as the AI market matures. That means Nebius does not only benefit if compute remains the most important bottleneck. It can also benefit if the market begins placing more value on the layers that determine whether compute produces trusted outcomes.

The AI stack is unlikely to remain static. In the early phase of a platform buildout, raw infrastructure scarcity tends to dominate the narrative. Over time, as more supply comes online, value often shifts upward toward workflow integration, software, trust, reliability, and the layers where switching costs deepen. Toloka gives Nebius a claim on that migration.

There is also an asymmetry here that investors should pay attention to. Infrastructure upside tends to require large execution success. More power needs to be energized. More clusters need to come online. More customers need to ramp. More utilization needs to stick. Toloka upside can emerge through different mechanisms. It can come through external financing rounds, through improved peer benchmarks, through tighter integration into the Nebius stack, through stronger revenue quality inside enterprise workflows, or eventually through some separate liquidity event. In other words, not all of the upside depends on the same machine running perfectly.

That diversification of upside pathways has real equity value. It means Nebius shareholders are not only betting on one lever. They are also holding exposure to an asset that could rerate through a different set of market conditions and a different peer group.

Toloka also changes the margin narrative around Nebius, at least conceptually. A pure infrastructure business can become very large, though it also tends to remain capital intensive and more exposed to competition over time. A business exposed to expert data, evaluation, and reliability sits closer to specialized services and software-like workflow value. That does not make it magically simple or high margin, though it does make the overall Nebius profile more balanced than a bare reading of the cloud ramp might suggest.

The practical result is that Toloka can amplify Nebius upside in several ways at once. It can improve the platform strategically. It can widen how investors think about the company. It can create optionality around future monetization. And it can potentially make the Nebius story feel less one-dimensional as the market begins looking beyond the first phase of the compute race.

10.3 How investors should think about strategic relevance versus reported financials

Investors need to stay disciplined. Strategic relevance and reported financials are not the same thing, and confusing them is how people talk themselves into stories that never cash out.

Toloka is strategically relevant already. That case is relatively strong. It sits in valuable parts of the AI stack. It has high-quality customers and partners. It has a growing role inside the Nebius AI Factory. It helps explain why Nebius is trying to build beyond raw GPU access and into a more complete production environment for enterprise AI and agentic systems.

Reported financial value is less clean. Toloka’s standalone economics are not disclosed with enough detail to support false precision. That means investors should resist the urge to pretend they have a crisp answer when the inputs are clearly incomplete. A hazy asset should not be modeled like a bond.

The better way to think about the gap is through tiers of confidence.

At the highest-confidence level, Toloka has strategic value and likely strengthens Nebius’s competitive position.

At the medium-confidence level, Toloka probably has meaningful standalone financial value due to the category it operates in, the external validation it has received, and the scarcity characteristics of the business it is building.

At the lowest-confidence level, the exact mark is still uncertain, and investors should be careful about assigning aggressive numbers without a stronger disclosure base.

That structure is important due to it lets investors benefit from the idea without turning it into dogma. Strategic relevance can justify spending more attention on the asset, asking better questions, and broadening the Nebius framework beyond compute alone. It does not justify pretending the valuation is settled.

There is also a timing element. Public markets often lag the operating reality of assets like this. First the asset becomes strategically useful. Then it becomes operationally integrated. Then it attracts external validation. Only later does the market begin assigning it cleaner value, often after some financing round, spinout discussion, reporting change, or public transaction makes the numbers harder to ignore. Toloka appears to be somewhere in the middle of that sequence.

That is why the distinction between strategic relevance and reported financials should be seen as a feature of the current setup rather than a flaw in the thesis. If everything were already clean, obvious, and fully marked, there would usually be less room for upside surprise.

For Nebius investors, the right posture is therefore balanced. Toloka should not be treated as a free lottery ticket thrown on top of the compute story. It should be treated as an increasingly important part of the architecture, one that may also carry standalone value the market has not yet fully organized around. The compute ramp remains the engine of the current equity narrative. Toloka introduces a second layer of optionality, one tied to the parts of AI that become more important once the machines are built and the real question shifts from can it run to can it be trusted.

11. Valuing Toloka within the Nebius sum of the parts

Choose your Northwise access

Keep reading with the path that fits you

Create a Free Account

Access all public research and personalized alerts.

Create a Free Account

Join Northwise Premium

Unlock valuation outputs, downloadable models, portfolios, action frameworks, and complete Premium research.

Join Northwise Premium

Reader discussion

Discuss the research

0 published

Premium access is required to join this report's discussion.

Join Northwise Premium

No comments yet. Start a thoughtful discussion.