Neocloud analysis
Standing syntheses the agent writes over each lab's captured pages, structured signals, and bounded web evidence — every material claim cited back to its source.
CoreWeaveCoreWeave3wCoreWeave is transitioning from a GPU-rental neocloud into a full-stack AI cloud platform with serious public-company scale and discipline. The evidence points to three reinforcing vectors: (1) infrastructure velocity — $5B annual revenue growing 168% YoY, 850MW active power across 43 data centers globally, and first-to-market NVIDIA Vera Rubin NVL72 bring-up; (2) software-stack deepening — a unified agentic AI…Cloudflare (Workers AI)Cloudflare (Workers AI)3wCloudflare is executing a multi-vector pivot from content-delivery infrastructure toward AI-native compute and agentic intermediation. The evidence shows three reinforcing moves: (1) Workers AI is being hardened as an inference platform for trillion-parameter open-weight models with commercial-grade routing and billing, (2) an Agents SDK and first-party agent framework (Flue) are being layered on top to capture…BasetenBaseten3wBaseten is in the midst of a breakout scaling phase fueled by a $1.5B Series F. The company is tripling headcount while simultaneously shipping at high velocity across three vectors: developer tooling (CLI, MCP server, Chains GA), model performance research (BEI embeddings, speculative decoding, timestep distillation), and GPU infrastructure breadth (B200, GH200, H100/H200, multi-node). The hiring pattern skews…ClarifaiClarifai3wClarifai is in wind-down mode as a standalone entity. The company's platform and inference IP have been acquired by Nebius (NASDAQ: NBIS), with all Clarifai services ceasing on July 17, 2026. New signups closed May 19 and new payments stop June 17, 2026. Recent GitHub activity reflects a final synchronization push of gRPC SDKs across seven languages to version 12.5.1, with no release notes published—consistent with…CerebrasCerebras3wCerebras is executing a hard pivot from wafer-scale training hardware specialist to full-stack inference cloud provider. The evidence shows a company that IPO'd in May 2026, closed an $850M revolving credit facility, and is aggressively building out inference datacenters across North America and Europe with a target of 20x aggregate capacity expansion. The wafer-scale architecture delivers inference speeds 10–20x…Databricks (DBRX)Databricks (DBRX)3wDatabricks is executing a three-front offensive: it is monetising enterprise data gravity through Lakebase (a serverless Postgres for operational/AI workloads), hardening an agent platform (Genie One, Genie Agents, Omnigent) that turns the Lakehouse into an operating system for enterprise agents, and building a specialised research pipeline — data agents trained with RL — that seeks to match frontier model…Eigen AIEigen AI4wEigen AI is a 2025-founded inference optimization company acquired by NASDAQ-listed neocloud Nebius for $643M, with the deal closing on 10 June 2026. The lab's optimization stack is being integrated into Nebius Token Factory to deliver production inference at scale. Active hiring for post-training/inference engineering and platform product management signals a dual buildout: continued technical R&D on model serving…WaferWafer4wWafer is a hardware-centric AI inference platform building competitive advantage through GPU kernel optimization expertise, with a distinctive multi-vendor strategy spanning NVIDIA and AMD accelerators. The evidence depicts a company vertically integrated from low-level kernel engineering up to a serverless inference product, using public benchmarks and developer education content as both recruiting and go-to-market…Public AIPublic AI4wPublic AI is not a frontier model lab; it is a neocloud-adjacent public infrastructure play building an inference utility positioned as an open, democratically governed alternative to commercial AI APIs. The org's GitHub activity reveals a concentrated push to operationalize a chat-based inference platform (chat.publicai.co) atop OpenWebUI, with CI/CD pipeline maturation, API gateway deployment, and multi-geography…CompactifAI (Multiverse Computing)CompactifAI (Multiverse Computing)4wCompactifAI (Multiverse Computing) is transitioning from a quantum-software R&D shop into a commercial AI infrastructure company anchored by model compression. Its proprietary CompactifAI technology applies quantum-inspired tensor-network mathematics to prune and restructure pre-trained LLMs, producing "Slim" variants that retain reasoning and tool-use capabilities at reduced inference cost. The lab is running a…Blackbox AIBlackbox AI4wBlackbox AI is not a frontier model builder; it is an inference-infrastructure and agent-orchestration platform that competes on serving others' models faster, cheaper, and more securely than anyone else. The company's public signals converge on a single bet: that enterprise and government adoption of coding agents will be won at the orchestration and inference layer, not at the model-training layer [W1, W5]. With…MakoraMakora4wMakora is a performance-engineering organization focused on automated GPU kernel generation and inference optimization. Its public surface spans four categories: (1) an AI-driven kernel generation system (MakoraGenerate) that produces optimized GPU kernels targeting NVIDIA H100/B200, AMD MI300X, and Tenstorrent hardware; (2) a lightweight multi-vendor GPU querying utility (gpuq) supporting CUDA and HIP runtimes; (3)…StreamLake (Kuaishou)StreamLake (Kuaishou)4wKuaishou's StreamLake is executing a dual-pronged open-source strategy: a video-native multimodal foundation model line (Keye-VL-2.0) aimed at long-video understanding and agentic capabilities, and an AI coding product suite (KAT-Coder, CodeFlicker, Vanchin) targeting the software engineering tool market. Both tracks are anchored in Apache 2.0 releases, rigorous public benchmarking against frontier models (GPT-5,…ParasailParasail4wParasail is an early-stage AI infrastructure company (Series A, $32M raised, $160M valuation) building a serverless inference cloud that aggregates distributed GPU supply into an OpenAI-compatible platform for open-weight models. The company positions itself as a GPU-network orchestration layer: workloads are automatically matched across a multi-provider GPU network, freeing developers from vendor lock-in and…GMI CloudGMI Cloud4wGMI Cloud is an inference-optimized neocloud building a full-stack platform tightly coupled to NVIDIA's hardware roadmap. The evidence shows a company transitioning from bare-metal GPU provisioning to a managed platform layer: 10 open roles are clustered around a named "Inference Engine" product, AgentBox has shipped as an agent marketplace and hosting platform, and every public post ties GMI's infrastructure…ScalewayScaleway4wScaleway is executing a three-pillar strategy to differentiate as Europe's sovereign AI cloud: (1) serving frontier open-weight models through its Generative APIs platform as a managed alternative to proprietary hyperscalers, (2) embedding sustainability and CSRD compliance tooling directly into its cloud product portfolio, and (3) investing in European AI infrastructure sovereignty through consortia spanning…Novita AINovita AI4wNovita AI is executing a two-pronged evolution: it operates a commercial model API and agent sandbox platform for third-party frontier models, while simultaneously building deep inference infrastructure—most visibly pegaflow, a Rust-based KV cache storage engine with vLLM integration—that targets the performance bottleneck of large-scale LLM serving. The pattern of GTM hiring in San Mateo alongside a relentless…SiliconFlowSiliconFlow4wSiliconFlow is building a "Token Factory" — an AI inference infrastructure layer that normalizes heterogeneous compute into standardized token output. The GitHub evidence reveals a two-track product strategy: (1) a deep acceleration stack (OneDiff, Nexfort) that compiles diffusion and LLM workloads for faster inference, and (2) BizyAir, a cloud-hosted ComfyUI service that wraps model access into a managed developer…HyperbolicHyperbolic4wHyperbolic is in a post–Series A scaling sprint, pivoting from its early Web3/decentralized microservices roots into a full-stack GPU marketplace aggregator. The evidence shows a company simultaneously hiring for infrastructure depth (GPU orchestration, bare-metal provisioning, SRE), commercial operations (supply, finance, GTM), and developer tooling (CLI, MCP, AI SDK, Gradio). The recent Forge launch crystallizes…FriendliAIFriendliAI4wFriendliAI is an AI inference infrastructure company entering an aggressive commercialization phase, signaled by a $20M funding round, a rapid SDK iteration cadence with breaking API changes across all serving tiers, the launch of a public OpenAPI schema, and day-zero support for frontier open-weight models. The dual-hub (Seoul/San Francisco) hiring pattern reveals simultaneous investment in core inference engine…DigitalOcean (GradientAI)DigitalOcean (GradientAI)4wDigitalOcean (GradientAI) is executing a concentrated pivot into agentic AI infrastructure as a managed cloud service, building the full stack from GPU inference to hosted agent runtimes. The evidence reveals a coordinated three-pronged buildout: (1) an Inference Engine now generally available with frontier model support across OpenAI, Anthropic, and fal; (2) a Codex plugin in Public Preview that provisions…Lightning AILightning AI4wLightning AI is in the midst of a structural transformation from developer-framework shop into a vertically integrated neocloud. The merger with Voltage Park [P2, P3] has reshaped the company's operational DNA: it now owns and operates physical data centers across at least three US geographies (Quincy WA, Fort Worth TX, Lisle IL) [P22, P23], is building bare-metal GPU compute, storage, and observability…ReplicateReplicate4wReplicate is a post-acquisition platform operating as an inference API aggregator, not a model builder. Following its acquisition by Cloudflare, its activity centers on platform engineering — evidenced by an intense cog release cadence across v0.16–v0.21 — ecosystem integration (SDKs in Python and JavaScript, MCP, LangChain, agent skills), and positioning as the hosted API layer for third-party frontier and…DeepInfraDeepInfra4wDeepInfra is an inference-cloud provider exploiting the open-weight model boom, not a model-building lab. Its GitHub footprint reveals a company systematically forking and maintaining the full inference-serving stack — from CUDA kernels to serving engines to client SDKs — while its $107M Series B and targeted hiring confirm a bet on inference infrastructure as a standalone business. The org tracks frontier…Snowflake (Arctic)Snowflake (Arctic)4wSnowflake is executing a deliberate convergence play: its Arctic model family — specialized for SQL, code generation, and enterprise retrieval — is being positioned not as a standalone frontier contender but as the AI inference layer inside a governed, agentic data platform. The firm's public writing, hiring, and releases all orbit a single narrative: "the agentic enterprise". Arctic now spans speculators…SambaNova SystemsSambaNova Systems4wSambaNova Systems is executing a decisive pivot from AI training hardware toward becoming an inference cloud provider purpose-built for agentic AI workloads. The evidence pack captures a company compressing its stack around three interlocking bets: (1) disaggregated/hybrid inference pairing its own SN40 RDU with NVIDIA GPUs for prefill-decode splitting [E29, E53, W2]; (2) "premium inference" as a differentiated…Fireworks AIFireworks AI4wFireworks AI is a Series C ($4B valuation) generative AI infrastructure platform transitioning from inference-speed leader to full-stack AI cloud provider, with training, fine-tuning, serverless and dedicated inference, multi-LoRA serving, and agentic orchestration all built on proprietary infrastructure. The most recent evidence — spanning June 2026 — reveals a company in an intensive GTM buildout phase, anchored…GroqGroq4wGroq is rebuilding as a pure-play AI inference cloud after a transformative non-acquisition by Nvidia that took its founding CEO, president, and key engineers. A $650M raise in June 2026 aims to scale GroqCloud to 200MW by 2027 and serve 5M developers on its purpose-built LPU chip architecture. The evidence pack shows Groq rapidly maturing SDK tooling (Python v1.5.0, TypeScript v1.3.0), building an evaluation and…Together AITogether AI4wTogether AI is consolidating its positioning as the AI-native cloud — an inference-first infrastructure platform that competes on raw speed and cost per token. The evidence pack shows the company simultaneously building out in three directions: (1) deepening the infrastructure surface from GPU clusters into managed storage, networking, and observability, (2) layering enterprise trust and access-control primitives…NebiusNebius4wNebius is executing a multi-front AI cloud scaling thesis: it is simultaneously building out physical data center capacity across the US and Europe, expanding its GPU orchestration software stack, commercializing a new agentic search product (Tavily), and deepening its research bench through an acqui-hire (Clarifai). The hiring pattern reveals a company transitioning from infrastructure provider to full-stack AI…