Neocloudfresh 16h

Baseten

Signal timeline484 total
Jul 26, 2026
Jul 24, 2026
1dJobGRC ManagerSan Franciscosource
Jul 23, 2026
3dReleasebasetenlabs/truss v0.18.23basetenlabs/trusssource
Jul 22, 2026
Jul 21, 2026
5dReleasebasetenlabs/baseten-go v0.2.0basetenlabs/baseten-gosource
5dReleasebasetenlabs/baseten-python v0.10.0basetenlabs/baseten-pythonsource
5dReleasebasetenlabs/baseten-js v0.1.0basetenlabs/baseten-jssource
Jul 17, 2026
1wJobPeople Business Partner, GTM San Francisco - Routine job postingsourcenotability 1.0/10
1wReleasebasetenlabs/truss v0.18.22basetenlabs/trusssource
Jul 16, 2026
1wReleasebasetenlabs/truss v0.18.21basetenlabs/trusssource
Jul 15, 2026
1wReleasebasetenlabs/truss v0.18.21rc0basetenlabs/trusssource
Jul 14, 2026
1wReleasebasetenlabs/truss v0.18.20basetenlabs/trusssource
Jul 13, 2026
1wReleasebasetenlabs/baseten-cli v0.3.0basetenlabs/baseten-clisource
Jul 10, 2026
2wReleasebasetenlabs/truss v0.18.19basetenlabs/trusssource
Jul 9, 2026
2wJobEngineeringSan Franciscosource
2wReleasebasetenlabs/truss v0.18.19rc1basetenlabs/trusssource
Jul 8, 2026
Jul 7, 2026

Top signals

  1. #1WritingSota Performance For Gpt Oss 120b On Nvidia Gpus8.0
  2. #2WritingHow We Made The Fastest Gpt Oss On Nvidia Gpus 60 Percent Faster7.0
  3. #3WritingKimi K2 Explained The 1 Trillion Parameter Model Redefining How To Build Agents7.0
  4. #4WritingNvidia Nemotron 3 Nano7.0
  5. #5WritingNvidia Nemotron 3 Nano Omni7.0

Agent answer

Baseten has 484 loaded public signals: 90 hiring, 68 forks, 78 releases or model cards, 204 talking, and 44 repos. Latest signal: How We Built The New Fastest Api For Glm 52. Data-business radar is currently scoped to frontier labs, so this category does not expose radar lanes. The standing analysis was generated with deepseek-v4-pro and 90 evidence refs.

Baseten

has loaded 484 public signals

Baseten

has hiring signal count 90

Baseten

has fork signal count 68

Baseten

has release signal count 78

Analysis — agent synthesisfull report →generated July 3, 2026

Thesis

Baseten is in the midst of a breakout scaling phase fueled by a $1.5B Series F W2E2. The company is tripling headcount W2 while simultaneously shipping at high velocity across three vectors: developer tooling (CLI, MCP server, Chains GA), model performance research (BEI embeddings, speculative decoding, timestep distillation), and GPU infrastructure breadth (B200, GH200, H100/H200, multi-node). The hiring pattern skews heavily toward GTM/commercialization (SDRs, partnerships, field enablement, revenue ops, customer marketing) alongside select infrastructure hires (capacity engineering, TPM-infra), signaling a transition from pure platform build to aggressive enterprise go-to-market. Fork activity clusters around container runtimes, GPU interconnects, and DeepSeek tooling — suggesting ongoing investment in low-level inference infrastructure and multi-node serving. The public writing focuses on GPU buyer's guides, model benchmarks, and inference performance, positioning Baseten as the pragmatic infrastructure layer for companies switching from closed-source APIs to open models at production scale.

Signal desks

Hiring

  • Sales Development Representative - Inbound — San Francisco, Sales and Business Development team. First point of contact for inbound prospects; qualifies against ICP and builds pipeline. Signals GTM team buildout post-Series F. E4P5
  • Recruiting Coordinator — San Francisco, Talent team. Operational hire supporting headcount tripling. E5
  • Ecosystem Partnerships Product Marketings Manager — San Francisco, G&A team. Drives partner storytelling, co-marketing campaigns, and enablement programs with ecosystem partners. Signals partnership-led growth strategy. E9P9
  • Field Productivity & Enablement Lead — San Francisco, Revenue Operations team. Builds enablement programs for field teams. Signals maturing revenue organization. E22
  • Analyst, Revenue Strategy & Operations — San Francisco, GTM team. Revenue analytics and strategy. E24
  • Frontend Engineer — San Francisco, Dedicated Inference team. Frontend hire for the core inference product UI. E25
  • Partnerships Product Marketing Manager — San Francisco, Product Marketing team. Joint GTM with technology partners. E27
  • Capacity Strategy & Operations Lead — San Francisco, Compute team. GPU capacity planning and operations, critical for a capital-intensive inference business. E42
  • Software Engineer - Capacity — San Francisco, Internal Platform (Dev Tooling) team. Building internal capacity tooling. E44
  • Customer Marketing Manager — San Francisco, Marketing team. Customer advocacy and marketing. E45
  • Product Manager, Developer Experience — San Francisco, Product team. PM focused on DevEx, aligned with CLI/MCP/Chains investments. E51
  • Technical Program Manager, Infrastructure — San Francisco, Infrastructure team. TPM for infrastructure programs. E54

Pattern: Hiring is concentrated in San Francisco, overwhelmingly GTM and revenue operations roles. Engineering hires are targeted at capacity and infrastructure teams, not broad platform engineering. Tripling headcount W2 with emphasis on commercialization suggests Baseten is entering a revenue-scaling phase.

Forks

  • basetenlabs/go-containerregistry — Fork of google/go-containerregistry (Go). Container registry manipulation library. Core infrastructure plumbing for model container management. E8P7
  • basetenlabs/rlm — Fork of alexzhang13/rlm. RL/ML library. E12
  • basetenlabs/ucx — Fork of openucx/ucx. High-performance communication framework for GPU interconnects. Critical dependency for multi-node inference. E14
  • basetenlabs/ucxx — Fork of rapidsai/ucxx. C++ bindings for UCX. Pairs with ucx fork for GPU communication layer work. E15
  • basetenlabs/container-debug-support — Fork of GoogleContainerTools/container-debug-support. Container debugging tooling. E31
  • basetenlabs/DeepGEMM — Fork of deepseek-ai/DeepGEMM. DeepSeek's optimized GEMM library, pointing to DeepSeek model serving optimization work. E35
  • basetenlabs/ideogram4 — Fork of ideogram-oss/ideogram4. Image generation model work. E55
  • basetenlabs/tinker-cookbook — Fork of thinking-machines-lab/tinker-cookbook (1 star). RL/agent cookbook. E56
  • basetenlabs/runc — Fork of opencontainers/runc. Core container runtime, fundamental to model serving infrastructure. E57
  • basetenlabs/compact-rl — Fork of PrimeIntellect-ai/prime-rl (7 stars). Reinforcement learning, notably higher community engagement than other forks. E58
  • basetenlabs/TorchSpec — Fork of lightseekorg/TorchSpec. PyTorch specification/optimization work. E59
  • basetenlabs/mcore-bridge — Fork of modelscope/mcore-bridge. Megatron-core bridge for model parallelism. E60
  • basetenlabs/action-junit-report — Fork of mikepenz/action-junit-report (0 stars, 6 open issues). CI/CD test reporting for internal dev workflows. P2
  • basetenlabs/run-report-action — Fork of moonrepo/run-report-action (1 star, 7 open issues). CI run reporting, likely used with moonrepo monorepo tooling. P3

Pattern: Forks cluster around three themes: (1) container and GPU interconnect infrastructure (runc, go-containerregistry, ucx/ucxx, container-debug-support), (2) model optimization and parallelism (DeepGEMM, mcore-bridge, TorchSpec), and (3) CI/CD developer productivity (junit-report, run-report). The ucx/ucxx pair and DeepGEMM fork suggest active work on multi-node inference optimization. The compact-rl fork (7 stars) hints at RL research interest.

Releases

  • basetenlabs/langchain-baseten v0.2.1 — LangChain integration package release. E13
  • basetenlabs/truss v0.18.8 → v0.18.17 — High-cadence Truss model packaging framework releases (10+ versions across the evidence window). Indicates active development on the model deployment toolchain. E20E26E28E30E34E37E40E41E43E48E52
  • basetenlabs/baseten-cli v0.2.0 — New CLI release. Paired with homebrew tap and changelog announcement. E21E6P6P10
  • basetenlabs/run-report-action v1 — CI report action release. P1

Pattern: Truss is under extremely active development with near-daily releases, suggesting it is a core piece of Baseten's model deployment workflow. The baseten-cli v0.2.0 and homebrew tap represent a new DevEx surface area.

Talking

  • Series F announcement — $1.5B raise led by Altimeter Capital, Conviction Partners, Spark Capital. Capital allocated to talent, compute, enterprise GTM. Tripling headcount. W2E2
  • GPU buyer's guides and benchmarks — H100 vs H200 vs B200 comparison P4E3; GH200 inference testing with Lambda Cloud P13; B200 acceleration benchmarks P26P27; Qwen 3 day-zero SGLang benchmarks P28; GLM 5.2 fastest API E1; Llama 3.3 70B GH200 testing P13. Consistent positioning as the pragmatic GPU inference expert.
  • Developer tooling launches — New baseten CLI with Homebrew distribution P6E6P10; MCP server for coding agents with model deployment capabilities P8E7; CLI log streaming and filtering E33P24; async log downloads E17; docs refresh P22; new sidebar navigation E47; vLLM and SGLang metrics E53; container restart tracking E49.
  • Platform infrastructure — Baseten Chains GA for compound AI systems P12P14; BEI embeddings inference engine with 2x throughput claims P15P19; multi-node inference for DeepSeek-R1 on 16 H100s P17; rolling deployments for zero-downtime updates E46; configurable scale-down rate P11E10; flexible instance types per deployment P25; full OpenAI API compatibility P18.
  • Model launches and integrations — GLM 5.2 available E38E1; Kimi K2.7 Coder E39; Mercury 2 E50; NVIDIA BioNeMo Agent Toolkit E29; Chroma vector database integration P23; FLUX.2 timestep distillation (HuggingFace release) W1; DeepSeek-V3.1 and MiniMax M2.5 deprecation E19.
  • Research & methodology content — Live draft model training for speculative decoding E18; how BEI was built with TensorRT-LLM P19; AI training vs inference explainer E16; checklist for switching to open-source models P20; deployment guide for open-source text embedding models P21; best open-source LLMs roundup E36; running GLM 5.2 in any harness E23.

Pattern: Public writing is high-volume and spans three registers: (1) infrastructure performance content that drives developer SEO and positions Baseten as GPU inference experts, (2) product changelogs demonstrating shipping velocity, and (3) model launch announcements that signal breadth of supported models. HN traction is modest (5-6 points) but consistent.

Shipping

Baseten's shipping cadence across the evidence window is exceptionally high. Key shipped artifacts include:

  • Baseten CLI v0.2.0 with Homebrew distribution — deploy, call models, stream logs, check metrics, manage deployments, all with --output json and --jq for scripting P6E6P10E21
  • Baseten MCP server — Connect coding agents to manage workspaces, deploy/promote models, tune autoscaling, pull logs, launch training jobs, with read-only/mutating labeling P8E7
  • Baseten Chains GA — SDK for compound AI systems with independent hardware and autoscaling per step, ultra-low-latency data exchange P12P14
  • Baseten Embeddings Inference (BEI) — 2x throughput and 10% lower latency vs alternatives, built on TensorRT-LLM, supporting embedding/reranker/classifier models P15P19
  • NVIDIA B200 GPU support (early access) — 5x higher throughput, 50%+ lower cost per token, 38% lower latency for large models P26P27
  • Full OpenAI-compatible APIs for chat completions and completions P18
  • Configurable autoscaling scale-down rate (1-50%) P11E10
  • Rolling deployments for zero-downtime model updates E46
  • Flexible instance types per deployment with environment-aware promotion P25
  • Multi-node inference in production for models like DeepSeek-R1 (16 H100s across 2 nodes) P17
  • FLUX.2 timestep distillation — 2.5x faster image generation, model on HuggingFace W1
  • Chroma vector database integration with BEI for embedding workflows P23
  • Truss v0.18.8 → v0.18.17 — sustained rapid iteration on model packaging framework E20E26E28E30E34E37E40E41E43E48E52
  • LangChain integration v0.2.1 E13

Research themes

1. Inference runtime optimization — BEI built on TensorRT-LLM with custom optimizations for embedding, reranker, and classifier workloads, achieving 2x throughput improvements P19P15. SGLang and vLLM integration for LLM serving with published metrics P28E53.

2. Speculative decoding — Live draft model training for speculative decoding, a technique that uses a smaller "draft" model to accelerate LLM token generation E18.

3. Timestep distillation for diffusion models — Applied Distribution Matching Distillation (DMD) to FLUX.2, reducing sampling steps from 20 to 4-8 while preserving image quality. Model released on HuggingFace W1.

4. Multi-node inference — Production system for serving models too large for single 8-GPU nodes (e.g., DeepSeek-R1 at 671B parameters across 16 H100s), addressing both infrastructure (interconnects, multi-cloud abstractions) and model performance (tensor parallelism optimization) P17.

5. GPU benchmarking and architecture analysis — Systematic evaluation of H100/H200/B200/GH200 for inference tradeoffs, including SXM vs PCIe, NVLink bandwidth, KV cache offloading on GH200, and FP4 support on B200 P4P13P26.

6. Compound AI orchestration — Chains SDK for multi-model workflows with independent hardware and autoscaling per step, addressing the latency/reliability/cost challenges of chained inference P12.

Hiring & scaling

Baseten is in an aggressive post-Series F scaling phase. The $1.5B raise is explicitly allocated to "talent, compute, and accelerating enterprise go-to-market," with stated plans to triple headcount W2.

Current open roles (12 identified in evidence):

| Function | Roles | Signal | |---|---|---| | GTM / Revenue | SDR-Inbound, Revenue Strategy Analyst, Field Productivity & Enablement Lead, Customer Marketing Manager | Commercialization ramp; building inbound-to-pipeline engine | | Partnerships & Ecosystem | Ecosystem Partnerships PMM, Partnerships PMM | Co-marketing and joint GTM with technology partners | | Infrastructure & Capacity | Capacity Strategy & Operations Lead, SWE-Capacity, TPM-Infrastructure | GPU supply chain and compute platform scaling | | Product & DevEx | PM-Developer Experience, Frontend Engineer (Dedicated Inference) | CLI, MCP, and UI investments | | Talent Operations | Recruiting Coordinator | Supporting tripled headcount |

All roles are based in San Francisco with hybrid/remote flexibility. The hiring skew — heavy GTM, selective infrastructure — indicates Baseten is transitioning from a product-led engineering org to a commercialization-stage company. The dedicated Capacity roles point to the capital-intensive nature of GPU inference, where supply chain and capacity planning are existential.

Category implications

Inference-as-a-Service is consolidating around a full-stack platform model. Baseten's simultaneous investments in low-level GPU optimization (ucx forks, DeepGEMM), model serving frameworks (Truss, BEI, Chains), developer tooling (CLI, MCP server), and enterprise GTM signal that the inference platform category is maturing past "GPU rentals" into integrated solutions spanning research, infrastructure, and developer experience P16P12P6P8.

GPU supply chain operations are becoming a core competency. The Capacity Strategy & Operations Lead and SWE-Capacity hires, combined with detailed public GPU benchmarking content (H100/H200/B200/GH200), suggest that inference providers must build dedicated capacity planning functions to manage GPU procurement, allocation, and multi-cloud abstraction E42E44P4P13. The B200 early access program shows Baseten is competing on hardware access speed P26P27.

Agent infrastructure is an emerging platform surface. The MCP server launch, which allows coding agents to deploy models, tune autoscaling, and launch training jobs, positions Baseten as infrastructure for the agent-building ecosystem. Combined with the NVIDIA BioNeMo Agent Toolkit blog E29 and Chains for compound AI P12, Baseten is layering agent orchestration on top of raw inference P8E7.

Open-source model migration is a core GTM motion. Multiple blog posts guide users through switching from closed-source APIs (GPT, Claude) to open-source models on Baseten, with content covering model selection, GPU sizing, and optimization checklists P20P21E36. The full OpenAI-compatible API removes a friction point for migration P18.

The embeddings market is a distinct and growing inference subcategory. BEI's dedicated optimization for embedding/reranker/classifier workloads — with published 2x throughput claims — and the Chroma integration suggest Baseten sees embeddings as a separate, high-volume inference workload worthy of specialized infrastructure, not just a side feature of LLM serving P15P19P23.

DevEx tooling is a competitive moat. The rapid Truss release cadence (10+ versions), new CLI with Homebrew distribution, MCP server, and dedicated DevEx PM hire indicate Baseten believes developer tooling quality — not just GPU pricing or availability — determines platform stickiness P6P8E51.

Traction highlights

  • $1.5B Series F led by Altimeter Capital, Conviction Partners, and Spark Capital W2E2. Previously raised $75M Series C P16.
  • Named customers include Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer P5P9.
  • Tripling headcount to meet demand W2.
  • HN traction is modest but consistent: 5-6 points on major posts E1E2.
  • GitHub forks with community engagement — compact-rl fork has 7 stars, notably above other forks E58; tinker-cookbook has 1 star E56. Most forks have 0 stars, indicating internal-use patterns rather than community-building.
  • Truss is Baseten's highest-activity public repo with sustained rapid release cadence .
  • Model breadth covering LLMs (GLM 5.2, Kimi K2.7, DeepSeek V4, GPT OSS 120B, Qwen 3), embeddings, speech (Whisper Large V3), and image generation (FLUX.2 distilled) P6E38E39P28W1.