New in Together GPU Clusters: Reliability and control for production GPU clusters
Captured source
source ↗New in Together GPU Clusters: Reliability and control for production GPU clusters Webflow Analyze/Optimize tracking bridge -->
💰 Announcing our Series C. Intelligence should be abundant, not expensive →
📊 Delivering 31% more TPS than the next-fastest OSS engine for production coding agent workloads →
🇫🇷 Join us at RAISE 2026 in Paris →
⚡ On-demand B200s now available on Together GPU Clusters →
🚀 Now serving MiniMax-M3 for efficient inference →
All blog posts
GPU Clusters
Published 7/15/2026
New in Together GPU Clusters: Reliability and control for production GPU clusters
Authors
Pavneet Ahluwalia
Table of contents
40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...
We’ve spent the last several weeks shipping a set of changes to Together GPU Clusters aimed at the operational reality of running training and inference at scale: hardware fails, schedulers leak, and teams outgrow the single-admin-kubeconfig workflow they started with. This post walks through what we shipped, why we built it the way we did, and what it means for the workloads you’re running on Together today. The changes group into two themes. The first is platform health : passive health checks, auto node repair, and Slinky 1.0, focused on catching and recovering from the failure modes that actually take down jobs. The second is operational control : a new cluster details view, external OIDC, startup scripts, and an acceptance-test opt-out, focused on giving your team the visibility, access, and customization hooks needed to run clusters as your organization grows.
Catching and fixing failures as they happen If you’ve run a multi-day training job on a large cluster, you know the pattern. A GPU falls off the PCIe bus. An Xid error takes a node out of rotation. Thermal throttling silently caps a job’s throughput and you don’t notice until the loss curve flattens. These are steady-state failure modes at scale. What matters is how quickly you catch them and how cleanly you recover. We already ran active health checks , synthetic tests that exercise the hardware against a known-good baseline. Active checks are useful at provisioning time and on idle nodes. Passive checks extend that coverage to failures that appear while real workloads are running.
Fig: Health checks tab shows all the active health-checks currently running and the historical information So we built passive health checks , which work like alerts and run continuously across every node in your cluster, observing real workloads, logs, and metrics to surface degradation as it happens. The coverage list today includes GPUs falling off the bus, thermal throttling, Xid errors, Slurm node drains on failure, and an expanding set of hardware and software failure signals. The checks observe live workloads with near-zero overhead on running jobs. Learn more about health checks here .
Fig: Repair tab for node repair recommendations and repair history of the cluster Detection on its own is just a dashboard, so we paired passive checks with auto node repair . When our monitoring system detects a node-level issue, it generates a recommended remediation and surfaces it for an operator to review. There are four repair actions, mapped automatically based on the failure signature: Reboot: In-place restart that preserves local data. The default for transient issues Reprovision: Rebuilds the node from a clean image and clears local data Failover: Moves the workload to a fresh bare-metal node and clears local data Remove: Pulls the node out of the pool and sends it to RMA
We believe in automation with a human touch: Your training checkpoints and inference replicas are too valuable to risk on automated drains. Our 'human-in-the-loop' approach keeps you in control — the system detects and recommends, you approve, and Together handles the graceful recovery. It’s the perfect balance of intelligence and safety. We’re already building fully automated node-repair options for specific failure modes, but for now, we prioritize keeping your production workloads uninterrupted. Learn more about node auto repair here . The result? You’ve traded hours of support tickets for minutes of in-product workflow. From detection to resolution, you’re back in the pool faster than ever. Together Slurm-on-K8s 2.0: The future of Slurm on Kubernetes Running Slurm on Kubernetes shouldn't feel like a constant battle against crashing daemons, zombie processes, or scheduler drift. We’ve rebuilt our Slurm-on-Kubernetes stack from the ground up, based on our fork of OSS project Slinky, to make those 'scale-at-failure' headaches a thing of the past.
Here’s what our upgraded stack brings to your cluster: Self-healing worker daemons: Transient failures are inevitable, but they shouldn't take down your node. Our new stack supervises workers, auto-restarting them in place so long training runs stay resilient. No more zombie processes: Forget the days of orphaned processes clogging your PID tables and blocking new jobs. Our stack automatically reaps orphans, ensuring your nodes stay clean and ready for work, every single time. Durable job accounting: sacct history used to live on ephemeral storage, which meant a pod restart could wipe your entire accounting database. We've moved accounting to durable, PVC-backed storage, so restarts and reschedules no longer touch your data. Billing reconciliation and post-hoc job analysis stay intact across the lifecycle of the cluster. Reliable process cleanup: When jobs ended, daemonized child processes used to slip out from under Slurm's view and stay running — holding GPU memory and /dev/shm segments hostage between runs, sometimes for days until someone cleaned them up by hand. The new stack tracks every descendant of a job at the kernel level and cleans them all up reliably at job end. The next job on the node starts on a clean machine every time. Accurate GPU state after reschedules: After a pod reschedules, Slurm's view of which GPUs existed used to drift from reality — stale GPU identifiers from the previous incarnation would mismatch the real hardware, and affected GPUs would silently drop out of the schedulable pool. The new stack rebuilds Slurm's GPU view fresh on every node start, so the schedulable pool always matches the hardware actually in the node. Beyond reliability, our new stack exposes DCGM metrics in your cluster's Grafana...
Excerpt shown — open the source for the full document.