Story Sentient
Captured source
source ↗Sentient & Fireworks Powers Decentralized AI At Viral Scale
GLM 5.2 is live! Opus-level intelligence at open-source rates. Pay per token on serverless. Try it today.
Blog
Story Sentient Sentient & Fireworks Powers Decentralized AI At Viral Scale
PUBLISHED 7/17/2025
Table of Contents 50% Higher Throughput per GPU, Scaling Without Cost Inflation
Key Outcomes at a Glance: Meet Sentient: Visionary Builders with Big Stakes What Sentient Built The Infrastructure Challenge The Fireworks Solution: Performance without Complexity
Timeline: 30 Days from Hackathon to Viral Launch What Fireworks Delivered
⚡ Fast, Consistent Inference
🌐 High-Concurrency Resilience
🛠 Dedicated Custom Deployments Results That Moved the Needle
Real-World Stats at Launch A Strategic Partnership
Table of Contents
50% Higher Throughput per GPU, Scaling Without Cost Inflation
Key Outcomes at a Glance:
• 1.8M+ Waitlisted Users in 24 Hours : Viral launch built for massive early interest for Sentient's products. • 25-50% More Concurrent Users per GPU : Industry-leading efficiency vs. tested alternatives. • Enterprise-Grade Performance without Overhead : Cost-effective scale, even under extreme concurrency. • Rapid Iteration & Launch: From hackathon to public release of a 70B model powering multi-agent chat in weeks, not months.
Meet Sentient: Visionary Builders with Big Stakes
Backed by $85 million from Founders Fund, Pantera, Framework Ventures, and Polygon Labs, Sentient unites Sandeep Nailwal (Polygon), Himanshu Tyagi (Witness Chain), and Princeton professor Pramod Viswanath at the helm. Their Princeton-driven research team is chasing a single, audacious goal: deliver the ultimate AI experience by fusing the planet’s collective intelligence into one open, decentralized network. Powered by blockchain and open-source models, Sentient turns transparency into a feature and democratizes AI for everyone. At the helm of product, Technical Product Manager Oleg Golev leads the charge in bringing that vision to life – starting with Dobby , an open-source family of LLMs showcasing AI loyalty at the model layer, fine-tuned to be loyal to personal freedom and the crypto community. The models possess unique qualities (distinct personality traits and human-like tone) that make it a perfect choice for content virality, while maintaining academic breakthroughs in post-training value and safety alignment. “The Open World is the world we want to live in, but it is only possible by leveraging blockchain to make AI more transparent and just.” - Sandeep Nailwal, Cofounder of Polygon & Sentient source Their flagship app, Sentient Chat , initially integrated 15 specialized AI agents to deliver fast, complex workflows for research, productivity, and search — alongside Open Deep Search (ODS), a fast, transparent search alternative that outperformed closed-source systems on search benchmarks (SimpleQA and FRAMES) that outperformed Perplexity and ChatGPT on search benchmarks. These efforts target systems like ChatGPT and Gemini as part of a broader vision to scale community-driven AI products that compete with closed-source incumbents.
What Sentient Built
Dobby Arena : An early experimental platform for community feedback that supported initial tuning of the Dobby models and user experience—though the real impact comes from the Sentient Chat, multi-agent framework, and underlying model innovations. Sentient Chat : A production-grade multi-agent assistant powered by Dobby-70B, launched virally at the Open AGI Summit during ETH Denver with over 1.8 million users waitlisted in 24 hours. Open Deep Search : A complementary project which enables transparent, high-speed search that supports Sentient’s vision of decentralized AI infrastructure. Achieving SOTA benchmarks on SimpleQA and FRAMES benchmarks, ODS is built to challenge opaque algorithms and pairs naturally with Sentient Chat to deliver fast, explainable search in multi-agent workflows. These products required infrastructure that could handle real-time inference, extreme concurrency, and unpredictable traffic—all without compromising on latency or reliability. The Infrastructure Challenge
Sentient’s products went viral fast. But viral success brings infrastructure pain, such as: Concurrency Bottlenecks : Multi-agent reasoning, multiple LLM calls, and real-time search required low latency performance to maintain user trust, especially during multiturn conversations. Unpredictable Spikes: Waitlist-gated product launches opened to users via random access code leads to sudden spikes, sometimes thousands of concurrent users with no time for manual scaling. Costly Inefficiencies at Scale: Internal GPU clusters and infra like vLLM would have required more GPUs for less throughput. No Room for Downtime : Slowdowns meant lost momentum against fast-moving competitors like ChatGPT and Claude. The Fireworks Solution: Performance without Complexity
Sentient benchmarked multiple infra providers, including custom silicon options. Fireworks outperformed all, delivering up to 50% more throughput per GPU and more consistent performance under real-world load. This translated into fewer GPUs, lower costs, and seamless launches. “In our first app, we recorded 1.5 million responses in five days , from 90,000 unique users , with up to 1,000–2,000 active users at any given time . That was with a query cap of 10–20 per user .” Fireworks provided a custom-engineered, high-performance infrastructure platform platform built on NVIDIA Blackwell and designed specifically for high-concurrency, burst-tolerant AI workloads for top performance under extreme load. Starting in January 2025, they used: Serverless endpoints for fast iteration, testing, and deployment Custom-dedicated deployments for real-time inference in Sentient Chat and Dobby Arena FP8-optimized hardware enables high throughput for tasks like summarization and sentiment analysis while efficiently utilizing GPU resources.This setup let Sentient iterate rapidly and scale confidently—without building and managing hyperscale infrastructure themselves. Fireworks evolved from a technical solution to a strategic growth multiplier. Timeline: 30 Days from Hackathon to Viral Launch
Fireworks began working with Sentient in early January 2025 to support the Dobby model rollout and community engagement campaigns. 📅 Date 🚀 Launch Milestone
Jan 25 Sentient x Fireworks Hackathon (150–200 attendees) with early access to Dobby-Mini Jan 27...
Excerpt shown — open the source for the full document.
Notability
notability 3.0/10Routine post without traction evidence.