WritingFireworks AIFireworks AIpublished Feb 12, 2026seen Jun 26

Story Sentient

Open original ↗

Captured source

source ↗
published Feb 12, 2026seen Jun 26captured Jun 28http 200method plain

Sentient & Fireworks Powers Decentralized AI At Viral Scale

GLM 5.2 is live! Opus-level intelligence at open-source rates. Pay per token on serverless. Try it today.

Blog

Story Sentient Sentient & Fireworks Powers Decentralized AI At Viral Scale

PUBLISHED 7/17/2025

Table of Contents 50% Higher Throughput per GPU, Scaling Without Cost Inflation

Key Outcomes at a Glance: Meet Sentient: Visionary Builders with Big Stakes What Sentient Built The Infrastructure Challenge The Fireworks Solution: Performance without Complexity

Timeline: 30 Days from Hackathon to Viral Launch What Fireworks Delivered

⚡ Fast, Consistent Inference

🌐 High-Concurrency Resilience

🛠 Dedicated Custom Deployments Results That Moved the Needle

Real-World Stats at Launch A Strategic Partnership

Table of Contents

50% Higher Throughput per GPU, Scaling Without Cost Inflation

Key Outcomes at a Glance:

• 1.8M+ Waitlisted Users in 24 Hours : Viral launch built for massive early interest for Sentient's products. • 25-50% More Concurrent Users per GPU : Industry-leading efficiency vs. tested alternatives. • Enterprise-Grade Performance without Overhead : Cost-effective scale, even under extreme concurrency. • Rapid Iteration & Launch: From hackathon to public release of a 70B model powering multi-agent chat in weeks, not months.

Meet Sentient: Visionary Builders with Big Stakes

Backed by $85 million from Founders Fund, Pantera, Framework Ventures, and Polygon Labs, Sentient unites Sandeep Nailwal (Polygon), Himanshu Tyagi (Witness Chain), and Princeton professor Pramod Viswanath at the helm. Their Princeton-driven research team is chasing a single, audacious goal: deliver the ultimate AI experience by fusing the planet’s collective intelligence into one open, decentralized network. Powered by blockchain and open-source models, Sentient turns transparency into a feature and democratizes AI for everyone. At the helm of product, Technical Product Manager Oleg Golev leads the charge in bringing that vision to life – starting with Dobby , an open-source family of LLMs showcasing AI loyalty at the model layer, fine-tuned to be loyal to personal freedom and the crypto community. The models possess unique qualities (distinct personality traits and human-like tone) that make it a perfect choice for content virality, while maintaining academic breakthroughs in post-training value and safety alignment. “The Open World is the world we want to live in, but it is only possible by leveraging blockchain to make AI more transparent and just.” - Sandeep Nailwal, Cofounder of Polygon & Sentient source Their flagship app, Sentient Chat , initially integrated 15 specialized AI agents to deliver fast, complex workflows for research, productivity, and search — alongside Open Deep Search (ODS), a fast, transparent search alternative that outperformed closed-source systems on search benchmarks (SimpleQA and FRAMES) that outperformed Perplexity and ChatGPT on search benchmarks. These efforts target systems like ChatGPT and Gemini as part of a broader vision to scale community-driven AI products that compete with closed-source incumbents.

What Sentient Built

Dobby Arena : An early experimental platform for community feedback that supported initial tuning of the Dobby models and user experience—though the real impact comes from the Sentient Chat, multi-agent framework, and underlying model innovations. Sentient Chat : A production-grade multi-agent assistant powered by Dobby-70B, launched virally at the Open AGI Summit during ETH Denver with over 1.8 million users waitlisted in 24 hours. Open Deep Search : A complementary project which enables transparent, high-speed search that supports Sentient’s vision of decentralized AI infrastructure. Achieving SOTA benchmarks on SimpleQA and FRAMES benchmarks, ODS is built to challenge opaque algorithms and pairs naturally with Sentient Chat to deliver fast, explainable search in multi-agent workflows. These products required infrastructure that could handle real-time inference, extreme concurrency, and unpredictable traffic—all without compromising on latency or reliability. The Infrastructure Challenge

Sentient’s products went viral fast. But viral success brings infrastructure pain, such as: Concurrency Bottlenecks : Multi-agent reasoning, multiple LLM calls, and real-time search required low latency performance to maintain user trust, especially during multiturn conversations. Unpredictable Spikes: Waitlist-gated product launches opened to users via random access code leads to sudden spikes, sometimes thousands of concurrent users with no time for manual scaling. Costly Inefficiencies at Scale: Internal GPU clusters and infra like vLLM would have required more GPUs for less throughput. No Room for Downtime : Slowdowns meant lost momentum against fast-moving competitors like ChatGPT and Claude. The Fireworks Solution: Performance without Complexity

Sentient benchmarked multiple infra providers, including custom silicon options. Fireworks outperformed all, delivering up to 50% more throughput per GPU and more consistent performance under real-world load. This translated into fewer GPUs, lower costs, and seamless launches. “In our first app, we recorded 1.5 million responses in five days , from 90,000 unique users , with up to 1,000–2,000 active users at any given time . That was with a query cap of 10–20 per user .” Fireworks provided a custom-engineered, high-performance infrastructure platform platform built on NVIDIA Blackwell and designed specifically for high-concurrency, burst-tolerant AI workloads for top performance under extreme load. Starting in January 2025, they used: Serverless endpoints for fast iteration, testing, and deployment Custom-dedicated deployments for real-time inference in Sentient Chat and Dobby Arena FP8-optimized hardware enables high throughput for tasks like summarization and sentiment analysis while efficiently utilizing GPU resources.This setup let Sentient iterate rapidly and scale confidently—without building and managing hyperscale infrastructure themselves. Fireworks evolved from a technical solution to a strategic growth multiplier. Timeline: 30 Days from Hackathon to Viral Launch

Fireworks began working with Sentient in early January 2025 to support the Dobby model rollout and community engagement campaigns. 📅 Date 🚀 Launch Milestone

Jan 25 Sentient x Fireworks Hackathon (150–200 attendees) with early access to Dobby-Mini Jan 27...

Excerpt shown — open the source for the full document.

Notability

notability 3.0/10

Routine post without traction evidence.