WritingFireworks AIFireworks AIpublished Feb 12, 2026seen Jun 26

Why Gpus On Demand

Open original ↗

Captured source

source ↗
published Feb 12, 2026seen Jun 26captured Jun 27http 200method plain

GPUs on-demand: Not serverless, not reserved, but some third thing

GLM 5.2 is live! Opus-level intelligence at open-source rates. Pay per token on serverless. Try it today.

Blog

Why Gpus On Demand GPUs on-demand: Not serverless, not reserved, but some third thing

PUBLISHED 6/3/2024

Table of Contents Tl;dr Intro Fireworks Offerings On-demand GPUs performance details Conclusion

Table of Contents

Note: Fireworks A100 and H100 prices have since been reduced to $2.90 and $5.80! Tl;dr

• On-demand GPUs are a great option for scaling companies who need reliability and speed but cannot yet commit to long-term enterprise reservations of GPUs • “Graduating” from serverless to on-demand deployments starts to make sense economically when you are running ~100k+ tokens per minute • On identical H100 hardware, Fireworks’ software advantage provides 53% cost reduction and 60% latency reduction compared to running vLLM with GPUs on competing platforms. This allows you to simultaneously serve more users, at lower cost, with faster response times • Fireworks on-demand deployments require no software installation. Simply choose your model and GPU configuration and get started in seconds, without paying for boot times. Your GPU even scales to zero when you are not using it!

Intro

One of the most rewarding things at Fireworks is being a part of the scaling journey for many, exciting AI start-ups. Over the last few months, we’ve seen an explosion in the number of companies beginning to productionize generative AI. A question that we commonly get is: “How should I think about serving LLMs via (1) A Serverless, token-based options vs (2) A GPU, usage time-based option? “ We’re writing this post to help explain the tradeoffs of serverless vs dedicated GPU options. Fireworks Offerings

Fireworks has 3 offerings for LLM serving: (1) Serverless (2) On-demand (3) Enterprise. Serverless Fireworks hosts our most popular models “serverless”, meaning that we operate GPUs 24/7 to serve these models and provide an API for any of our users to use the models. Our serverless offering is the fastest, widely available platform and we’re proud of its production-readiness. Serverless is the perfect option for running limited traffic and experimenting with different LLM set-ups. However, serverless has limitations: Serverless is not personalized for you Your LLM serving can be made faster, higher-quality or lower cost based on personalization on several levels: Our serverless platform is designed to excel at a variety of goals and support hundreds of base models and fine-tuned models. However, a personally-tailored stack still may provide better experiences. • GPU configuration - The quantity, type and configuration of GPUs can be tailored based on business needs, like optimizing for speed vs cost or providing more capacity for product launches • Serving stack - The software itself ****can be tuned based on factors like prompt length and UX goals (latency vs throughput, etc) • Models - Serverless model selection is curated by Fireworks, so niche models may not be available

Serverless performance is affected by others - Other Fireworks users share our serverless deployment, so speeds vary depending on overall usage. If you happen to use the deployment at the emptiest hours, you’ll experience the fastest speeds and vice versa. Our public platform is still independently benchmarked to have the lowest variation in latency, but consistent performance is paramount for certain use cases, like live voice chat agents. Serverless has volume constraints - Generally, when businesses have the volume to use significant GPU capacity, it doesn’t make sense to use serverless because: • Rate limits don’t allow it - Since GPU deployments are shared, we employ rate limits to ensure that a few actors can’t greatly affect everyone else’s experience. Our rate limits (600 RPM by default) are significantly higher than rate limits of other serverless inference providers (often below 50 RPM) but businesses with significant volume could struggle to run their app entirely on Fireworks • Private GPUs may be cheaper with large volume - Given (a) Efficiency improvements from personalized set-ups and (b) GPU pricing that builds in “bulk discounts” compared to serverless, large businesses generally receive cheaper pricing by reserving their own GPU

Enterprise Reserved GPUs Given these constraints, companies with large usage volumes often reserve their own private GPU(s) for set periods of time. This commitment also enables Fireworks to help companies personally configure their serving set-up and provide SLAs and guaranteed support. However, many scale-up companies in the midst of prototyping are unable to commit to enterprise reserved capacity. On-demand GPUs To make it easier for scaling teams to benefit from Fireworks, we offer on-demand, dedicated GPUs with our FireAttention stack. Users pay per hour and can scale their GPU usage up or down automatically based on traffic. Configurations can automatically scale up and down from 0 GPUs, so developers pay nothing during idle periods. Compared to serverless, using your own GPU can provide: • Guaranteed speed and reliability by using private GPUs • More model choice - import your custom base model or use dozens of Fireworks-provided base models • Improved speed - configure the number of set-up of GPUs based on your goals and optimize the Fireworks serving stack based on your prompt length • Lower costs, especially with high volume. Costs reduce as volume increases. • Capacity for higher request volume , with no hard rate limits

On-demand GPUs performance details

An important consideration in deciding to use on-demand deployments is expected performance vs price, so we have included some performance details and FAQs. What latency improvements can I expect compared to vLLM or hosting my own GPU? Generally, we see that Fireworks is ~40-60% faster compared to open-source solutions but performance varies depending on model and workload. We obtained the below results using the min worldsize (minimum # GPUs to host a specific model) on H100 GPUs on Fireworks vs vLLM software. The latency was calculated with heavy throughput, so the GPU on Fireworks simultaneously delivered significantly better speed and throughput. Mixtral 8x7b Prompt Lengths (Tokens) Fireworks Latency vLLM Latency Very long prompt (32000 input, 100 output) 3319 ms (at 0.293 QPS) 6049 ms (at 0.165 QPS) Long prompt (4000 input, 200...

Excerpt shown — open the source for the full document.

Notability

notability 1.0/10

Low traction promotional blog post.