Llama32 With Fireworks
Captured source
source ↗Partnering with Meta: Bringing Llama 3.2 to Fireworks for Fine-Tuning and Inference
GLM 5.2 is live! Opus-level intelligence at open-source rates. Pay per token on serverless. Try it today.
Blog
Llama32 With Fireworks Partnering with Meta: Bringing Llama 3.2 to Fireworks for Fine-Tuning and Inference
PUBLISHED 9/25/2024
Table of Contents Key Highlights Llama 3.2 Adds New Multimodal Capabilities, Expanding Use Cases For Developers Pricing and Deployment Options: The Great News
Serverless
On-Demand
Enterprise Reserved Customize Llama 3.2 With Fine-Tuning Try the models out directly in our playground Next, try using the models with our inference APIs
Install the Fireworks AI Python package
Accessing Llama 3.2 on Serverless Inference API
Table of Contents
We are excited to announce support for the newest additions to the Llama collection from Meta. With the addition of Llama 3.2 , developers gain access to new tools that enable the creation of sophisticated multi-component AI systems that combine models, modalities, and external tools to deliver advanced real-world AI solutions. Llama 3.2: Seeing The World More Clearly (And Quickly)
The release of Llama 3.2 1B , Llama 3.2 3B , Llama 3.2 11B Vision , and Llama 3.2 90B Vision models brings a range of text-only and multimodal models designed to enhance modular AI workflows. These models provide deep customization, allowing developers to tailor solutions and accelerate specific tasks in compound AI systems. Get started today on Fireworks: • Llama 3.2 1B (text-only) : Ideal for retrieval and summarization tasks such as personal information management , multilingual knowledge retrieval , and rewriting tasks . • Llama 3.2 3B (text-only) : Optimized for query and prompt rewriting , it supports applications like mobile AI-powered writing assistants and customer service tools running on edge devices . • Llama 3.2 11B Vision and Llama 3.2 90B Vision : These models extend capabilities with image understanding and visual reasoning for tasks such as image captioning , visual question answering , and document visual analysis .
The instruct variants of these models are available serverless (pay-per-token for models on Fireworks-configured GPUs). Both the instruct and non-instruct variants of these models are available on-demand (private GPU instances billed per GPU second). Meta’s Llama Guard 3 models for detecting violating content are also available on-demand. Key Highlights
Text Multimodal Recommendation Use Case Llama 3.2 1B ✅ Ideal for retrieval and summarization tasks. This model can be effectively deployed for personal information management, multilingual knowledge retrieval, and rewriting tasks running locally on edge devices, offering efficiency in handling data close to the user. Llama 3.2 3B ✅ Recommended for query and prompt rewriting and optimized for mobile AI-powered writing assistants and edge devices. This model is perfect for customer service applications and mobile AI-powered writing assistants, allowing businesses to deploy high-performance models that deliver real-time AI solutions at scale. Llama 3.2 11B Vision ✅ ✅ Optimized for image understanding and visual reasoning, it excels in image captioning, image-text retrieval, visual grounding, and document visual question answering. This model is critical for any application needing advanced visual and text-based comprehension, such as enterprise search and expert copilots in areas like coding, math, and medicine. Llama 3.2 90B Vision ✅ ✅ Similar to the 11B model, this larger model offers exceptional performance in visual question answering, visual reasoning, and other complex multimodal tasks. Its ability to handle both image and text inputs allows for a range of applications, from image captioning to document analysis, making it ideal for industries like healthcare, legal, and finance.
Llama 3.2 Adds New Multimodal Capabilities, Expanding Use Cases For Developers
The release of multimodal models unlocks exciting new production use cases for developers, from enterprise to everyday applications. Examples of use cases for Llama 3 models on Fireworks includes: • Visual Question Answering and Reasoning : In healthcare, clinicians can use multimodal systems to ask questions about medical images, like "Is there a fracture in this X-ray?" The system analyzes the image, provides a precise answer, and highlights key areas, enabling faster, more accurate diagnoses and reducing human error in time-sensitive situations. • Document Visual Question Answering : For document-heavy fields like legal and finance, visual-language models can extract specific information from PDFs or charts, such as "What is the total amount due?" This reduces manual effort, speeds up analysis, and boosts accuracy in reviewing complex documents. • Image Captioning : In retail, compound-AI systems can automatically generate product descriptions from images, such as "A sleek black leather handbag with gold hardware." The system analyzes the product image and creates a detailed, engaging caption that enhances customer experience and boosts metrics like conversion rates. By eliminating the need for manual captioning, this approach enables businesses to quickly scale as their product catalogs grow, while maintaining consistency and accuracy.
See how customers like AlliumAI are supporting multimodal models in production with Fireworks in this blog post . Start Small Then Scale Quickly and Efficiently with Llama 3.2 Models on Fireworks
Llama models , fine-tuned and deployed through Fireworks, offer developers the flexibility to build personalized AI systems tailored to specific needs. With Fireworks handling the fine-tuning and inference , developers can leverage these powerful tools to accelerate innovation and bring their AI solutions to market faster. For example, Fireworks can serve Llama 3.2 1B in approximately 500 tokens/second and Llama 3.2 3B in 270 tokens/second. Pricing and Deployment Options: The Great News
There’s no one-size-fits-all approach to developing compound AI systems, which is why Fireworks offers a number of different options for using and deploying models like Llama-3.2 for production AI (including serverless, on-demand, and enterprise reserved ). We’re also happy to announce new, competitive pricing for text and multimodal models, especially for the Llama 3.2 - 11B and Llama 3.2 - 90B multimodal models which will be the same price as the text-only models . Images will be...
Excerpt shown — open the source for the full document.
Notability
notability 4.0/10Routine integration post, low traction.