WritingFireworks AIFireworks AIpublished Feb 12, 2026seen Jun 26

Traces Are All You Need

Open original ↗

Captured source

source ↗
published Feb 12, 2026seen Jun 26captured Jun 28http 200method plain

Fireworks AI

GLM 5.2 is live! Opus-level intelligence at open-source rates. Pay per token on serverless. Try it today.

Blog

Traces Are All You Need Traces Are All You Need (to rank LLMs)

PUBLISHED 9/22/2025

Table of Contents Introduction TL;DR A Step-by-Step Explanation

Step 1: Starting with Raw Production Data

Step 2: Deconstructing Conversations for Better Comparisons

Step 3: Generating New Responses with a Rollout Processor

Step 4: The Judgment: Pairwise Comparison and Scoring ⚖️

Step 5: Synthesizing a Final Score The Real-World Proof — Validating the Results Conclusion

Table of Contents

From your existing observability platform logs to a data-driven model leaderboard in minutes – quickly compare candidate models with an LLM judge. Introduction

Choosing the right AI model is a critical decision, yet it’s often a guess. Public benchmarks don't reflect the real-world trade-offs between cost, speed, and quality on your data. What if you could find the optimal model by building a leaderboard from your production logs in just five minutes? This post shows you how to find out using Eval Protocol , an open-source toolkit for building your internal model leaderboard. We’ll demonstrate a quick, no-ground-truth-required method and validate it by showing our results correlate strongly with the official Tau Bench Airline benchmark . While our Quickstart Guide covers the code, this article goes under the hood to explore the step-by-step methodology—inspired by Arena-Hard-Auto research—for turning raw logs into a validated model leaderboard. TL;DR

Our LLM Judge results (left), generated automatically from production traces, show a strong correlation with the official Tau Bench scores (right). This validates a fast, label-free method for accurately ranking models on your specific use case. Keep reading to see how you can make your own leaderboard! A Step-by-Step Explanation

Step 1: Starting with Raw Production Data

Everything begins with real user interactions. For this analysis, we used a dataset of airline customer service conversations that were already being logged to Langfuse . Each log, or trace, contains the full back-and-forth between a user and the AI assistant, including any tool calls. This raw, unfiltered data is the perfect source material because it reflects how our models perform in the wild. In this case, these traces were generated using kimi-k2-instruct . This is important because kimi-k2-instruct will serve as our baseline, and all the comparisons later on will be checking if a model is better or worse than this baseline. Trace Input: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 { "messages" : [ { "role" : "system" , "content" : "\nYou are a customer service agent that helps the user according to the provided below..." } , { "role" : "user" , "content" : "Hi, I need to talk to someone about my flight HAT039 from Atlanta to Seattle. It’s been delayed and I’m really upset about it. Can you tell me why this happened?" } , { "role" : "assistant" , "content" : "I understand you're frustrated about the delay on flight HAT039 from Atlanta to Seattle. Let me check the current status of your flight to see what's happening." , "tool_calls" : [ { "id" : "functions.get_flight_status:0" , "type" : "function" , "function" : { "name" : "get_flight_status" , "arguments" : { "flight_number" : "HAT039" , "date" : "2024-05-15" } } } ] } , { "role" : "tool" , "tool_call_id" : "functions.get_flight_status:0" , "content" : { "observation" : "delayed" } } , { "role" : "assistant" , "content" : "I can see that flight HAT039 from Atlanta to Seattle is indeed showing as delayed today. However, I don't have access to the specific reasons for the delay - that information would typically come from the airline's operations team or announcements at the airport.\n\nSince you're upset about this delay, I'd be happy to help you with any changes you might need to make to your reservation. If you have a reservation on this flight and would like to explore options like rebooking or if you're considering cancellation, I can assist with that.\n\nTo help you further, could you please provide:\n1. Your user ID\n2. Your reservation ID (if you have it)\n\nThis will allow me to look up your specific reservation and discuss what options are available to you due to this delay." } ] , "tools" : [ { "type" : "function" , "function" : { "name" : "cancel_reservation" , "description" : "Cancel the whole reservation." , "parameters" : { "properties" : { "reservation_id" : { "description" : "The reservation ID, such as 'ZFA04Y'" , "title" : "Reservation Id" , "type" : "string" } } , "required" : [ "reservation_id" ] , "title" : "cancel_reservationArguments" , "type" : "object" } } } { ... } , { ... } , ] }

Trace Output: 1 2 3 4 5 6 { "content" : "I'd be happy to help you book a flight from San Francisco to New York for three people! To get started, I'll need some information from you.\n\nFirst, could you please provide your user ID so I can access your profile?\n\nAlso, I'll need to know:\n- What type of trip are you looking for - one way or round trip?\n- What date(s) for the flight(s)?\n- What cabin class would you prefer - basic economy, economy, or business?\n\nOnce I have this information, I can search for available flights and help you complete the booking." , "role" : "assistant" , "tool_calls" : null , "function_call" : null }

Step 2: Deconstructing Conversations for Better Comparisons

A single multi-turn conversation isn't a great test case on its own. If we just give the whole chat history to a new model, we're only testing its ability to write the final message. We need to test how a challenger model would have behaved at every single turn of the conversation. Here’s how it works: Every time the original assistant replied in a conversation, the function creates a new test case. Each test case consists of: messages : The history of the conversation right before the assistant's turn. ground_truth : The exact response the original assistant gave.

This effectively turns one long conversation into multiple, independent challenges. Step 3: Generating New Responses with a Rollout Processor

With a clean set of test cases prepared, the next step is to generate responses using our...

Excerpt shown — open the source for the full document.

Notability

notability 7.0/10

Substantive research post from Fireworks AI, potentially notable.