Solar Pro 4
Captured source
source ↗Solar Pro 4: The Agentic Model That Finishes the Job
Upstage Studio — Build and deploy your agent. Learn more → Upstage Studio — Build and deploy your agent. Learn more →
Solutions
Resources
Company
Pricing
Try demo Try demo Contact us
Solar Pro 4 is built to carry real work to the finish, reading the documents, running the tools, producing the deliverable, and to stop and say so when the evidence runs out. The work you hand to an AI rarely ends with a single question. It's reviewing a contract, reconciling numbers across files, and verifying intermediate results before passing them to the next step. One wrong value or one stalled step, and you're re-checking everything. We built Solar Pro 4 against that bar. Compared with Solar Pro 3, its agent capability took a major step up, and the gains are largest on evaluations that resemble real work: long documents, terminal tasks, multi-turn tool use. And when the evidence isn't there, it says it can't verify instead of making something up. [Start with Solar Pro 4 on Upstage Console] To mark the launch, Solar Pro 4 is 90% off on Upstage Console and OpenRouter through September 10. Solar Pro 4 supports a 512K context with up to 128K output tokens, the baseline for agent work that loads several contracts, reports, and data files into a single session without splitting them. It handles English, Korean, and Japanese for both input and output, and response speed is a dial: set reasoning effort high for deep analysis, low for real-time interaction, where responses come back at everyday-chatbot speed. Full specs are in the developer docs . Built for Agent Work Solar Pro 4's strongest results land on the evaluations that decide whether real work gets done. Terminal tasks (Terminal-Bench v2.1): 57 What this means: complete multi-step jobs in a live shell, not just generating commands.
Multi-turn tool use (τ³-Banking): 23 What this means: find the right policy in a large knowledge base and act on it correctly across a multi-turn, tool-calling conversation.
Long-document reasoning (AA-LCR): 71 What this means: reason across ~100k tokens of reports and filings at once, synthesizing answers scattered across multiple documents.
- Scores from Artificial Analysis ( artificialanalysis.ai ) as of August 2026; public listing of our result upcoming.
Each of these is a several-fold step over Solar Pro 3, but the more useful read is what the scores mean: the model finishes terminal tasks, keeps its footing across many tool calls, and stays accurate deep into long documents. Solar Pro 4 finishes work because it was trained on finished work. OfficeVerse, Upstage's pipeline since Solar Open 2, synthesizes office tasks from real public data across 11 industry domains and 12 task types and grades each one pass or fail on the final deliverable. Solar Pro 4 was trained and validated on work in the same shape it takes in the real world. Ko-GDPval, our Korean office-work benchmark, came from the same effort.
Real Work, Start to Finish: From Excel to Report to Slides A high benchmark score means little if the model can't finish the job. So we handed Solar Pro 4 an entire assignment: a store-location analysis for Solarbean Coffee, a fictional coffee brand. The input was one store-opening policy document and six market-data files. In three prompts, Solar Pro 4 screened ten candidate sites against the policy and produced three deliverables in sequence.
Three prompts turned market data into an Excel workbook, a review report, and a slide deck. From data analysis to three final deliverables, in one session. Ten candidate sites screened against the policy, with PASS/FAIL marked by conditional formatting. Organized into a five-sheet workbook. A number computed in the workbook carries into the review report unchanged. That's the point of this job: values from one step survive into the next deliverable. We're releasing seven work agents in the same family as the Solar Pro 4 Agent Cookbook : Excel workbook generation, evidence-based Word reports, and PowerPoint decks with charts and speaker notes, each with its system prompt, actual outputs, and pass criteria. The cookbook also includes an English version of this exact assignment — a fictional Austin site-selection review with the company standards, the raw data, and the pass criteria — so you can run the job yourself. Answers Built on Evidence The most dangerous answer in real work is not a plainly wrong one. It's an unverified claim stated as fact. Once a number or clause that isn't in the document gets invented, it flows downstream untouched, and eventually someone has to re-review everything. So Solar Pro 4 is built to say it can't verify rather than fill the gap. We gave it a fictional service agreement and a quotation, then asked ten questions. Half have answers in the documents. The other half are traps: clauses that don't exist, a premise the contract contradicts, and an amount that differs between the two documents.
Each question gets its own verdict. Clauses that exist are answered with citations (grounded), clauses that don't are called out as absent (not-in-document), and a figure that differs between the two documents is reported as a discrepancy (mismatch). Only half of this question is in the document. The model answers liquidated damages from Article 16 and states that the early-termination penalty isn't provided for. Find what's there. Don't invent what isn't. Doing both at once is what matters. A model that always answers will fabricate; a model that always hedges is safe but useless. A trustworthy model isn't the one with an answer every time, but the one that first checks whether an answer has grounds. You can see the same behavior in the Solar Pro 4 Agent Cookbook' s research agent, which links every sentence to a source and labels anything it can't ground as unverified. Solar Open 2 and Solar Pro 4: Which One, When Upstage ships two current models with different jobs. Solar Open 2 is a general-purpose open-weights model you deploy yourself. It fits organizations that run on their own hardware, operate on-premises, and can't let data leave the building. Solar Pro 4 targets longer, more complex commercial agent workloads: jobs that move across multiple documents, run tasks in a terminal, and chain tool calls over many steps. You use it through an API with no GPUs or serving...
Excerpt shown — open the source for the full document.
Notability
notability 7.0/10New LLM release from Upstage