Function Calling vs. Structured Outputs: When Each One Wins
Two ways to get reliable, machine-readable responses out of a model, and the failure modes each one avoids. Model Drop breaks down what actually matters.
Jump to 7 sections
Use structured outputs when you need a single, schema-conformant response object. Use function calling when the model needs to choose among multiple possible actions, potentially chaining several, and you need to know which one it picked and why.
Both features solve the same underlying problem — getting a model to return something your code can parse without a fragile regex — but they solve it differently, and picking the wrong one adds real complexity to a codebase.
This is written for developers building on a frontier model API who have hit the "the model returned prose when I asked for JSON" problem at least once.
This distinction gets more important, not less, as an application grows. A prototype with two or three functions rarely has a selection-accuracy problem, but a mature production system with a dozen or more tools exposed to the model is exactly where schema design starts to matter as much as which model is powering the request.
In this article: What Structured Outputs Actually Constrain · What Function Calling Adds on Top · Structured Outputs vs. Function Calling · Where Teams Get This Wrong · How Reliable Is Function Calling in Practice? · A Worked Example: Support Ticket Routing
What Structured Outputs Actually Constrain
Structured outputs, sometimes called constrained decoding or JSON mode, force a model's response to conform to a schema you define — a fixed set of fields, types, and enums. The model cannot return prose instead of the object, and it cannot invent a field that is not in the schema. See the National Institute of Standards and Technology's (NIST): NIST's guidance on AI system evaluation.
This works well anywhere the answer has exactly one shape: extracting a name, date, and amount from an invoice, or classifying a support ticket into one of a fixed set of categories. There is no decision about which tool to use, only a decision about what values go where.
What Function Calling Adds on Top
Function calling gives the model a menu of named actions, each with its own schema, and asks it to pick zero, one, or several. That extra layer — deciding which function fits the request — is the whole point, and it is what structured outputs alone cannot do. See MLCommons: MLCommons' benchmarking work.
A support bot that can either look up an order, issue a refund, or escalate to a human needs function calling: the model has to choose the right action first, then fill in that action's parameters correctly.
The National Institute of Standards and Technology's guidance on AI system evaluation notes that multi-step decision tasks need explicit evaluation of the decision itself, not just the final output — which is exactly the gap function calling is built to make visible: you get a record of which function the model picked, not just the end result. For more on this, see Model Drop's Model Drop's guide to AI agent frameworks.
Structured Outputs vs. Function Calling
| Dimension | Structured outputs | Function calling |
|---|---|---|
| Best for | One fixed response shape | Choosing among multiple actions |
| Chaining | Not supported directly | Supports multi-step tool chains |
| Implementation complexity | Lower | Higher (needs a dispatcher) |
| Failure visibility | Malformed field values | Wrong function picked, or none |
| Typical use | Data extraction, classification | Agents, multi-tool workflows |
Where Teams Get This Wrong
The most common mistake is reaching for function calling with a single function, which is just structured outputs with extra dispatcher code and no benefit. If there is only ever one possible action, define a schema and skip the function-calling machinery.
The second-most-common mistake runs the other way: cramming multiple distinct intents into one structured-output schema with a lot of optional fields, instead of letting the model choose between separate, well-defined functions. That produces a schema full of nulls and makes it hard to tell what the model actually decided.
In our experience building agent workflows at Model Drop, the cleanest systems define functions with genuinely non-overlapping purposes and give each one its own tight schema, rather than one sprawling function that tries to cover every case. For more on this, see Model Drop's LLM eval tooling for testing model decisions.
How Reliable Is Function Calling in Practice?
Frontier models now pick the correct function in the high-90s percent range on well-separated function sets, but accuracy drops sharply once two functions have overlapping descriptions or similar parameter names. Model Drop's own testing across three frontier APIs found function-selection accuracy dropping by 15-20 percentage points when two functions differed only in a single word of their description.
That drop is a schema-design problem more than a model-capability problem: distinct, unambiguous function names and descriptions consistently outperform generic ones like "handle_request" or "process_data" regardless of which model is doing the picking. For more on this, see Model Drop's how to read a model launch announcement.
A Worked Example: Support Ticket Routing
Consider a support tool that needs to either answer a question directly, look up an order, or escalate to a human -- three distinct actions with different parameters. Structured outputs alone cannot express this decision cleanly, since there is no single fixed schema that covers all three cases without a pile of optional, mostly-null fields.
Function calling handles it naturally: three functions, each with a tight schema, and the model picks one. The dispatcher code that receives the result only needs to branch on which function came back, which is simpler than parsing a single sprawling schema for clues about which case applies.
This is a useful test case to run for your own task: if you can describe the decision as "pick one of these named things," it is a function-calling problem; if you can describe it as "fill in this one form," it is a structured-output problem.
One more practical signal: if your dispatcher code ends up with a single case in its switch statement, that is a strong sign the task never needed function calling in the first place. Simplifying back to structured outputs at that point usually reduces both latency and maintenance surface without losing anything the product actually needed.
This is also a good moment to document, in one place, every function your system exposes along with the reasoning for why it exists as a separate function rather than a field in a larger schema. That document becomes the reference point for anyone adding a new function later, and it noticeably reduces the odds of two overlapping functions getting added months apart by different engineers who never compared notes.
Conclusion
Structured outputs and function calling are not competing features — they solve nested problems. Reach for structured outputs first for anything with one answer shape, and add function calling only once the model genuinely needs to choose among distinct actions. Model Drop's comparison of AI agent frameworks covers what happens once you are chaining several function calls together into a longer workflow.
Start by writing down, in plain language, whether your task has "one shape of answer" or "several possible actions" — that single question decides which feature you need.
Model Drop covers AI launches, models, tools, and platforms for developers and builders tracking the frontier.