Structured outputs: prompt engineering for answers your code can trust
How I design LLM prompts and schemas that return reliable JSON for analytics, reporting and automation — plus the validation layer that catches everything else.
Chatting with a language model is easy. Building a system where a model's answer feeds directly into code — a report, a CRM update, a dashboard — is a different discipline. The difference is structure.
Most of my AI automation work at Mendoza Solutions comes down to one pattern: the model returns data in a defined shape, code validates it, and only valid data moves on.
Why free text breaks automations
Ask a model "Is this chart bullish or bearish?" and you might get "Bullish," "bullish.", "It appears moderately bullish, although..." or a three-paragraph essay. A human handles that fine. A parser does not.
The fix is to stop asking for prose when you need data.
Step 1: design the schema first
Before writing the prompt, write the output you want:
{
"trend": "bullish | bearish | neutral",
"confidence": 0.0,
"key_levels": [{"type": "support | resistance", "price": 0.0}],
"summary": "max 280 characters"
}
Every field has a type, allowed values and limits. This is the contract between the model and the rest of the system.
Step 2: enforce it with the platform and with code
Most current LLM APIs support some form of structured output or tool calling, where you hand the model a JSON Schema. Use it. Then validate anyway. In Python, Pydantic makes this short:
from typing import Literal
from pydantic import BaseModel, Field
class Level(BaseModel):
type: Literal["support", "resistance"]
price: float = Field(gt=0)
class ChartRead(BaseModel):
trend: Literal["bullish", "bearish", "neutral"]
confidence: float = Field(ge=0, le=1)
key_levels: list[Level] = []
summary: str = Field(max_length=280)
result = ChartRead.model_validate_json(raw_response)
If validation fails, the system can retry with the error message, fall back to a different model, or route the item to a human queue. What it never does is pass bad data downstream.
Step 3: write prompts like specs
Good prompts read like a specification for a careful contractor:
- Role and goal: what the output will be used for.
- Inputs: exactly what's provided and what isn't.
- Rules: "If the data is insufficient, set
trendtoneutraland say why insummary." - Examples: one or two complete, correct outputs.
The "what to do when unsure" rule is the most important line. Models fill silence with confidence. Give them a sanctioned way to say "I don't know."
Step 4: keep the model's job small
Multi-step workflows work better when each step does one thing. In my pipelines, a model might extract fields in one call, deterministic Python calculates in the next, and a model explains the result in plain language at the end. Math stays in code. Language stays in the model.
Step 5: log and evaluate
Every request, response and validation result is logged. A small set of reference inputs with known-good outputs gets re-run whenever a prompt or model changes. If you route between models with a tool like OpenRouter or LiteLLM, this is how you know a cheaper model is actually good enough.
The payoff
Structured outputs turned LLMs from a novelty into a dependable component for analytics, reporting and onboarding workflows. The model is still creative where it should be — summaries, explanations — and boring where it must be: types, enums and numbers that code can trust.
Some links in these notes are affiliate links. If you buy through one, I may earn a commission at no extra cost to you. I only link to tools I use or would recommend.