Roger Mendoza

Structured outputs: prompt engineering for answers your code can trust

How I design LLM prompts and schemas that return reliable JSON for analytics, reporting and automation — plus the validation layer that catches everything else.

Chatting with a language model is easy. Building a system where a model's answer feeds directly into code — a report, a CRM update, a dashboard — is a different discipline. The difference is structure.

Most of my AI automation work at Mendoza Solutions comes down to one pattern: the model returns data in a defined shape, code validates it, and only valid data moves on.

Why free text breaks automations

Ask a model "Is this chart bullish or bearish?" and you might get "Bullish," "bullish.", "It appears moderately bullish, although..." or a three-paragraph essay. A human handles that fine. A parser does not.

The fix is to stop asking for prose when you need data.

Step 1: design the schema first

Before writing the prompt, write the output you want:

{
  "trend": "bullish | bearish | neutral",
  "confidence": 0.0,
  "key_levels": [{"type": "support | resistance", "price": 0.0}],
  "summary": "max 280 characters"
}

Every field has a type, allowed values and limits. This is the contract between the model and the rest of the system.

Step 2: enforce it with the platform and with code

Most current LLM APIs support some form of structured output or tool calling, where you hand the model a JSON Schema. Use it. Then validate anyway. In Python, Pydantic makes this short:

from typing import Literal
from pydantic import BaseModel, Field

class Level(BaseModel):
    type: Literal["support", "resistance"]
    price: float = Field(gt=0)

class ChartRead(BaseModel):
    trend: Literal["bullish", "bearish", "neutral"]
    confidence: float = Field(ge=0, le=1)
    key_levels: list[Level] = []
    summary: str = Field(max_length=280)

result = ChartRead.model_validate_json(raw_response)

If validation fails, the system can retry with the error message, fall back to a different model, or route the item to a human queue. What it never does is pass bad data downstream.

Step 3: write prompts like specs

Good prompts read like a specification for a careful contractor:

  • Role and goal: what the output will be used for.
  • Inputs: exactly what's provided and what isn't.
  • Rules: "If the data is insufficient, set trend to neutral and say why in summary."
  • Examples: one or two complete, correct outputs.

The "what to do when unsure" rule is the most important line. Models fill silence with confidence. Give them a sanctioned way to say "I don't know."

Step 4: keep the model's job small

Multi-step workflows work better when each step does one thing. In my pipelines, a model might extract fields in one call, deterministic Python calculates in the next, and a model explains the result in plain language at the end. Math stays in code. Language stays in the model.

Step 5: log and evaluate

Every request, response and validation result is logged. A small set of reference inputs with known-good outputs gets re-run whenever a prompt or model changes. If you route between models with a tool like OpenRouter or LiteLLM, this is how you know a cheaper model is actually good enough.

The payoff

Structured outputs turned LLMs from a novelty into a dependable component for analytics, reporting and onboarding workflows. The model is still creative where it should be — summaries, explanations — and boring where it must be: types, enums and numbers that code can trust.

Some links in these notes are affiliate links. If you buy through one, I may earn a commission at no extra cost to you. I only link to tools I use or would recommend.