L4 · Oct 9, 2026
45% of our LLM calls failed: what the error strings told us
- llm-reliability
- structured-output
azure-openai
- testing
- case-study
About 45% of the screening evaluations in my batch returned nothing usable. Not a low fit score. No result.
This was early in building an AI CV screening platform at Metrodata. The batch had 22 candidate and job pairs: 2 job posts by 11 candidates.
The percentage told me how often the pipeline broke, not why. Reading the actual errors gave me two bugs I could work on.
Where the result was getting lost
The platform helps recruiters screen CVs. A recruiter defines a job post, the system extracts a weighted requirement rubric, and an LLM evaluates each candidate against the job. It returns structured JSON with a fit score, evidence and a written rationale. Today the recruiter makes the decision; the AI supplies evidence.
Each evaluation ran as an Azure Durable Functions activity. It called a reasoning model on Azure AI Foundry through the Responses API.
The parsing path was simple: strip markdown code fences from the reply, then call json.loads. If parsing worked, everything downstream worked. If it failed, that candidate and job pair was lost.
Two error strings, two places to look
"The LLM sometimes returns bad JSON" is not much of a diagnosis. It invites another round of prompt tweaks without telling you what to fix.
I grouped the failures by error message instead:
Response missing required keys: {'fitScore'}Invalid control character at: line 15 column 95
There were only those two strings. Two different messages usually point to two different mechanisms.
String one: the answer ran out of room
fitScore was the last key in the output schema. That position mattered. Randomly dropped fields would mean random missing keys. A consistently missing final key pointed to truncation: the answer stopping before it got there, rather than the model "forgetting" a field. That was my reasoning, not a measurement.
The configuration explained it. Reasoning effort was high, the prompt asked for 10 to 15 sentences of rationale, and I had not set an explicit output token ceiling.
On a reasoning model, hidden reasoning tokens and the visible answer share the output budget. Long reasoning plus a long rationale used it up mid-JSON. The last field never arrived.
That made reasoning effort a budget setting, not just a quality setting. The output ceiling needed to be explicit too.
There was also an easy prompt fix: I had told the model not to use code fences, then shown it a fenced example.
String two: a newline where JSON forbids it
Python's json module raises the second error in strict mode when a string value contains a raw control character.
The model had written literal newlines inside JSON string values instead of escaping them. The reply looked close to valid JSON, but json.loads could not accept it.
The layered fix
I stopped guessing at the prompt and mapped each error to its cause. No single change covered both, so the fix had several layers:
- JSON mode. I set
response_formattojson_object. The API constrains the reply to JSON rather than leaving that entirely to the prompt. - Explicit output ceilings. I set limits of 6,000 and 16,000 tokens for the two call types. The budget became a deliberate choice.
- Medium reasoning effort. I lowered it from high so less of the budget goes to hidden reasoning before the answer starts.
- A quote-aware JSON extractor. I replaced pattern-based fence stripping with an extractor that balances braces and tracks whether it is inside a string. A brace in a quoted value no longer ends the object early.
- Escape control characters inside strings only. Raw newlines and tabs get escaped inside string values. Between tokens they stay untouched, because they are legal whitespace there.
- Bounded retries with usage accounting. Each call gets up to three attempts. Token usage is summed across all attempts, so retries remain visible in the accounting. After the third failure, the error is re-raised for the Durable Functions retry policy to take over.
- Consistent prompt instructions. I removed the fenced example from the prompt that said "no fences".
Layer 5 looks roughly like this. This is a generic illustration, not the production code:
ESC = {"\n": "\\n", "\r": "\\r", "\t": "\\t"}
def escape_controls(raw: str) -> str:
out, in_str, esc = [], False, False
for ch in raw:
if esc:
esc = False
elif in_str and ch == "\\":
esc = True
elif ch == '"':
in_str = not in_str
elif in_str and ord(ch) < 0x20:
ch = ESC.get(ch, f"\\u{ord(ch):04x}")
out.append(ch)
return "".join(out)
The important part is in_str. Replacing every newline would also change valid whitespace between keys. Tracking quotes and escapes keeps the repair inside string values.
JSON mode reduces failures; it does not replace defensive parsing. The quote-aware extractor and bounded retries handle the rest. Retries also need to stay honest: count every attempt's tokens, cap the attempts, and re-raise the error so the next layer can decide what to do.
Turning the failures into tests
I shipped the fix with 7 unit tests that reproduce the exact error strings from the batch and run through the new parsing path. All 7 passed. The core test suite passed too, 12 of 12 at the time.
The platform had only started getting tests a few days earlier. These were among its first.
A synthetic malformed-JSON fixture covers what I imagine might go wrong. A fixture built from the model's actual reply covers what already did. It also fails loudly if someone later "simplifies" the parser back to json.loads. Saving the real failures as tests was cheap insurance against undoing the fix.
The number I do not have
My notes support the original failure rate, about 45% in that 22-pair run, the two error strings, their root causes, the layered fix, and the 7 passing regression tests.
They do not contain a rerun of the same 22 pairs after the fix. I cannot give you a "45% to X%" result for that batch. The claim I can make is narrower: I identified and fixed two failure mechanisms, with tests covering both.
Months later, the project had an offline evaluation harness. Its baseline measured a final parse failure rate, after retries, of 0.13%: 1 of 800 non-adversarial calls.
That used a newer model and a different prompt version. It tells me where reliability stood at the later baseline, not what this particular fix achieved. Putting the two percentages side by side would not make them a before-and-after measurement.
Parsing was not the last reliability problem
Three later problems on the same platform had a similar pattern:
- Retries that did not reliably fire. The Durable Functions retry I relied on in layer 6 was not reliable, and outcome counters could double count. The fix was a single retryable versus terminal classification, with the outcome marker and counter written in one transactional batch.
- Evals counting the wrong failures. Throttling responses, HTTP 429, were counted as schema failures. One eval run also measured an older code revision. Transport failures are now counted separately, and every run records a fingerprint of the code under test.
- Calls that quietly stalled. About 3% of model calls stalled until the SDK's default 600 s timeout. Explicit per-call deadlines now bound every call: 90 s for screening and 240 s for deep analysis.
Reasoning effort came up again later, this time for latency. In controlled evaluation rounds, screening p95 was 21.1 s at high effort and 12.1 s at medium. The fairness and prompt-injection gates still passed at medium. Medium shipped.
Before changing the prompt for a failing batch, I would read the error strings. Which field went missing? Where did the parser stop? In this run, the useful clues were the final key and a newline inside a string, not the percentage at the top.