1. Core Architecture & Migration Context
Anthropic’s developer migration guide for Claude Opus 5.5 moves away from legacy integration patterns. While your existing Opus 5 prompts will still execute, direct porting introduces performance degradation, token inflation, and unexpected agent stalls.
The Opus 5.5 optimization model is organized around "failure modes"—diagnostic error states mapping directly to parameter adjustments and official drop-in prompt primitives.
The 7 Essential Terminology Shifts
effort(Compute Tier / Reasoning Depth): Controls pre-computation inference depth. Higher tiers scale execution latency and token expenditure.max_tokens(Output Ceiling): Maximum token generation limit per invocation. Crucial: Internal reasoning tokens count against this ceiling regardless of whether the reasoning block is rendered to the UI.stop_reason(Execution Terminator): API state return label. Examples includeend_turn(successful handover),tool_use(execution pending), andrefusal(safety constraint violation).thinkingBlock: Isolated structural array containing internal pre-computation logic, decoupled from user-facingtextstreams.- Prompt Caching: Server-side ephemeral caching for static system instructions and leading context. Global mid-session parameter or tool-definition mutations invalidate the cache prefix.
- Harness (Agent Orchestration Runtime): Backend scheduling logic. Manages round-trips, executes CLI/database tools, and drives the agent step loop.
- Subagent (Delegation Mesh): Isolated child instances spawned by a primary agent for parallelized task execution.
2. Technical Highlights: Speed, Cost, and Reasoning Overhaul
Performance & Default Scaling
Opus 5.5 generates output tokens >30% faster than Opus 5 while consuming fewer total tokens per completed task. However, under identical effort configurations, Opus 5.5 computes deeper than Opus 5 (particularly at the xhigh and max tiers).
- Default Tier Downgrade: The default
effortshifts fromhightomedium. Opus 5.5mediumachieves parity with (or exceeds) Opus 5highbenchmarks in programming and knowledge workloads, rendering legacy configurations excessively slow and expensive. - Permanent Reasoning Layer: The legacy
thinking: {"type": "disabled"}override is deprecated. Reasoning is always active; the model determines its own inference depth dynamically. - Token Allocation Rule: Long-running agentic coding tasks require
max_tokensscaling up to 128,000 to account for hidden internal reasoning overhead.
Thinking Display & Observability Matrix
thinking.display Value |
Behavioral Output | UI Rendering Impact |
|---|---|---|
"omitted" (Default) |
Inference executes, tokens are billed, but the thinking payload field is empty. |
Zero-progress renders during long tasks; UI appears frozen. |
"summarized" |
Returns a structural summary of the internal reasoning block. | Direct drop-in replacement for prompts that historically forced internal reasoning into the text stream. |
"updates" (Beta) |
Emits incremental status summaries between tool executions (Requires header: thinking-display-updates-2026-08-18). |
Displays live operational updates ("Discovered X, executing Y") to mitigate user abandonment. |
3. Practical Tradeoffs & Symptom-Driven Troubleshooting
Symptom 1: Unattended Agents Stall After Progress Reports
- Root Cause: The harness misinterprets conversational status updates as a terminal completion state rather than an intermediate execution step.
- Resolution: Configure your harness loop to intercept text output as a progress signal, verifying completion against the task manifest before terminating the loop (limit autonomous retries to 2–3 iterations).
Symptom 2: Requests Trigger stop_reason: "refusal"
- Root Cause: Legacy prompts explicitly instruct the model to write its internal reasoning or chain-of-thought directly into the main text response.
- Resolution: Remove chain-of-thought generation commands from system prompts. Rely on
thinking.display: "summarized"for transparent reasoning payloads.
Symptom 3: Prompt Injection Vulnerabilities via Pasted Inputs
- Root Cause: Unsanitized user inputs containing control tokens or instruction overrides ("Ignore previous instructions...").
- Resolution: Implement system-level boundary demarcation using explicit XML framing to isolate user-supplied payloads from execution directives.
4. Quickstart / Verdict & Best Practices
- Drop Obsolete Prompt Patterns: Strip redundant "Think carefully" instructions from system prompts. Opus 5.5 manages allocation autonomously via the
effortparameter. - Optimize Tier Selection: Transition production workloads from
hightomediumto capitalize on the 5.5 speedup without incurring latency penalties. - Harden Harness Loops: Ensure your backend state machine evaluates agent return values against strict execution criteria rather than relying on natural language termination cues.
