Claude Code Effort Benchmark: When to Use Low and When to Scale Up

A hands-on engineering guide to navigating the new effort parameters in Claude Code, balancing execution cost, token latency, and code accuracy across real-world workflows.


1. Core Architecture: What Does "Effort" Actually Change?

The effort parameter instructs the model on how much dynamic compute to allocate to a given task. It is neither a hard time-out nor a simple output token limiter.

When you scale the effort parameter, you are modifying the agent's internal compute budget and behavioral loops:

  • Low Effort: Prioritizes fast time-to-first-token and immediate execution paths. The model moves quickly to generate a baseline that meets the explicit prompt requirements, minimizing autonomous exploration and self-correction cycles.
  • High/Max Effort: Expands the agentic loop. Claude allocates additional compute cycles to proactively research internal context, test alternative code paths, run self-audits on its own implementations, and edge-case test requirements that weren't explicitly enumerated in the prompt.

Think of it like allocating developer resources: giving a human engineer 1 hour yields an early iteration for design feedback, whereas giving them 12 hours yields a hardened, thoroughly tested, and independently polished feature.


2. Technical Highlights: Behavioral Matrix Across Tiers

As you toggle the effort slider, Claude's execution strategy shifts across the lifecycle of a task:

Effort Tier Primary Objective Verification & Edge Cases Developer Workflow Synergy
Low Fast baseline generation; quick directional alignment. Minimal self-correction; defers edge-case validation to external test suites. Ideal for early-stage prototyping, rapid iterations, and tight feedback loops.
Medium Balanced execution; structured problem-solving. Moderate checks; verifies core logic against standard failure modes. Good for standard feature additions where specs are partially defined.
High Deep implementation with proactive refactoring. Comprehensive self-review; tests non-obvious failure modes and system boundaries. Suited for mission-critical logic, performance optimization, and architectural scaffolding.
Max Maximum autonomy; exhaustive implementation and polish. Extensive exploratory testing, boundary handling, and autonomous product decisions. High token latency/cost; best when you want zero-touch end-to-end delivery of well-defined specs.

Note: Changing effort tiers within Claude Code adjusts the model's compute allocation dynamically without invalidating or breaking your active prompt cache.


3. Practical Benchmarks: How Requirement Clarity Dictates Tier Selection

To evaluate how effort translates into wall-clock time and output variance, Thariq Shihipar ran comparative benchmarks across three distinct engineering scenarios using Opus 5.5.

(Note: Execution times reflect single-run observations to illustrate tradeoffs rather than strict, reproducible SLAs.)

Scenario A: Ambiguous Requirements (Open-Ended Exploration)

  • Prompt: "Build a personal fitness and workout tracking app."
  • Execution Variance:
    • Low (~1.5 min): Delivered a foundational tracker with basic historical logging and simple charts.
    • Medium (~4 min): Added structured intermediate views and session management.
    • High (~11 min): Built a more robust application structure with integrated state handling.
    • Max (~67 min): Delivered an expansive, near-production-ready app complete with heatmaps, navigation frameworks, and proactive UI/UX decisions.
  • Engineering Takeaway: At higher effort levels, Claude spends significant compute making product decisions on your behalf. If requirements are fluid, a Low effort baseline lets you iterate on direction faster than waiting for a Max-effort speculative build.

Scenario B: Design Exploration (UI/UX Refactoring)

  • Prompt: "Redesign the Claude Code /config menu."
  • Execution Variance:
    • Low (~1 min): Generated functional interactive wireframes that established a clear conceptual direction, though structurally distinct from existing patterns.
    • Max (~28 min): Produced polished prototypes tightly aligned with the core product's design system, complete with multi-path interaction walk-throughs.
  • Engineering Takeaway: For early structural exploration, Low effort wins by providing rapid directional feedback. Max effort is only efficient once the design paradigm is locked in and you require high-fidelity execution.

4. Verdict: A Four-Step Framework for Production Workflows

To optimize token spend, latency, and engineering velocity, implement this four-step selection process for your daily development cycles:

  1. Phase 1: Exploration & Scaffolding (Use Low)
    • When: Exploring greenfield ideas, mocking up initial architectures, or testing raw feasibility.
    • Why: Keeps token latency low and prevents the model from over-engineering unvalidated assumptions.
  2. Phase 2: Rapid Iteration (Use Low to Medium)
    • When: Refining features where you are actively driving the feedback loop.
    • Why: Fast iteration cycles matter more than autonomous edge-case handling during active co-design.
  3. Phase 3: Implementation of Clear Specs (Use Medium to High)
    • When: Requirements are fully specified, API contracts are locked, and logic needs to be written right the first time.
    • Why: Allows the agent to catch standard bugs and enforce internal consistency.
  4. Phase 4: Hardening & Deep Verification (Use High to Max)
    • When: Building security filters, storage engine primitives, payment flows, or mission-critical code paths.
    • Why: This is where the highest effort tiers pay off—exercising boundary conditions, self-auditing implementations, and catching the corner cases your prompt missed.