Auto Mode: The AI Judges Its Own Risk (Experimental)

This is the most cutting-edge part of the permissions system, controlled by Feature Flag TRANSCRIPT_CLASSIFIER and currently available only to internal Anthropic employees.

Core Idea

Instead of having users manually judge whether each operation is safe, use another Claude to judge.

Claude A (main agent): I want to execute rm -rf ./build
  ↓
Claude B (classifier): Analyzes conversation history + this command → verdict: safe / unsafe
  ↓
safe   → Auto-execute, don't interrupt the user
unsafe → Show confirmation dialog with explanation

Preventing an Overly Strict Classifier

The Auto Mode AI classifier may be too conservative, causing the user to be interrupted too often. The code contains a denial tracker (denialTracking.ts) as a safety valve:

// src/utils/permissions/denialTracking.ts
const DENIAL_LIMITS = {
  maxConsecutive: 3,   // 3 consecutive denials
  maxTotal: 20,        // 20 total denials
}

// Over the limit → automatically fall back to manual prompting
function shouldFallbackToPrompting(state): boolean {
  return (
    state.consecutiveDenials >= DENIAL_LIMITS.maxConsecutive ||
    state.totalDenials >= DENIAL_LIMITS.maxTotal
  )
}

Product logic: If the AI classifier denies 3 times in a row or 20 times cumulatively, that signals the classifier's judgment may be unreliable for this scenario, and the system automatically reverts to human confirmation. This is a classic circuit breaker design.