Auto Mode: The AI Judges Its Own Risk (Experimental)
This is the most cutting-edge part of the permissions system, controlled by Feature Flag TRANSCRIPT_CLASSIFIER and currently available only to internal Anthropic employees.
Core Idea
Instead of having users manually judge whether each operation is safe, use another Claude to judge.
Claude A (main agent): I want to execute rm -rf ./build
↓
Claude B (classifier): Analyzes conversation history + this command → verdict: safe / unsafe
↓
safe → Auto-execute, don't interrupt the user
unsafe → Show confirmation dialog with explanation
Preventing an Overly Strict Classifier
The Auto Mode AI classifier may be too conservative, causing the user to be interrupted too often. The code contains a denial tracker (denialTracking.ts) as a safety valve:
// src/utils/permissions/denialTracking.ts
const DENIAL_LIMITS = {
maxConsecutive: 3, // 3 consecutive denials
maxTotal: 20, // 20 total denials
}
// Over the limit → automatically fall back to manual prompting
function shouldFallbackToPrompting(state): boolean {
return (
state.consecutiveDenials >= DENIAL_LIMITS.maxConsecutive ||
state.totalDenials >= DENIAL_LIMITS.maxTotal
)
}
Product logic: If the AI classifier denies 3 times in a row or 20 times cumulatively, that signals the classifier's judgment may be unreliable for this scenario, and the system automatically reverts to human confirmation. This is a classic circuit breaker design.