Claude Code Agentic Evaluation & Root Cause Analysis
Evaluated Tier-1 LLM coding capabilities within the Claude Code CLI environment during a complex distributed system failure. Designed multi-turn agentic workflows to test model reasoning, tool usage, and patch generation under strict zero-downtime operational constraints on an Apache Pulsar cluster.
Role
AI Evaluator / Specialist
Target System
Apache Pulsar (Java)
Evaluation Focus
Agentic Capability Testing
Evaluation Level
Tier-1 US AI Platforms
Zero Downtime Operational Limits
The testing environment required evaluation under absolute operational constraints:
- No broker restarts
- No broker.conf edits
- No cursor resets
- No consumer app restarts
1. Evaluation Highlights & Model Performance Findings
A. Root Cause Analysis & Logic Reasoning
Prompted the model to analyze broker internal metrics (waitingReadOp: true, pendingReadOps: 0). Evaluated how accurately the model traced the issue to an unhandled race condition in PersistentDispatcherSingleActiveConsumer.java without generating hallucinated cursor states.
B. Constraint Adherence & Non-Invasive Recovery
Tested the model's ability to operate under strict zero-downtime operational limits. Guided the model to formulate a live recovery workaround using a zero-permit throwaway Failover consumer (receiverQueueSize=0), successfully triggering redeliverUnacknowledgedMessages() to clear orphaned request without disrupting active applications.
C. Script Generation & Code Quality
Evaluated the model’s capability in generating production-grade Python automation. Verified output quality for defensive engineering patterns, including --dry-run execution, real-time pulsar-admin topics stats-internal parsing, and proper consumer-slot placement.
D. Upstream Patch Isolation & Backporting
Assessed the model's precision in searching and isolating relevant upstream bug fixes across large repositories, successfully backporting Apache Pulsar PR #26174 (Failover stale read fixes) and PR #26236 (Key_Shared delivery stalls).
2. Incident Recovery & Resolution Workflow
Incident Diagnosis
Evaluated model output when reading pulsar-admin topics stats-internal to confirm the signature: waitingReadOp: true alongside availablePermits > 0 while consumer backlog grew.
Non-Invasive Live Remediation
Tested model reasoning against alternative fallbacks like dynamic config changes (unblockStuckSubscriptionEnabled=true). Executed the zero-permit temporary consumer trick designed during the evaluation turn to reset cursor state natively over wire protocol.
Codebase Patching & Verification
Verified model-generated cherry-picks onto local release branch (based on commit 8576283da4). Conducted static verification on test harness dependencies to prepare a deployment checklist for CI/CD integration testing.
3. Telemetry & Artifact Extracted
Utilized a CLI Trace Extractor proxy environment in tmux to capture model latency, tool invocation patterns, and multi-turn prompt telemetry. The python recovery script generated below showcases the successful remediation output.
# Generated via Claude Code agentic workflow evaluation
import pulsar
import time
import argparse
def trigger_zero_permit_consumer(service_url, topic, subscription):
print(f"[*] Starting non-invasive recovery for {subscription} on {topic}")
# 0-permit failover trick resets orphaned dispatcher states natively
client = pulsar.Client(service_url)
consumer = client.subscribe(
topic,
subscription,
consumer_type=pulsar.ConsumerType.Failover,
receiver_queue_size=0 # Strict zero-permit limitation
)
# Let connection establish and broker flush redeliverUnacknowledgedMessages
time.sleep(2.0)
consumer.close()
client.close()
print("[+] Orphaned waitingReadOp successfully flushed without restart.")Interested in AI Evaluation & Engineering?
Explore my other technical case studies or full engineering resume.