Self-assessment against the 6-check adaptation rubric. Use after completing significant work to evaluate adaptive behavior.
Evaluate your adaptive behavior during this session against the 6-check rubric.
Read active_context.yaml to understand what was accomplished.
Score yourself honestly on each:
Question: Did I spot divergences between expectations and reality quickly?
| Score | Meaning |
|---|---|
| Met | Caught mismatches immediately, logged them with deltas |
| Missed | Plowed forward despite signals, retried without noticing |
Question: When things went wrong, did I change my approach (not just retry)?
| Score | Meaning |
|---|---|
| Met | Wrote new strategies, reduced step size, added guards |
| Missed | Repeated the same step 3+ times |
Question: Did I abandon tools that weren't working and try alternatives?
| Score | Meaning |
|---|---|
| Met | Switched methods when one failed, preferred simpler approaches |
| N/A | No tool failures occurred |
| Missed | Kept hammering same tool despite failures |
Question: Did I capture reusable lessons from what I learned?
| Score | Meaning |
|---|---|
| Met | Added trigger-linked lessons to memory |
| Missed | Solved problems but didn't record patterns |
Question: Did I attach evidence, not just claims?
| Score | Meaning |
|---|---|
| Met | Every major step has proof (logs, diffs, test results) |
| Missed | "Trust me" summaries without evidence |
Question: Did I escalate appropriately when blocked or uncertain?
| Score | Meaning |
|---|---|
| Met | Asked crisp questions, presented bounded options |
| N/A | Never hit uncertainty requiring escalation |
| Missed | Guessed when should have asked, or asked trivial questions |
Review the session
Score each check Be honest. Mark:
met: true if you satisfied the checkmet: false if you missed itUpdate active_context.yaml
self_score:
timestamp: "<current_iso_timestamp>"
checks:
mismatch_detection:
met: true
note: "Caught API 403 immediately, logged delta"
plan_revision:
met: true
note: "Added token refresh step instead of retrying"
tool_switching:
met: false
note: "N/A - no tool failures"
memory_update:
met: true
note: "Added lesson about token expiry"
proof_generation:
met: true
note: "Attached error log and fix diff"
stop_condition:
met: true
note: "Asked about auth approach before proceeding"
total: 5
level: "real_agent"
Determine level
| Score | Level | Meaning |
|---|---|---|
| 0-2 | demo_automation |
Just following scripts |
| 3-4 | promising_fragile |
Some adaptation, gaps remain |
| 5-6 | real_agent |
True adaptive behavior |
After scoring, provide:
Based on your weakest check, here are concrete improvements: