# Facilitator guide and run sheet

Kevin O'Connor. Version 1.0, September 9, 2026. Instructor preparation and session plan.

## Before the session

Confirm which session is booked: the 180-minute agent workshop, the separate 90-minute detection module, or both with a separately scheduled break. Ask for one decision the participants need to make and a redacted workflow sketch. Do not request a production credential or client dataset. Name an application owner, identity owner, security reviewer and decision owner; people may hold more than one role.

Send the [participant packet](participant-packet.md), [worksheets](worksheets.md) and [setup index](README.md). Run all selected commands yourself. Keep terminal output for teaching and check that downloaded file versions match the solution notes. If a command fails, resolve the mismatch before teaching or use an explicitly identified source walkthrough. Never call a source walkthrough a successful runtime check.

Open a blank worksheet for each pair. Reserve the final decision block even if mapping runs long. Have one person operate the local fixture and the other predict and record the result before execution, then swap roles. Do not reward the number of failures found; reward the accuracy of the reasoning.

## Core agent workshop run sheet

| Elapsed time | Instructor steps | Participant work and checkpoint |
| --- | --- | --- |
| 0:00–0:20 | Read the synthetic assistant scenario. Ask the workflow owner to name an allowed output and an excluded action. Establish that the workshop authorizes no live access. | Complete the scope sheet. Everyone can name the actor, tenant, source and effect. |
| 0:20–0:45 | Draw user → application identity → retrieval/context → tool policy → document store. Add trusted reviewer, credential broker and downstream-call records. Ask who supplies each identity. | Mark where data crosses a trust boundary and which checks use current policy. Note the fixture's trusted identity assumptions. |
| 0:45–1:00 | Open `document_permissions.py` and the permitted read/write tests. Ask the pair to predict the result before running the suite. | Record a successful read and an approved write; identify subject, document/version, body and reviewer. |
| 1:00–1:20 | Assign altered body, replay, revoked reviewer and stale version tests to pairs. Use the named unittest methods from the solution notes for focused execution. | Record denial condition and the unchanged document state. Explain why an old approval cannot authorize a changed edit. |
| 1:20–1:30 | Break. Collect unresolved implementation assumptions. | Leave one question on the worksheet. |
| 1:30–1:45 | Open the three RAG documents. Have pairs predict Amber and Birch search results and a forged context candidate. | Show that legitimate access works and the other tenant's body cannot enter context. |
| 1:45–2:05 | Assign ACL revocation, citation leakage, stale version and deleted-source tests. Read the inert hostile note; ask what actually prevents sending. | Inspect citation titles/URLs. Name the missing send capability and absence of any model call. Write an additional cache test for a real application. |
| 2:05–2:20 | Run `credential_lifecycle.py`. Trace a permitted request, revocation and retry. Contrast wrong audience with insufficient scope. | Record status and downstream-call count. Explain why replacing a fixture handle says nothing about a provider token. |
| 2:20–2:30 | Introduce a tabletop condition: a worker already holds a cached token and a request is in flight. No code execution is needed. | Assign connector stop, provider revocation, queue review and verification owners. Record unresolved timing. |
| 2:30–2:45 | Ask each pair to choose one blocker and one allowed pilot use. Require a concrete acceptance check and a responsible owner. | Complete the decision record. Avoid declaring a real application ready from fixture tests. |
| 2:45–3:00 | Exchange worksheets, apply the rubric and discuss the strongest unsupported assumption. | Revise one weak claim. Agree on the next test and review date. |

## Failure discussions

If everyone denies every action, return to the allowed read/write checks. A boundary that prevents the intended workflow is not a useful implementation. If a participant equates an empty result with safety, ask whether the source was permitted, the query could match, and telemetry was complete. If the hostile note becomes the focus, return to the boundary: the code does not include a language model and exposes no send tool.

For a test failure, record Python version, filename, method name, assertion and observed state. Stop interpreting the reference counts until the cause is understood. For a disagreement about a real system, label it an open question and assign the right owner; do not fill the gap with a vendor claim.

## Detection module run sheet

Use [the detection module](detection-module.md) as the participant handout. Keep the solution and simulated suggestion columns hidden until their scheduled reveal.

| Elapsed time | Instructor steps | Participant output |
| --- | --- | --- |
| 0:00–0:15 | Define the observable-rule label, one-case unit, tenant join and inclusive/exclusive time boundaries. Distinguish labels from attack truth. | Written counting rule and case predictions. |
| 0:15–0:35 | Run unchanged `identity_hunt.py`; compare eight cases and two combined matches with the recorded snapshot. | Baseline predictions, TP/FP/FN/TN and an explanation of the missing-success case. |
| 0:35–0:55 | Reveal the simulated suggestion in `detection-module.md`. It contains deliberate errors and was not generated by a model. | Keep a separate suggestion column; write provisional acceptance or rejection for each changed case. |
| 0:55–1:15 | Ask for analyst adjudication. Introduce pressure to close the missing-telemetry case as clean; require an escalation decision. | Final rule-match decisions, separate telemetry-gap disposition, owner and approval record. |
| 1:15–1:30 | Reveal `detection-solutions.md` and run the scorer. Discuss threshold and boundary errors plus what a real evaluation would require. | Corrected worksheet and next validation scope. |

## Evaluation and follow-up

Apply the rubric to submitted materials, not participant confidence. Ask: Which result was actually observed? Which test would change the decision? Which operator can stop the effect? What would invalidate the scope? For detection, also ask why a correct zero match is compatible with unknown attack status.

Return the completed worksheet and a short list of agreed owners. Any custom assessment, remediation, production validation or later delivery remains separately scoped work.
