# Detection facilitator solution

Kevin O'Connor. Version 1.0, September 9, 2026. Facilitator copy; reveal after baseline and assisted review.

## Reproduce

Download and inspect the original [identity_hunt.py](https://www.kevinbytes.com/examples/security-writing/identity_hunt.py), this package's [evaluate_detection.py](evaluate_detection.py) and [detection-fixture.json](detection-fixture.json). Keep them in one directory:

```sh
python3 identity_hunt.py
python3 evaluate_detection.py
```

The original script asserts eight expected case results and two combined matches. Its 35 rows produce the boundary and positive sequences. The scorer runs that same local script, checks its case names and expected outcomes, then scores the supplied simulated and analyst predictions. It uses the standard library with no network, model, credentials or persistent database. An assertion failure or missing file must be resolved rather than reported as a passing reproduction.

## Expected counts

The observable-rule label is positive for `positive_sequence` and `boundary_grant`, negative for the other six. It is not an attack label. Each named case contributes exactly one count.

| Stage | Positive predictions | TP | FP | FN | TN | Precision | Recall |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Unchanged script baseline | positive_sequence, boundary_grant | 2 | 0 | 0 | 6 | 1.0 | 1.0 |
| Simulated suggestion | positive_sequence, two_failures, missing_success_telemetry | 1 | 2 | 1 | 4 | 1/3 | 1/2 |
| Example analyst-approved rule decision | positive_sequence, boundary_grant | 2 | 0 | 0 | 6 | 1.0 | 1.0 |

These counts are reproduced against the supplied fixture. They are not estimates of operational precision, attack recall, model quality or analyst performance. The example analyst decisions were prepared with the answers visible; they are not participant results. Matching the baseline counts shows that the original rule was restored. It does not demonstrate a benefit from assistance.

## Case-by-case rationale

| Case | Baseline | Suggestion | Example analyst | Rationale and disposition |
| --- | --- | --- | --- | --- |
| positive_sequence | 1 | 1 | 1 | All required records join within the windows. Send the correlation for authorized analyst context review. |
| ordinary_admin | 0 | 0 | 0 | No preceding failures. No rule match; this alone does not establish every aspect of the activity is benign. |
| two_failures | 0 | 1 | 0 | The suggestion lowers the required threshold without changing the rule specification. Reject this suggestion for the current rule. |
| late_grant | 0 | 0 | 0 | 501 is greater than 200 + 300. A different window needs a separately evaluated rule. |
| boundary_grant | 1 | 0 | 1 | 500 equals 200 + 300; `<=` includes the endpoint. Reject the suggestion's endpoint interpretation. |
| cross_tenant_join | 0 | 0 | 0 | The grant cannot join to the other tenant's sequence despite the same user identifier. |
| duplicate_delivery | 0 | 0 | 0 | Identical duplicate rows collapse to one failure; they do not satisfy the three-failure threshold. |
| missing_success_telemetry | 0 | 1 | 0 | The observed rule lacks a required success record. Do not invent one. Keep a separate telemetry-gap review open and assign its owner. |

For the simulated stage: `positive_sequence` supplies the TP; `two_failures` and `missing_success_telemetry` supply the two FPs; `boundary_grant` supplies the FN; the other four supply TNs. Calling the missing-success suggestion an FP is valid only for the stated observable-rule label. It says nothing about whether an unobserved real attack took place.

## False assurance and approval discussion

A quiet rule may mean a missing source, not absent suspicious activity. Distinguish “no rule match on supplied records” from “safe to close.” The analyst must not generate an event to fill a gap, silently move a tenant boundary, lower a threshold after seeing labels or execute account containment from a suggestion alone.

An acceptable approval record identifies the rule/version, input source and time range, case, suggestion, final determination, reason, source limitations, reviewer, next owner and review deadline. Escalate missing collection to the telemetry owner and material security concerns through the authorized response process. Real operational decisions need permitted corroboration and the organization's actual escalation policy.

For a more demanding future exercise, hide independent labels, add held-out positive and benign-but-matching cases, record review time and evaluate a specified real assistant under appropriate data controls. Do not reuse these eight teaching cases to claim generalized improvement.
