# Evaluating AI-assisted threat detection: participant module

Kevin O'Connor. Version 1.0, September 9, 2026. A separate 90-minute exercise for analysts and detection engineers.

The decision is whether an assisted review suggestion preserves the intent of a detection rule and how the analyst should handle gaps. This module uses KevinBytes' public [identity_hunt.py](https://www.kevinbytes.com/examples/security-writing/identity_hunt.py), [recorded synthetic results](https://www.kevinbytes.com/examples/security-writing/identity-hunt-results.json) and [reference description](https://www.kevinbytes.com/examples/security-writing/README.md). The original recorded run contains eight cases, 35 rows and two combined matches. No Entra or Sentinel deployment, real attack dataset or AI model was tested.

## Baseline and scope

Compare the unchanged SQL rule in `identity_hunt.py` with a deliberately flawed, simulated suggestion, then review the differences as an analyst. No AI model produced the suggestion. This exercise practices review decisions; it does not measure productivity or model performance.

The synthetic schema is `(event_id, tenant, user_id, ip, time_s, kind)`. A sequence matches when there are at least three distinct failure records from the same tenant, user and IP in the 600 seconds before a success, followed by a role grant for the same tenant and user within 300 seconds after that success. The grant join does not require the same IP. Failure time is inclusive at success minus 600 and exclusive at success. Grant time includes the success timestamp and success plus 300.

The SQL uses `SELECT DISTINCT` over complete event rows. It removes identical duplicate deliveries; it does not resolve conflicting rows that reuse an event ID. That is an untested data-quality case for a real pipeline.

## Small labeled fixture

Time is invented seconds, with the usual success at 200. Most cases start with failures at 100, 101 and 102 and a role grant at 250. `events()` in the original script constructs these rows. Score each of the eight cases once. The 35 input rows are not separate classification examples.

| Case | Difference from the ordinary matching sequence | Observable-rule label |
| --- | --- | --- |
| positive_sequence | Three failures, success at 200, grant at 250 | 1 |
| ordinary_admin | No failures | 0 |
| two_failures | Only two distinct failures | 0 |
| late_grant | Grant at 501, outside the inclusive 500 upper bound | 0 |
| boundary_grant | Grant at 500, exactly on the included boundary | 1 |
| cross_tenant_join | Grant is in tenant-b, preceding records in tenant-a | 0 |
| duplicate_delivery | One failure is delivered three times identically | 0 |
| missing_success_telemetry | The success record is absent | 0 |

Label 1 means the supplied observable records meet this particular rule. Label 0 means they do not. The fixture contains no ground-truth attack labels. In particular, missing success telemetry cannot establish that the underlying activity was benign. Keep an independent operational disposition such as `telemetry gap; review required` alongside its zero rule-match label.

## Phase A: baseline review, minutes 0–35

Agree on the rule and one-case counting unit in the first 15 minutes. Predict each result, then run `python3 identity_hunt.py` and record the baseline by case. Compare the combined match identifiers with the public snapshot. Find the tenant join, exact time boundaries and identical-row deduplication in SQL. Explain the missing-success outcome before calculating scores.

Use these exact rules for every stage:

- Prediction 1 if that stage flags the named case as matching the rule; otherwise 0. Collapse any multiple returned tuples for the same case to one positive prediction.
- TP: label 1 and prediction 1. FP: label 0 and prediction 1.
- FN: label 1 and prediction 0. TN: label 0 and prediction 0.
- Count each of the eight cases exactly once. TP + FP + FN + TN must equal 8.
- Precision is TP/(TP+FP); recall is TP/(TP+FN). A zero denominator is undefined. These measures show how well predictions match the stated rule on these eight cases.

## Phase B: reveal the simulated suggestion, minutes 35–55

Record baseline results before reading this section. The suggestion below is deliberately imperfect teaching material, not model output:

> Flag positive_sequence, two_failures and missing_success_telemetry. Do not flag boundary_grant or the remaining cases. Two failures look suspicious enough; a missing success can be assumed from the later grant; the 300-second endpoint should be excluded.

Keep this in its own prediction column. Count its FP and FN results without changing the labels or rule to make it look correct. Explain which parts alter the threshold, invent missing data or misread the endpoint. Different detection hypotheses may be worth evaluating, but changing a rule requires a new specification and validation set.

## Phase C: analyst approval and escalation, minutes 55–75

The analyst reviews changed suggestions against actual rows and SQL, then records an explicit final rule decision and reason per case. An example final decision matches `positive_sequence` and `boundary_grant`, rejects the other rule-match suggestions and retains the missing-success case as an operational telemetry gap.

A flagged correlation goes to the authorized analyst for context and possible escalation; it is not an automatic account-disable command. The analyst records case ID, scope, source rows, rule/version, accepted or rejected suggestion, rationale, confidence limits, approving person and next owner. The synthetic module performs no containment action.

Do not close missing telemetry as clean. Assign the collection gap to its owner, identify the missing source/time range, request additional records through approved channels and set a review deadline. If proposing a new rule without the success requirement, document that as separate work instead of fabricating a success event.

## Phase D: reproduce and debrief, minutes 75–90

Open [the solution notes](detection-solutions.md) and run `python3 evaluate_detection.py` after downloading the scorer, JSON fixture and original identity script as described in [setup](README.md). Compare the counts with your worksheet. The scorer checks supplied predictions; it does not measure participants' independent judgment.

Discuss how incomplete collection, ingestion delay, clock skew, changed schemas, privileged but legitimate activity, conflicting duplicate IDs and cross-tenant identity collisions could affect an operational rule. Before evaluating a real assistant, agree on separately labeled data, data permissions, independent adjudication, a fixed baseline, held-out cases, time measurement and change approval. This workshop alone establishes none of those results.
